3 points, 0 comments on Hacker News
Source: [Hacker News](https://www.bloomberg.com/news/articles/2026-07-29/creator-of-test-that-openai-models-tried-to-cheat-sounds-alarm)
Most CI pipelines assume a function called with the same input twice returns the same output. That assumption breaks the moment an LLM call enters your test suite. Ask GPT-4 or Claude the same question twice and you can get two different (both correct) answers.
Building Multi-Model AI Applications: Why One Model Won't Be Enough If you ask most developers what model powers their AI application, you'll usually get one answer. "We're using GPT-4o. " Or Claude.
1 points, 0 comments on Hacker News
1 points, 0 comments on Hacker News
1 points, 0 comments on Hacker News
1 points, 0 comments on Hacker News