We ran seven Google models through the same task, in the same multi-agent harness. Improvements to reasoning and tool calls are the main reasons for the improvements. We saw less and less reasoning and some of the lowest invalid tool calls across the entire evaluation suite.

Source: [Hacker News](https://github.com/awesoftsolutions/google_gemini_flash_comparison)

Sponsored