Originally published at nlqdb. com/blog Our small NL-to-SQL benchmark — twenty questions, one gold query each, scored by executing the SQL and comparing result sets — came back 17/20 on a greedy pass. An immediate second run on the same commit, drawing three samples per question instead of one, ...

Source: [Dev.to](https://dev.to/omer_hochman/your-offline-llm-eval-isnt-measuring-your-model-its-measuring-your-rate-limits-2ph0)

Sponsored