16 local model configurations, 56 hidden-test coding tasks, 36 full runs, one 16 GB card Every number below was recounted from the committed SCORES-*. tsv files, not transcribed from notes. The raw data - 36 rows, one per run, each with its full failure list - is RESULTS-q56.
Source: [Dev.to](https://dev.to/mrdushidush/a-75b-model-beat-a-24b-on-my-coding-benchmark-30o4)