How to build an app where two models get the same CSV, write matplotlib code, get sandboxed and executed, and a third model judges the actual rendered output — measured, not vibes. The problem with AI comparisons Ask two models the same question and you get two confident answers. Which is better?

Source: [Dev.to](https://dev.to/harishkotra/graphite-making-two-llms-draw-their-own-comparison-13p1)

Sponsored