Last month I shipped an eval that ranked two prompt variants. Variant A won by four points. A teammate reran the same eval the next morning and Variant B won.
Source: [Dev.to](https://dev.to/maya_andersson_dev/an-llm-judge-is-a-biased-instrument-not-a-measurement-2b15)