We talk about agents drifting. We almost never talk about the thing we measure them against drifting. But your golden dataset — the fixtures, expected outputs, and "known good" traces your evals grade against — is code that ships to production and then never gets a code review again.
Source: [Dev.to](https://dev.to/saurav_bhattacharya/your-golden-dataset-is-rotting-the-eval-oracle-nobody-re-validates-4id3)