Here's an output from a RAG system asserting a pricing claim it was never given, for a question its context couldn't answer. I ran it past the two most popular LLM-as-judge faithfulness metrics, five times each, judges on gpt-4o at temperature 0 — the setting most favourable to judge stability. ...

Source: [Dev.to](https://dev.to/nickjlamb/we-gave-the-same-fabricated-answer-to-ragas-and-deepeval-one-scored-it-00-the-other-scored-it-10-57om)

Sponsored