Your fine-tune gains three points of acc_norm on HellaSwag and loses two points of acc . Same checkpoint, same harness, same seed. Nothing about the model's commonsense reasoning moved in two directions at once — you changed its average per-token entropy, and one of those two metrics is partly ...

Source: [Dev.to](https://dev.to/ji_ai/acc-vs-accnorm-why-length-bias-skews-llm-eval-scores-p6e)

Sponsored