Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
Updated
Updated · arxiv.org · Oct 2
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
1 articles · Updated · arxiv.org · Oct 2
Summary
A new audit reveals that large language model (LLM) judges often misalign with human workers when assessing AI-generated workplace outputs.
While LLM judges can rank responses similarly to humans, they frequently misestimate acceptance rates and occupational performance aggregates, sometimes by large margins.
Researchers urge that LLM judges be validated against the actual acceptance rates and aggregates they are used to estimate, not just ranking agreement.