Updated
Updated · arxiv.org · Oct 2
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
Updated
Updated · arxiv.org · Oct 2

Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

1 articles · Updated · arxiv.org · Oct 2

Summary

  • A new audit reveals that large language model (LLM) judges often misalign with human workers when assessing AI-generated workplace outputs.
  • While LLM judges can rank responses similarly to humans, they frequently misestimate acceptance rates and occupational performance aggregates, sometimes by large margins.
  • Researchers urge that LLM judges be validated against the actual acceptance rates and aggregates they are used to estimate, not just ranking agreement.