A new study has compared five proper scoring rules as training objectives for large language model (LLM) forecasters on real-world binary events.
The research found that reward choice affects not just overall forecasting accuracy, but also calibration, discrimination, and the structure of forecasting errors.
These findings suggest that different scoring rules can shape LLM forecasting profiles in distinct ways, with implications for model selection and ensemble design.