Updated
Updated · arxiv.org · Sep 25
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
Updated
Updated · arxiv.org · Sep 25

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

1 articles · Updated · arxiv.org · Sep 25

Summary

  • Researchers have developed a new self-supervised training method, ConfSFT, that improves the efficiency of reasoning in language models without explicitly optimizing for shorter outputs.
  • ConfSFT trains models to predict their own confidence at intermediate reasoning steps, leading to up to 25% fewer generated tokens while maintaining accuracy across multiple benchmarks.
  • This approach suggests that teaching models to assess their confidence can naturally lead to more efficient reasoning, potentially reducing computational costs in AI applications.