Open Source RadarNotícias G
Open Source RadarNotícias
Open Source RadarInteligência de Notícias Entrar Projetos
Todos os eventos
Pesquisa Pesquisa Médio prazo 1 matérias

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or earl…

Evento consolidado
Gerenciar alertas
Análise do Radar

O que aconteceu e por que importa

Inteligência do evento

Por que este sinal merece atenção

Comparar com Em Alta
85Relevânciaforça do sinal no contexto atual
34Tendênciavelocidade e recorrência do movimento
100Novidadequanto o sinal adiciona informação nova
Não identificadoImpacto Brasilabrir contexto nacional
Primeiro sinal25/09/2026 17:59
Último sinal25/09/2026 17:59
0horas em evolução
1fontes distintas
Evidências

Timeline do evento

1 matéria(s)
25/09 17:59
arXiv cs.AIGlobal
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or earl…
Abrir fonte original
Receba o Radar Diário grátis
Todo dia às 8h: as notícias e os projetos open source de IA que importam, com contexto em português e a análise completa em PDF.
Prefere começar pelo PDF de hoje? Baixe o panorama.