Open Source RadarNotícias G
Open Source RadarNotícias
Open Source RadarInteligência de Notícias Entrar Projetos
Todos os eventos
Pesquisa Pesquisa Médio prazo 1 matérias

Strategically Diverse Sampling for Self-Training

Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform…

Evento consolidado
Gerenciar alertas
Análise do Radar

O que aconteceu e por que importa

Inteligência do evento

Por que este sinal merece atenção

Comparar com Em Alta
85Relevânciaforça do sinal no contexto atual
34Tendênciavelocidade e recorrência do movimento
100Novidadequanto o sinal adiciona informação nova
Não identificadoImpacto Brasilabrir contexto nacional
Primeiro sinal25/09/2026 17:35
Último sinal25/09/2026 17:35
0horas em evolução
1fontes distintas
Evidências

Timeline do evento

1 matéria(s)
25/09 17:35
arXiv cs.CLGlobal
Strategically Diverse Sampling for Self-Training
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform…
Abrir fonte original
Receba o Radar Diário grátis
Todo dia às 8h: as notícias e os projetos open source de IA que importam, com contexto em português e a análise completa em PDF.
Prefere começar pelo PDF de hoje? Baixe o panorama.