为什么重要

该研究揭示了 RLVR 训练中推理轨迹入口处的关键瓶颈,为改进强化学习算法和提升模型推理能力提供了重要见解。

关键事实

事实 1

RLVR 导致 Countdown 任务的解题覆盖下降最多 67%

来源与依据

单一来源

Across both training setups, solution coverage falls by up to 67%

Hugging Face Daily Papers · 第一方证据 · 支持

查看 Hugging Face Daily Papers 原文

事实 2

首次算术运算之前的每 token 似然偏移比下游推理大 11 至 16 倍

来源与依据

单一来源

per-token likelihood shifts are 11x--16x larger prior to the first arithmetic operation

Hugging Face Daily Papers · 第一方证据 · 支持

查看 Hugging Face Daily Papers 原文

事实 3

仅提供未选择的入口前缀即可将低入口族的完成率从 0.018 提升到 0.212

来源与依据

单一来源

Supplying only an unselected entrance prefix restores completion rates in low-access families by over an order of magnitude (0.018 -> 0.212 under PPO)

Hugging Face Daily Papers · 第一方证据 · 支持

查看 Hugging Face Daily Papers 原文