PAPER

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

LLM SFT Other LLM Agent Agent Planning Other Agents ML RL IOCIL
在线策略蒸馏(OPD)是一种极具前景的语言智能体训练方法,它能为学生模型生成的轨迹提供密集的教师监督信号。然而,在多轮交互场景中,一旦学生采取了错误动作,便会改变其后续所处的状态,导致错误在各轮次间不断累积。我们在三款Qwen3模型(参数量分别为8B、72B与235B)上开展的初步实验发现:超过半数失败的推理轨迹中均存在一种“关键性错误”——即某个动作不仅未能推进任务进展,反而使智能体离任务完成目标更远;且此类关键性错误通常发生在交互早期。值得注意的是,这些关键性错误往往仍具备可恢复性:只要在关键错误发生后的若干轮次内对模型加以正确引导,即可使其重新回到成功完成任务的路径上。为此,我们提出PivotOPD——一种新型在线策略蒸馏框架,旨在联合训练学生模型,使其既能在关键节点前主动规避关键性错误,又能在错误发生后有效从其所导致的不良状态中恢复。具体而言,每当学生出现关键性错误时,教师模型不仅为其提供该步的黄金动作(gold action),还会在随后若干轮次中逐一指定“恢复动作”(recovery action)。其中,“预防式蒸馏”(preventive distillation)利用黄金动作,通过逆向KL散度(reverse KL)引导学生模型远离关键性错误;而“恢复式蒸馏”(recovery distillation)则借助恢复动作,采用正向KL散度(forward KL),将学生模型极少采样的、有效的恢复行为迁移至其策略中。在ALFWorld、WebShop和基于搜索的问答(Search-based QA)三大基准任务上,面对13种基线方法,PivotOPD在Qwen3-1.7B与Qwen3-8B两类学生模型上均取得最强的平均性能表现;其中,Qwen3-1.7B学生模型在ALFWorld上的性能较最强基线提升达+5.5%。此外,该方法的增益效果还可迁移到其他模型家族——在软件工程领域,PivotOPD将Nemotron-3.5学生模型在SWE-Bench Verified基准上的问题解决率提升了+3.2%。项目主页:https://research.nvidia.com/labs/lpr/pivotopd/
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
许愿