PAPER

Harness Learning Enables Generalizable Test-Time Adaptation

Agent Agent Planning Autonomous Workflows ML RL ML
语言模型智能体由其基础模型及其“运行框架”(harness)共同定义;该运行框架是一个可执行程序,负责统筹模型调用、工具使用以及信息流转。由于不同任务对上述操作的组织方式要求各异,运行框架必须依据当前任务所提供的反馈进行动态适配。为此,我们提出“运行框架学习”(harness learning)方法:训练一个“提议模型”(proposer model),使其能基于执行反馈对求解器(solver)的运行框架进行修订。我们将这一过程建模为面向可执行程序的元学习(meta-learning),其中运行框架的修订扮演了梯度式适应中权重更新的角色。我们采用强化学习训练该提议模型,以修订后运行框架在具体任务上的性能作为奖励信号。在测试阶段,提议模型无需更新任何参数,仅需利用新任务上连续多次执行所获得的反馈,即可逐步优化并精炼运行框架。在推理任务与多跳问答任务上的实验表明,运行框架学习显著提升了修订质量,且这种在测试阶段即时自适应的能力可泛化至未见过的新任务。此外,在单次修订样本上训练所得的策略,能够通过多轮迭代持续改进运行框架;而针对修订序列整体进行训练所带来的增益,则因具体任务设置不同而有所差异。这些发现为构建持续学习型智能体指明了一条可行路径——此类智能体可将过往积累的经验转化为具有泛化能力的、可持续演进的系统性提升。
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.
许愿