Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks

AK's Picks LLM Other LLM AI Safety / AI Ethics PIAAP
随着大语言模型能力的持续提升,其被滥用的潜在风险也在同步加剧。封闭源代码模型通常依赖外部防御机制,而开源权重模型则主要须依靠内置安全防护措施来抑制有害行为。以往的红队测试研究大多聚焦于基于输入的越狱攻击(jailbreaking)和参数层面的操控手段。然而,开源权重模型原生支持“预填充”(prefilling)功能,即攻击者可在模型生成响应前预先设定初始输出词元(tokens)。尽管这一攻击路径潜力巨大,却长期缺乏系统性研究与关注。本文开展了迄今为止规模最大的预填充攻击实证研究,针对多个主流模型家族及当前最先进的开源权重模型,全面评估了二十余种既有及新型预填充攻击策略。实验结果表明,预填充攻击对所有当前主流的开源权重模型均表现出稳定且显著的有效性,暴露出一个关键性、此前未被充分认识的安全漏洞,对模型的实际部署具有重大影响。尽管某些大型推理模型对通用型预填充攻击展现出一定鲁棒性,但面对量身定制、针对特定模型设计的预填充策略时,仍显脆弱。我们的研究发现凸显出一个紧迫需求:开源权重大语言模型的开发者必须将抵御预填充攻击的防护能力建设置于优先地位。
As the capabilities of large language models continue to advance, so does their potential for misuse. While closed-source models typically rely on external defenses, open-weight models must primarily depend on internal safeguards to mitigate harmful behavior. Prior red-teaming research has largely focused on input-based jailbreaking and parameter-level manipulations. However, open-weight models also natively support prefilling, which allows an attacker to predefine initial response tokens before generation begins. Despite its potential, this attack vector has received little systematic attention. We present the largest empirical study to date of prefill attacks, evaluating over 20 existing and novel strategies across multiple model families and state-of-the-art open-weight models. Our results show that prefill attacks are consistently effective against all major contemporary open-weight models, revealing a critical and previously underexplored vulnerability with significant implications for deployment. While certain large reasoning models exhibit some robustness against generic prefilling, they remain vulnerable to tailored, model-specific strategies. Our findings underscore the urgent need for model developers to prioritize defenses against prefill attacks in open-weight LLMs.
许愿