PAPER

Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

LLM PTSAD ML LM
预预训练(PPT)在合成的非自然语言数据上进行,可提升语言模型预训练(PT)阶段的词元利用效率。以往研究将这一增益归因于“语法先验”——即模型在PPT阶段习得的一种结构性归纳偏置,该偏置可迁移到自然语言的语法规则中。然而,此前PPT仅在参数量至多为10亿(1B)的模型上开展验证,且预训练预算低于20亿(2B)词元,所用数据也以网络文本为主。目前尚不清楚:当模型规模进一步扩大、预训练数据构成更贴近真实场景(例如混合代码、数学等多样化来源)时,PPT是否依然有效。为此,我们开展了迄今最全面的PPT实证研究,系统考察了五种PPT任务、四种预训练数据组合、四种模型参数规模(5亿至70亿,即500M–7B),以及最高达1000亿(100B)词元的预训练预算。结果表明,PPT带来的下游任务性能提升与词元利用效率增益在更大规模下依然稳健——例如,在30亿(3B)参数规模下,至少可节省210亿(21B)预训练词元。但与先前研究结论相反,我们并未发现一致证据支持这些增益源于“语法先验”:下游性能在不同模型规模下,并未与语法可接受性指标呈现稳定一致的相关性。相反,我们发现,真正带来下游性能提升的PPT任务,是那些能增强长程信息检索能力的任务。此外,PPT的性能增益对预训练数据组合方式具有较强鲁棒性,仅当训练数据中完全不含网络文本时,其增益才显著减弱。总体而言,PPT是一种低成本、高效益的预训练补充手段;未来PPT任务的设计应聚焦于提升长程检索能力,而非刻意模拟自然语言语法。
Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.
许愿