Bridging Compute- and Data-Optimal Pretraining

LLM PTSAD ML Other ML
经典计算最优缩放定律假设预训练数据的供给是无限的,然而预训练正日益进入一个计算资源增长速度超过高质量数据可获得性的新阶段。为此,我们提出“计算–数据”(CD)缩放定律——一种统一框架,弥合了两类极端情形:一类是计算最优缩放(即数据量可随计算资源自由增长),另一类是数据最优缩放(即语料库规模固定,而计算资源可无限制增加)。CD缩放定律在经典缩放定律基础上进行了拓展,引入了一个“词元有效性函数”(token-effectiveness function)$η$,用以量化一个衍生词元(例如通过多轮遍历训练或改写生成的词元)相对于一个全新词元的实际价值;该值介于“完全等效替代”与“毫无价值”之间。我们基于Dolma-3语料库,在模型参数量从1400万至6亿的范围内,针对两种数据扩展策略——多轮遍历重复与文本改写——分别拟合了$η$函数。结果表明,词元有效性远非恒定不变:它同时依赖于模型规模、每参数对应词元数(tokens-per-parameter ratio)以及衍生数据的总量,并且随着语料库不断扩展而趋于饱和。$η$函数的具体形式揭示:当以计算资源替代数据时,其边际收益递减——无论模型规模增大,还是数据供给增加,均呈现这一规律。此外,该框架将训练过程划分为三个运行区间:计算受限区、数据受限区和模型受限区;并进一步指出,在绝大多数实际相关场景中,传统计算最优的资源分配方式并非最优解。
Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $η$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $η$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $η$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.
许愿