用信息论解释合成数据为何有效,关键在是否引入外部反馈信号。
An Information-Theoretic Criterion for Efficient Data Synthesis
- 通过信息开放/封闭循环判断合成数据有效性
- 粗粒度反馈(如对错)能提升泛化能力
- 简单信号易引发奖励黑客,需警惕伪相关
合成数据在大语言模型训练中日益重要,但其效果差异巨大。本文从信息论角度解释这一现象:只有当生成-训练循环是信息开放的(即由验证器、环境或评分标准等外部信号注入任务相关知识),合成数据才能提升模型性能;若循环为信息封闭(仅依赖模型自身输出),则根据数据处理不等式,任务相关信息只会衰减,导致性能崩溃。在信息开放的流程中,效率与泛化能力取决于监督的元层级:粗粒度信号(如二值正确性)将所有可接受输出视为等价,所学习的行为不绑定特定领域或表面形式,因而天然具备跨任务和跨领域的泛化能力。由此得出核心论点:学习会优先收敛到当前最信息高效的信号成分;若该成分正是目标,则加速学习;若为偶然出现的简化解,则引发奖励黑客。
原文摘要 · Abstract (English)
Synthetic data becomes crucial for large language model training, but its effectiveness is highly inconsistent. We provide an information-theoretic account of this inconsistency: synthetic data improves a model only when the generation-training loop is information-open, i.e., shaped by external signals (verifiers, environments, or rubrics) that inject task-relevant information beyond the model's current distribution. When the loop is information-closed (relying on the model's own outputs without such signals), the data processing inequality ensures that task-relevant information can only decrease, making collapse a predicted outcome. Among information-open pipelines, both efficiency and generalization hinge on the meta-level of supervision: a coarser signal such as binary correctness treats all acceptable outputs as equivalent, so the behavior it teaches is not tied to any particular domain or surface form and generalizes naturally across tasks and domains. These observations lead to a guiding thesis: learning preferentially converges to the most information-efficient signal component available, which accelerates learning when that component is the intended one, but causes reward hacking when a spurious pattern happens to be simpler.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。