arXiv:2605.26442cs.CLcs.AI2026-05ACL

从数据视角重构大模型对齐流程,揭示关键设计陷阱。

Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines

论文配图:Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines
图 1 · 摘自论文原文
  • 将对齐数据构建拆解为三阶段:生成回复、评估偏好、实例化偏好
  • 发现现有方法普遍存在的权衡与失效模式,提炼出优化信号设计原则
  • 适合研究对齐数据构建或想理解对齐失败原因的从业者

多数对齐调优研究聚焦于优化目标,而对齐数据的构建常被隐含处理。本文从数据中心视角出发,将对齐调优重新定义为流水线设计问题。我们将对齐数据构建分解为三个相互作用的阶段:响应生成、偏好评估和偏好实例化,并以此框架将现有对齐方法统一归类。通过这一视角,我们识别出跨方法的常见设计权衡与失败模式,提炼出影响最终优化信号的高层原则。最后,我们指出了对齐数据流水线的开放挑战,包括提示级对齐、代理场景下的对齐以及目标动态变化时的对齐问题。

原文摘要 · Abstract (English)

Much of the alignment tuning literature is organized around optimization objectives, while the construction of alignment data is often treated implicitly. In this survey, we adopt a data centric perspective and reframe alignment tuning as a pipeline design problem. We decompose alignment data construction into three interacting stages, response synthesis, preference evaluation, and preference instantiation, and use this framework to organize existing alignment methods into a unified taxonomy. Through this lens, we identify recurring design trade-offs and failure modes observed across prior alignment methods, and distill a set of high level principles that clarify how pipeline design choices influence the resulting optimization signal. Finally, we outline open challenges for alignment data pipelines, including prompt-level alignment, agentic settings, and alignment under evolving objectives.

对齐方法数据流水线大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。