用因果方法提升大模型开发与评估的可靠性
Causal methods for LLM development and evaluation

- 通过因果推断分析数据域、风格变化等干预对模型的影响
- 发现现有评估依赖有偏判官,导致结果不可靠
- 适合关注模型可解释性与科学验证的研究者
大语言模型(LLM)开发目前依赖大规模经验迭代,涉及数据混合、奖励模型、路由策略和评估流程。我们指出,许多核心问题本质上具有因果性:预训练中加入新数据域有何影响?当模型生成不同风格文本时,标注者偏好如何变化?在推理成本约束下,应将提示路由到大模型还是小模型?尽管因果方法适用于此类干预改变结果的场景,但在LLM开发中仍被严重低估。本文贡献三方面:(1)阐明因果方法如何改善现代LLM开发与评估——开发依赖日志数据,常受混杂与分布漂移影响;评估使用可能有偏的学得判官;部署环境非平稳。这些条件使纯预测方法脆弱,而因果推断提供了更稳健的识别与估计框架。(2)系统映射了因果方法在预训练、对齐、路由、智能体工作流和评估全流程中的应用机会。(3)探讨利用因果方法推动未来研究的新方向。总体而言,我们主张因果方法在LLM开发与评估中被潜在低估,但能确保更可靠、科学的设计。
原文摘要 · Abstract (English)
Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and evaluation pipelines. Here, we argue that many central questions in LLM development and evaluation are inherently causal: What is the effect of adding a data domain during pretraining? How do annotator preferences change when LLMs generate text in a different style? Should a prompt be routed to a larger or smaller model given inference cost constraints? In general, causal methods are well-suited to such settings where interventions change outcomes but, surprisingly, are underrepresented in LLM development. Our contribution is threefold: (1) We explain how causal methods can help develop modern LLM development and evaluation: LLM development relies heavily on logged data, which are often subject to confounding and distribution shifts; evaluation uses learned but potentially biased judges; and deployment environments are non-stationary. These conditions make purely predictive approaches fragile and create opportunities for principled identification and estimation methods from causal inference. (2) We further map opportunities for causal methods in the entire LLM development pipeline, including pretraining, alignment, routing, agentic workflows, and evaluation. (3) We discuss new research opportunities around leveraging causal methods for LLM development and evaluation. Overall, we argue that causal methods are potentially underutilized for the LLM development and evaluation pipeline, despite the fact that such methods can ensure a reliable and scientifically grounded design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。