arXiv:2506.00152cs.LGecon.EM2025-06被引 1

用观察数据微调大模型时,需消除混淆因素才能避免错误学习。

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective

  • 提出DeconfoundLM,从奖励信号中剥离已知混淆因素影响。
  • 实验证明直接使用观察数据会学到虚假关联,导致性能下降。
  • 适合想低成本利用历史数据优化模型但又担心偏差的研究者。

大型语言模型在生成内容以提升转化率等关键指标方面广泛应用,但预训练模型常难以对齐人类偏好或业务目标。为此,高质量标注数据的微调至关重要。尽管控制实验(如A/B测试)可提供此类数据,但成本高且实施困难。企业却拥有大量未被充分利用的历史观察数据。本文研究了基于观察数据微调大模型的挑战与机遇。我们发现,虽然观察结果能提供有用监督信号,但直接微调可能导致模型学习虚假相关性。通过多个真实世界数据集的实证分析,我们提出DeconfoundLM方法,显式消除已知混淆因素对奖励信号的影响。模拟实验表明,该方法能更好恢复因果关系,并缓解忽略或粗略处理混淆变量的方法所导致的失效问题。研究强调:若采取正确的因果修正,观察数据可成为大模型对齐的强大信号源。

原文摘要 · Abstract (English)

Large language models are being widely used across industries to generate content that contributes directly to key performance metrics, such as conversion rates. Pretrained models, however, often fall short when it comes to aligning with human preferences or optimizing for business objectives. As a result, fine-tuning with good-quality labeled data is essential to guide models to generate content that achieves better results. Controlled experiments, like A/B tests, can provide such data, but they are often expensive and come with significant engineering and logistical challenges. Meanwhile, companies have access to a vast amount of historical (observational) data that remains underutilized. In this work, we study the challenges and opportunities of fine-tuning LLMs using observational data. We show that while observational outcomes can provide valuable supervision, directly fine-tuning models on such data can lead them to learn spurious correlations. We present empirical evidence of this issue using various real-world datasets and propose DeconfoundLM, a method that explicitly removes the effect of known confounders from reward signals. Using simulation experiments, we demonstrate that DeconfoundLM improves the recovery of causal relationships and mitigates failure modes found in fine-tuning methods that ignore or naively incorporate confounding variables. Our findings highlight that while observational data presents risks, with the right causal corrections, it can be a powerful source of signal for LLM alignment. Please refer to the project page for code and related resources.

大模型对齐因果推断观察数据去混淆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。