提出统一框架,发现参数与上下文忠实性可协同优化但不对称。
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

- 设计统一接口,同时优化参数与上下文忠实性
- 参数优化能跨范式提升,上下文优化效果不稳定
- 现有上下文指标衡量不同方面,存在内在权衡
链式思维(CoT)忠实性指其是否真实反映大模型的内在行为,通常在两种分离范式下评估:通过扰动输入或CoT轨迹衡量上下文忠实性,通过干预模型参数知识评估参数忠实性。以往研究仅做描述性比较。本文提出FaithMATE,一种统一的偏好对齐接口,用于优化模型向任一忠实性范式靠拢。该方法使我们得以探究两范式间的相互作用,检验忠实性提升是否能在范式内或跨范式泛化。在三个模型、两个数据集、六个忠实性度量下,发现两范式呈正向耦合但不对称:优化参数忠实性可在两范式中稳定提升,而上下文优化则表现更不稳定;在同一范式内,某一度量的提升无法一致传递到其他度量,表明现有上下文指标捕捉的是忠实性的不同侧面,揭示了内在权衡。结果表明,CoT忠实性并非单一目标,需多维度优化与评估。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint paradigms: contextual faithfulness, measured by perturbing the input or CoT trace, and parametric faithfulness, assessed by intervening on a model's parametric knowledge. Yet prior work compares them only descriptively. We fill this gap by proposing FaithMATE, a unified preference-alignment interface for optimizing models towards either faithfulness paradigm. It enables us to investigate the interplay between the two paradigms, examining whether and to what extent faithfulness gains generalize within and across paradigms. Across three models, two datasets, and six faithfulness metrics, we find that the two paradigms are positively coupled, yet asymmetric: optimizing towards parametric faithfulness yields consistent gains across both paradigms, whereas the contextual counterpart delivers more variable gains. Within the contextual paradigm, faithfulness gains on one metric do not consistently transfer to others, implying that existing contextual metrics capture disjoint facets of faithfulness and exposing inherent trade-offs. These findings imply that CoT faithfulness is not a monolithic objective and therefore requires multifaceted optimization and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。