CoT推理未必不透明,缺提示词不等于不忠实。
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- 用因果中介分析证明:未说出口的提示也能影响推理结果
- 大推理预算下提示词出现率升至90%,原判不忠多因字数限制
- 建议结合因果与破坏性测试,别只看提示词是否出现
近期研究使用偏差特征(Biasing Features)指标,若链式思维(CoT)未包含影响预测的提示词,则判定为不忠实。我们指出该指标对忠实性的定义过于狭窄,混淆了不忠实与不完整——即把分布式变换器计算压缩成线性自然语言叙述时的必然损失。在指令微调和推理模型的多跳推理任务中,许多被该指标标记为不忠实的CoT,在其他指标下仍被判定为忠实,部分模型超过50%。通过新提出的faithful@k指标,我们发现更大的推理预算显著提升提示词的出现率(某些情况下达90%),表明多数看似不忠实的表现实因严格令牌限制所致。利用因果中介分析进一步证明,即使未显式表达的提示也能通过CoT因果性地影响预测。因此我们警告不应仅依赖提示词评估,并倡导采用更广泛的可解释性工具包,包括因果中介分析和破坏性指标。我们并非声称所有CoT都忠实,而是强调仅缺失提示词无法证明不忠实。
原文摘要 · Abstract (English)
Recent work, using the Biasing Features metric, labels a CoT as unfaithful if it omits a prompt-injected hint that affected the prediction. We argue this metric adopts a narrow notion of faithfulness and confuses unfaithfulness with incompleteness, the lossy compression needed to turn distributed transformer computation into a linear natural language narrative. On multi-hop reasoning tasks with instruct-tuned and reasoning models, many CoTs flagged as unfaithful by Biasing Features are judged faithful by other metrics, exceeding 50% in some models. With a new faithful@k metric, we show that larger inference-time budgets greatly increase hint verbalization (up to 90% in some settings), suggesting much apparent unfaithfulness is due to tight token limits. Using Causal Mediation Analysis, we further show that even non-verbalized hints can causally mediate prediction changes through the CoT. We therefore caution against relying solely on hint-based evaluations and advocate a broader interpretability toolkit, including causal mediation and corruption-based metrics. We do not claim all CoTs are faithful, only that the absence of hint words alone does not prove unfaithfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。