用相同检验结果测试模型对因果问题的判断能力,发现隐藏状态可解码证据与目标的匹配关系。
Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States
- 设计配对提示,同一证据不同因果目标,检验模型理解能力
- 在9类诊断任务中成功识别18-21组正确匹配,准确率约65.4%-65.9%
- 隐藏状态信息比答案概率或文本基线更能反映证据与因果目标的关系
相同的诊断结果可能支持或反驳一个因果主张,却对另一个无关,当这些主张涉及不同人群、结果、估计量、作用路径或识别假设时。当证据与目标同时变化,正确答案可能源于有利措辞、词汇重叠或熟悉模式,而非真正匹配证据与因果问题。我们引入成对提示:重复使用相同诊断证据,仅改变因果目标。每条提示标注为“支持”、“挑战”、“未解决”或“错误目标”,只有当一对提示均被正确分类时才视为成功恢复。利用在独立开发集上训练的线性读出器,分析Qwen2.5-7B-Instruct、Qwen3-8B和Llama-3.1-8B-Instruct的倒数第二层Transformer块的最终标记隐藏状态。在包含49对样本的主基准上,平衡准确率为0.654至0.659,成功恢复18-21对。两名独立人类评审员对98条提示中的95条达成一致(96.9%)。各检查点的平衡准确率与完整对恢复率均超过保持开发场景分组的置换零模型。在Qwen2.5中,全提示平衡准确率高于受限输入,且双侧自助区间均大于零。在未使用评估诊断家族开发样例的读出器中,仍恢复21对,涵盖全部九类。隐藏状态读出器在平衡准确率和恢复对数上优于基于答案选项逻辑值和文本基线的线性分类器。结果表明,隐藏状态中存在可线性解码的信息,能判断诊断证据是否支持、挑战或不适用于特定因果目标。
原文摘要 · Abstract (English)
The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions. When the evidence and target vary together, a correct answer may reflect favorable or adverse wording, lexical overlap, or a familiar diagnostic pattern rather than matching the evidence to the causal question. We introduce paired prompts that repeat the same diagnostic evidence verbatim while changing the causal target. Each prompt is labeled Favors, Challenges, Unresolved, or Wrong Target according to how the evidence bears on the causal question. A pair is recovered only when both prompts are classified correctly. Using linear readouts trained on a separate development set, we analyze the final-token hidden state from the penultimate transformer block of Qwen2.5-7B-Instruct, Qwen3-8B, and Llama-3.1-8B-Instruct. On the 49-pair primary benchmark spanning nine diagnostic families, balanced accuracy ranges from 0.654 to 0.659 and 18-21 pairs are recovered. Two independent human reviewers assigned the same label to 95 of the 98 prompts (96.9%). Across checkpoints, balanced accuracy and complete-pair recovery exceed permutation nulls that preserve development scenario groups. In Qwen2.5, full-prompt balanced accuracy exceeds both restricted inputs, with paired-bootstrap intervals for both differences above zero. Readouts trained without development examples from the evaluated diagnostic family recover 21 pairs, including at least one in each of the nine families. The hidden-state readout exceeds a linear classifier on answer-option logits and text baselines in balanced accuracy and recovered pairs. These results show that the hidden state contains linearly decodable information about whether diagnostic evidence favors, challenges, or fails to address the causal target.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。