不匹配≠错误:参考集不全会颠倒模型校准排名,影响可信度判断。
Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking

- 用有限参考集判断思维理论输出真伪,会因遗漏真实信念导致误判。
- 在6个场景中,参考集重编码使模型置信度排名反转,误差变化达0.227至0.152。
- 提出可修复校准问题的三源恢复框架,适用于需高可信度推理的系统。
开放式思维理论(ToM)追踪器会生成有限参考集中未包含的有效信念。采用有限参考集加匹配器的流程将未匹配输出标记为虚假,由此产生的代理标签可能在固定输出下颠倒正确得分的模型选择顺序。保持259个信念及其对应评分不变,参考集重编码使加权出现率从0.783降至0.295,且逆转了严格适当的布里尔风险:在参考标签下,原生置信度领先0.227;而在盲审判决下则落后0.152,六个作者场景均如此。仅基于参考集的Platt校准器进一步加剧反转。在已发布的301题NQ-open DPR-BERT流水线中,平均置信度基线在精确匹配下使实例级校准误差降低0.045,但在人类正确性标准下反而恶化0.074,两者区间均不包含零。在独立撰写的OpenToM叙事中,90%-96%被审计的未匹配信念为真实,且关联方向再次反转。精确分解表明失真是由遗漏真实信念所致,闭式准则能准确区分十二个发布系统的表现。冻结审计回放显示,50次尝试标注中,至少99.6%的概率可恢复正确排名方向。TriSource-Restore通过概率采样人类先验,锚定全帧参考标签与冻结自动判断,维持至少名义覆盖率,缩小置信区间,并在基础率部署门控下修复置信度。
原文摘要 · Abstract (English)
Open-ended Theory-of-Mind (ToM) trackers emit valid beliefs absent from finite references. A finite-reference-plus-matcher pipeline marks unmatched outputs false, creating proxy labels that can reverse proper-score model selection on fixed outputs. Holding 259 beliefs and paired scores fixed, reference recoding lowers weighted prevalence from 0.783 to 0.295 and reverses strictly proper Brier risk: a frozen source-prior rule leads native confidence by 0.227 under reference labels and trails by 0.152 under blinded adjudication, in all six authored scenarios. A reference-only Platt recalibrator reverses further. An ICE-specific reversal appears in a released 301-question NQ-open DPR-BERT pipeline: its average-confidence baseline improves instance-level calibration error by 0.045 under exact match but worsens it by 0.074 under human correctness, with both intervals excluding zero. On independently authored OpenToM narratives, 90-96% of audited unmatched beliefs are literally true and the paired direction again reverses. An exact decomposition attributes the distortion to omitted truths, and a closed-form criterion correctly classifies comparisons from twelve released systems. Frozen-audit retrospective replay shows 50 attempted annotations recover ranking direction with probability at least 0.996. TriSource-Restore anchors full-frame reference labels and frozen automatic judgments to a probability-sampled human pilot, maintains at least nominal coverage, narrows intervals, and repairs confidence subject to a base-rate deployment gate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。