发现证据顺序影响模型判断,提出可量化可靠性的新方法
Predictable Compression Failures: Order Sensitivity and Information Budgeting for Evidence-Grounded Binary Adjudication
- 将证据顺序视为干扰因素,建立预期与实际差距的数学模型
- 提出三类指标:可信度、幻觉风险和信息充足率,指导是否回答
- 在多个数据集验证,顺序打乱后幻觉率降至0.7%以下,支持精准决策
用于证据支撑型二元判定(如支持/反驳、是/否或验证器背书的通过/失败判断)的Transformer模型对可交换证据的呈现顺序敏感,导致不同排列下结果分散,且在验证器相对的伯努利判别下产生不可靠答案。本文将证据顺序视为干扰变量,形式化了期望-实现差距:下一词训练可最小化所有顺序下的条件描述长度期望,而固定顺序仍受位置影响。提出的量化鞅违规(QMV)界预测了相邻秩位置敏感引发的分散,其增长为$O(\log n)$阶;期望级解压缩定律(EDFL)将KL凸性/数据处理界特化至伯努利判别,导出比特转信任(B2T)、幻觉风险(RoH)及信息充分率(ISR)门控机制,用于决定回答或弃权。在来自FEVER、HotpotQA、NQ-Open、PopQA和对照组的3,059个有依据样本上,观察到对数级分散,并从均匀排列混合中获得正向詹森收益。在一次预设保留审计(528项)中,解析固定的ISR=1门控实现了0.0%-0.7%幻觉率,同时伴随20.6%-27.9%弃权率(95%置信区间),支持该操作点,但不主张跨所有模型族或无限制生成的普遍校准。
原文摘要 · Abstract (English)
Transformers used for evidence-grounded binary adjudication (e.g., support/refute, yes/no, or verifier-backed pass/fail decisions) can be sensitive to the order in which exchangeable evidence is presented, producing dispersion across permutations and unreliable attempted answers under a verifier-relative Bernoulli predicate. We treat evidence order as a nuisance variable and formalize an expectation-realization gap: next-token training can minimize expected conditional description length over orderings while a fixed ordering remains position-sensitive. Our Quantified Martingale Violation (QMV) bound predicts the dispersion induced by adjacent-rank positional sensitivity, with $O(\log n)$ growth in the harmonic regime; our Expectation-level Decompression Law (EDFL) specializes a KL convexity/data-processing bound to Bernoulli predicates, yielding Bits-to-Trust (B2T), Risk-of-Hallucination (RoH), and an Information Sufficiency Ratio (ISR) gate for answer/abstain decisions. On 3,059 grounded items from FEVER, HotpotQA, NQ-Open, PopQA, and Controls, we observe logarithmic dispersion and positive Jensen gains from uniform permutation mixtures. In one pre-specified held-out audit (528 items), the analytically fixed ISR$=1$ gate attains 0.0-0.7% hallucination with 20.6-27.9% abstention (95% CIs), supporting the operating point without claiming universal calibration across all model families or unrestricted generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。