arXiv:2509.11208stat.MLcs.LG2025-09被引 2

发现证据顺序影响模型判断,提出可量化可靠性的新方法

Predictable Compression Failures: Order Sensitivity and Information Budgeting for Evidence-Grounded Binary Adjudication

  • 将证据顺序视为干扰因素,建立预期与实际差距的数学模型
  • 提出三类指标:可信度、幻觉风险和信息充足率,指导是否回答
  • 在多个数据集验证,顺序打乱后幻觉率降至0.7%以下,支持精准决策

用于证据支撑型二元判定(如支持/反驳、是/否或验证器背书的通过/失败判断)的Transformer模型对可交换证据的呈现顺序敏感,导致不同排列下结果分散,且在验证器相对的伯努利判别下产生不可靠答案。本文将证据顺序视为干扰变量,形式化了期望-实现差距:下一词训练可最小化所有顺序下的条件描述长度期望,而固定顺序仍受位置影响。提出的量化鞅违规(QMV)界预测了相邻秩位置敏感引发的分散,其增长为$O(\log n)$阶;期望级解压缩定律(EDFL)将KL凸性/数据处理界特化至伯努利判别,导出比特转信任(B2T)、幻觉风险(RoH)及信息充分率(ISR)门控机制,用于决定回答或弃权。在来自FEVER、HotpotQA、NQ-Open、PopQA和对照组的3,059个有依据样本上,观察到对数级分散,并从均匀排列混合中获得正向詹森收益。在一次预设保留审计(528项)中,解析固定的ISR=1门控实现了0.0%-0.7%幻觉率,同时伴随20.6%-27.9%弃权率(95%置信区间),支持该操作点,但不主张跨所有模型族或无限制生成的普遍校准。

原文摘要 · Abstract (English)

Transformers used for evidence-grounded binary adjudication (e.g., support/refute, yes/no, or verifier-backed pass/fail decisions) can be sensitive to the order in which exchangeable evidence is presented, producing dispersion across permutations and unreliable attempted answers under a verifier-relative Bernoulli predicate. We treat evidence order as a nuisance variable and formalize an expectation-realization gap: next-token training can minimize expected conditional description length over orderings while a fixed ordering remains position-sensitive. Our Quantified Martingale Violation (QMV) bound predicts the dispersion induced by adjacent-rank positional sensitivity, with $O(\log n)$ growth in the harmonic regime; our Expectation-level Decompression Law (EDFL) specializes a KL convexity/data-processing bound to Bernoulli predicates, yielding Bits-to-Trust (B2T), Risk-of-Hallucination (RoH), and an Information Sufficiency Ratio (ISR) gate for answer/abstain decisions. On 3,059 grounded items from FEVER, HotpotQA, NQ-Open, PopQA, and Controls, we observe logarithmic dispersion and positive Jensen gains from uniform permutation mixtures. In one pre-specified held-out audit (528 items), the analytically fixed ISR$=1$ gate attains 0.0-0.7% hallucination with 20.6-27.9% abstention (95% CIs), supporting the operating point without claiming universal calibration across all model families or unrestricted generation.

模型可靠性证据排序幻觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。