arXiv:2608.21559cs.CL2026-08

提出证据可靠性评估框架,揭示大模型流水线中结构正确却证据失效的问题。

Evidence-State Reliability Under Controlled Degradation: Parser-Validity Divergence in a Multi-Stage LLM Pipeline

论文配图:Evidence-State Reliability Under Controlled Degradation: Parser-Validity Divergence in a Multi-Stage LLM Pipeline
图 1 · 摘自论文原文
  • 设计独立于结构校验的证据可靠性评估层,关注证据完整性与可用性。
  • 在9种退化条件下,阶段成功率全为负值,而结构合规性仍为正。
  • 首次分离出证据退化检测与恢复能力,适合系统安全与可信推理研究者。

多阶段大模型流水线在下游阶段面对不完整、压缩或冲突的证据时,仍可能保持结构有效。本文提出并实证了证据状态可靠性(ESR)评估框架,关注中间证据是否足够完整、有据、内部一致且可用于当前阶段任务。该框架与衡量结构符合性的解析器有效性分开评估。在GLM-5.2上对60个净化后的基础案例,在四种证据条件(清晰、压缩损失、部分缺失、噪声冲突)下进行决策、审计、升级三阶段处理,共执行720次调用,保留713条有效执行记录。在九组退化-清洁对比中,所有阶段成功估计均为负值,95%自举区间均低于零;所有解析器有效性点估计为正值,但三种部分缺失情境的区间包含零。在解析器有效的退化审计输出中,退化检测率为1.0,误保证率非零;在解析器有效的退化升级输出中,恢复率为0.0。结果表明:在所评估流水线中,结构合规性可改善方向性,而依赖证据的阶段成功则下降——存在可靠层与结构层的分离。结论限于特定模型配置、流水线设计、案例集、评分流程和单次缩放运行。

原文摘要 · Abstract (English)

Multi-stage LLM pipelines can remain structurally valid even when evidence available to downstream stages becomes incomplete, compressed, or conflicting. This paper introduces and operationalizes Evidence-State Reliability (ESR), an evaluation layer concerned with whether intermediate evidence remains sufficiently complete, grounded, internally consistent, and usable for a stage's assigned function. ESR is evaluated separately from parser validity, which measures structural conformance. We evaluate the framework using GLM-5.2 on 60 sanitized base cases under four evidence conditions: clean, compressed-lossy, partial-dropout, and noisy-conflicting. Each condition was processed through decision, audit, and escalation stages. The design comprised 720 planned and ledgered calls, with 713 retained, sanitized execution rows. Across nine matched degraded-minus-clean condition-stage comparisons, all operational stage-success estimates were negative, and all 95% bootstrap intervals remained below zero. All nine parser-validity point estimates were positive, although the three partial-dropout intervals included zero. Among parser-valid degraded audit outputs, degradation detection was 1.0 in each degraded condition, while false-assurance rates remained non-zero; among parser-valid degraded escalation outputs, recovery was 0.0 in every degraded condition. The results show a bounded reliability-layer divergence in the evaluated pipeline: structural conformance can improve directionally while evidence-sensitive stage success deteriorates under the same controlled intervention. They also separate detection of degraded evidence from recovery. The conclusions are limited to the evaluated model configuration, pipeline design, selected sanitized cases, scoring procedure, and single scaled run.

大模型可靠性评估流水线证据验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。