arXiv:2606.20634cs.AIcs.CY2026-06被引 1

评测智能体运行时证据是否足够支撑决策判断,发现现有系统普遍高估证据能力。

DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency

  • 基于决策证据成熟度模型,构建跨场景评估框架
  • 64个案例中,基线系统在75%情况下过度声称证据充足
  • 提出新评分器实现零误判,平均准确率达56.25%

智能体运行系统生成轨迹、账本、溯源图、策略日志、授权令牌、缓存事件和工具防火墙记录,但这些数据未必能回答特定决策的治理问题。DEMM-Bench 是一个基于决策证据成熟度模型(DEMM)的跨范式基准,用于衡量八类证据范畴中记录是否足以重建决策层级属性,而不仅仅是存在。该基准通过适配器标准化不同范式,针对行动者、权限、行为、策略、决策依据、资源接触、生命周期上下文和验证强度提出属性问题,并施加八种确定性退化条件。在64个样本案例中,仅含轨迹或模式的基线在75%案例中过度声称,仅含账本的基线在50%中过度声称;而经过删减的属性级候选评分器无过度声称,平均属性充分性准确率为56.25%。提交包包含64案例数据集、构造标注、基线与适配器,支持在异构智能体运行证据底座上可复现地评估决策证据成熟度。

原文摘要 · Abstract (English)

Agent-runtime systems emit traces, ledgers, provenance graphs, policy logs, delegation tokens, cache events, and tool-firewall records, but those containers do not necessarily answer governance questions about a specific decision. DEMM-Bench is a cross-regime benchmark for agent-runtime governance-evidence sufficiency, grounded in the Decision Evidence Maturity Model (DEMM): it measures whether records across eight evidence regimes are sufficient to reconstruct decision-level properties rather than merely present. The benchmark normalizes the regimes through adapters, asks property questions over actor, authority, action, policy, decision basis, resource touch, lifecycle context, and verification strength, and applies eight deterministic degradation conditions. Across 64 manuscript cases, trace-present and schema-present baselines overclaim on 75% of cases, ledger-present overclaims on 50%, and the redacted property-level candidate scorer has zero overclaim with 56.25% mean Property Sufficiency Accuracy. The deposited package provides the 64-case dataset, construction-oracle labels, baselines, and adapters, supporting reproducible evaluation of decision-evidence maturity across heterogeneous agent-runtime evidence substrates.

智能体治理证据评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。