评估六个厂商SDK中代理决策的可还原性,发现推理轨迹是跨框架共性短板。
Property-Level Reconstructability of Agent Decisions: An Anchor-Level Pilot Across Vendor SDK Adapter Regimes
- 用不变的追踪工具在六个厂商框架中测试决策还原能力
- 42.9%至85.7%的治理完整性差异揭示制度性差距
- 推理轨迹缺失是跨框架共性问题,适合系统设计者参考
代理智能失败后需事后还原其行为:做了什么、依据谁的授权、违反何种策略、基于何种推理。现有单一属性级框架下跨制度可行性未被量化。本研究将未经修改的决策追溯器应用于六个公开厂商SDK框架(涵盖云代理、可观测性、工具使用、遥测与协议日志)中的固定示例锚点,外加两组对照列。每个决策事件模式(DES)属性被分类为完全可填、部分可填、结构不可填或透明。在锚点层级上,代理决策的属性可还原性已在不同框架间显现差异。严格治理完整性划分为三个等级,介于42.9%至85.7%之间,暴露出一个跨框架共性缺口(推理轨迹)、四个框架特有缺口及一个混合属性;该初步研究为单标注者、每单元一锚点、描述性分析,输出可通过已存复现包校验。
原文摘要 · Abstract (English)
Agentic AI failures need post-hoc reconstruction: what the agent did, on whose authority, against which policy, and from what reasoning. Cross-regime feasibility remains unmeasured under one property-level schema. We apply the Decision Trace Reconstructor unmodified to pinned worked-example anchors from six public vendor SDK regimes spanning cloud-agent, observability, tool-use, telemetry, and protocol traces, plus two comparator columns. Each Decision Event Schema (DES) property is classified as fully fillable, partially fillable, structurally unfillable, or opaque. Per-property reconstructability of an agent decision already varies between regimes at this anchor scale. Strict-governance-completeness separates into three tiers ranging from 42.9% to 85.7%, yielding one regime-independent gap (reasoning trace), four regime-dependent gaps, and one Mixed property; the pilot is single-annotator, one anchor per cell, descriptive, with outputs checksum-verifiable from a deposited reproducibility package.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。