arXiv:2608.25067cs.AI2026-08

模拟成功不等于真实部署成功,该研究提出验证框架捕捉真实世界中的隐蔽失败。

SimVerity: When Does Simulated Agent Success Survive Physical Deployment?

论文配图:SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
图 1 · 摘自论文原文
  • 通过真实环境回放与独立物理见证交叉验证,检测模拟中未暴露的失败
  • 42次亚秒级失败在模拟中被忽略,但真实摄像头捕获,且可提前预测
  • 提升模型配置可改善任务匹配率,物理测量是唯一可靠验证方式

模拟评估广泛用于评测智能体,但模拟通过能否反映真实部署表现尚未系统量化。本文提出 SimVerity:一种证言转移保障框架,通过在目标智能家居环境中回放匹配场景,并与独立验证的物理见证交叉比对代理执行结果。评估显示,部署成功是现实过程而非模拟中的静态属性:同一执行中完成度、报告状态、可观测效应和最终结果存在分歧。尽管先进模拟器在240次灯光测试中全通过,摄像头仍发现42次亚秒级失败,这些失败无法通过最终状态检查发现。失败具有可预测性:基于已测实验构建的风险模型,在评估前锁定后,能预测其从未物理测量过的路径,在两个队列共11个保留会话中均优于无属性基线。代理可审计性也可量化:切换一个代理循环的模型-客户端/服务配置,使其场景匹配率从52%-88%提升至100%。第二台经认证的模拟器未提供独立交叉验证:它与第一台在所有重叠案例中一致,唯有物理测量揭示了它们共同的盲区。SimVerity将证言转移转化为明确决策:明确通过、放弃或升级前部署。

原文摘要 · Abstract (English)

Simulated evaluation is widely used to benchmark AI agents, yet how much evidence a simulated pass provides about physical deployment has not been systematically quantified. We present SimVerity, a verdict-transfer assurance framework: it replays matched scenarios on target smart home deployments and cross-validates agent execution against independently qualified physical witnesses. Our evaluation highlights that deployment success is a real-world process, not a static property in simulation: completion, reported state, observable effect, and settled outcome diverged within the same execution. Although an advanced simulator cleared all 240 light trials, a camera caught 42 sub-second failures invisible to settled-state checks. False clearance was predictable: a risk profile learned from measured trials and locked before evaluation predicted failures on a path it never physically measured, beating a property-blind baseline in all eleven held-out sessions across two cohorts. Agent auditability was also measurable: switching one agent loop's model-client/serving configuration raised its scenario-matching share from 52-88% to 100%. Finally, a second qualified simulator added no independent cross-check: it never disagreed on any overlapping case, and only physical measurement exposed their shared blind spots. SimVerity turns verdict transfer into an explicit decision: clear, abstain, or escalate before deployment.

AI验证仿真迁移智能体评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。