arXiv:2608.06657cs.AIcs.HC2026-08中稿 · presentation and p…

构建多层人机协同故障追踪基准,定位漂移来源与传播路径。

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

  • 从ALFRED数据集注入可控漂移,生成1918条时间对齐的多层轨迹
  • 漂移可准确识别至具体层级(宏观F1近0.70)、责任人(近0.85)和成因(近0.49)
  • 验证了简单模型在符号任务中已足够,无需复杂注意力结构

现代人机协同控制系统将操作员、AI决策模块与自动化控制器整合于同一控制回路中,其可信性取决于整个回路而非单一模型。然而,当前尚无标准基准能捕捉漂移与故障在多层间的时序对齐传播过程,导致无法诊断协作失效的位置、原因及恢复方式。本文聚焦此缺口的一方面:漂移——一种可起源于任意系统层且传统单模态监测难以定位层级或发生时间的偏差。我们通过向源自ALFRED(一个基于日常家庭任务的具身指令基准)的轨迹中注入受控漂移,构建了包含1,918条漂移轨迹的基准数据集。每条轨迹为跨五个执行层(状态、观测、决策、规则、控制)的时间对齐记录序列,标注了漂移类型、影响层级、起始时间、责任主体与因果机制,并经独立评审员验证,报告了评标者间一致性。我们配套提出去泄漏协议,消除近完美起始泄露问题,并在经典、循环与注意力模型族上开展基线研究。在该诚实协议下,漂移在各模型族中均显著优于随机与多数基线,受影响层级宏平均F1接近0.70,责任主体识别率约0.85,因果机制识别率约0.49;重注意力模型在该符号基准上并无优势。

原文摘要 · Abstract (English)

Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model. Yet no standard benchmark captures time-aligned, multi-layer traces of how drift and failures propagate across these layers, so we cannot diagnose where coordination breaks down, why, or how to recover. This paper targets one facet of that gap: drift, a deviation that can originate in any stack layer and that conventional single-modality monitoring cannot localize to a layer or pin to an onset time. We construct a benchmark by injecting controlled drift into traces derived from ALFRED, a grounded-instruction benchmark for everyday household tasks, yielding 1,918 drifted traces. Each trace is a time-aligned sequence of per-step records across five execution layers (state, observation, decision, rules, control), labeled with the drift type, affected layer, onset time, responsible actor, and causal mechanism, and validated by independent raters with inter-annotator agreement reported. We pair the dataset with a leak-aware protocol that removes a near-perfect onset leak, and a baseline study across classical, recurrent, and attention-based model families. Under this honest protocol, drift is identifiable and attributable well above random and majority baselines across every family (affected layer macro-F1 near 0.70, responsible actor near 0.85, causal mechanism near 0.49), and heavy attention offers no advantage over simpler models on this symbolic benchmark.

人机协同故障诊断多层追踪基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。