arXiv:2603.01189cs.ROcs.HC2026-03

用模拟实验揭示人机团队中信任如何随可靠性变化,指导安全部署。

Agent-Based Simulation of Trust Development in Human-Robot Teams: An Empirically-Validated Framework

  • 基于实证数据构建可模拟的人机信任模型,支持多角色协作测试。
  • 机器人可靠性影响最大,解释45.4%的信任差异,且显著提升任务成功率。
  • 发现信任与绩效可能脱钩,适合人机交互设计与安全评估研究者参考。

本文提出一个基于实证的代理模型,用于模拟人机团队中的信任动态、工作负荷分配及协作表现。模型在NetLogo 6.4.0中实现,模拟2–10个代理完成不同复杂度任务。通过与Hancock等(2021)元分析对比,8个信任前因类别中有4个达到区间有效性,等级有效性达Spearman ρ=0.833。采用OFAT与全因子设计(每条件50次重复)的敏感性分析显示,机器人可靠性对信任影响最强(η²=0.35),并主导任务成功(η²=0.93)和生产率(η²=0.89),与元分析一致。信任不对称比值范围为0.07–0.55,低于元分析基准1.50,表明事件级不对称不必然导致累积不对称,因信任修复机制持续作用。情景分析揭示信任-绩效解耦:在信任恢复场景中,尽管信任最低(38.2),生产力最高(4.29);在不可靠机器人场景中,信任最高(73.2),但任务成功率最低(33.4%),说明校准误差是独立于信任强度的关键诊断指标。方差分析确认可靠性、透明度、沟通与协作均有显著主效应(p<.001),共解释45.4%的信任变异。开源实现为部署前识别过度信任与不足信任提供了证据基础。

原文摘要 · Abstract (English)

This paper presents an empirically grounded agent-based model capturing trust dynamics, workload distribution, and collaborative performance in human-robot teams. The model, implemented in NetLogo 6.4.0, simulates teams of 2--10 agents performing tasks of varying complexity. We validate against Hancock et al.'s (2021) meta-analysis, achieving interval validity for 4 of 8 trust antecedent categories and strong ordinal validity (Spearman \r{ho}=0.833ρ= 0.833 \r{ho}=0.833). Sensitivity analysis using OFAT and full factorial designs (n=50n = 50 n=50 replications per condition) reveals robot reliability exhibits the strongest effect on trust (η2=0.35η^2 = 0.35 η2=0.35) and dominates task success (η2=0.93η^2 = 0.93 η2=0.93) and productivity (η2=0.89η^2 = 0.89 η2=0.89), consistent with meta-analytic findings. Trust asymmetry ratios ranged from 0.07 to 0.55 -- below the meta-analytic benchmark of 1.50 -- revealing that per-event asymmetry does not guarantee cumulative asymmetry when trust repair mechanisms remain active. Scenario analysis uncovered trust-performance decoupling: the Trust Recovery scenario achieved the highest productivity (4.29) despite the lowest trust (38.2), while the Unreliable Robot scenario produced the highest trust (73.2) despite the lowest task success (33.4\%), establishing calibration error as a critical diagnostic distinct from trust magnitude. Factorial ANOVA confirmed significant main effects for reliability, transparency, communication, and collaboration (p<.001p < .001 p<.001), explaining 45.4\% of trust variance. The open-source implementation provides an evidence-based tool for identifying overtrust and undertrust conditions prior to deployment.

人机协作信任建模仿真模拟可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。