arXiv:2512.18474cs.ROcs.HC2025-12中稿 · the ACM/IEEE Inter…

构建机器人拒绝对话的标准化测试平台,评估安全与信任平衡。

When Robots Say No: The Empathic Ethical Disobedience Benchmark

  • 设计多场景拒绝对话环境,让机器人权衡风险、情绪与信任
  • 解释性拒绝能维持信任,安全强化学习提升鲁棒性但易过度谨慎
  • 提供可复现的代码与参考策略,适合人机交互与伦理研究者

机器人需在服从指令与保障安全、符合社会期待间取得平衡,盲目服从可能造成伤害,过度拒绝则削弱信任。现有安全强化学习基准侧重物理风险,而人机交互信任研究规模小且难以复现。本文提出情感伦理抗拒基准(EED Gym),一个标准化测试平台,同时评估拒绝的安全性与社会可接受性。智能体在面对不同情境时,可选择服从、拒绝(带或不带解释)、澄清或提出更安全替代方案。该平台包含多种情景、人物角色设定及安全、校准度与拒绝行为的量化指标,其信任与责备模型基于情景实验构建。实验发现,动作掩蔽可消除不安全服从,解释性拒绝有助于维持信任;建设性风格最可信,共情风格最具同理心,而安全强化学习虽提升鲁棒性,但也导致代理更倾向过度谨慎。我们发布代码、配置和参考策略,支持可复现的人机交互研究。投稿时附匿名可复现包,承诺论文录用后全面开源。

原文摘要 · Abstract (English)

Robots must balance compliance with safety and social expectations as blind obedience can cause harm, while over-refusal erodes trust. Existing safe reinforcement learning (RL) benchmarks emphasize physical hazards, while human-robot interaction trust studies are small-scale and hard to reproduce. We present the Empathic Ethical Disobedience (EED) Gym, a standardized testbed that jointly evaluates refusal safety and social acceptability. Agents weigh risk, affect, and trust when choosing to comply, refuse (with or without explanation), clarify, or propose safer alternatives. EED Gym provides different scenarios, multiple persona profiles, and metrics for safety, calibration, and refusals, with trust and blame models grounded in a vignette study. Using EED Gym, we find that action masking eliminates unsafe compliance, while explanatory refusals help sustain trust. Constructive styles are rated most trustworthy, empathic styles -- most empathic, and safe RL methods improve robustness but also make agents more prone to overly cautious behavior. We release code, configurations, and reference policies to enable reproducible evaluation and systematic human-robot interaction research on refusal and trust. At submission time, we include an anonymized reproducibility package with code and configs, and we commit to open-sourcing the full repository after the paper is accepted.

人机交互伦理决策强化学习信任建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。