arXiv:2606.12147cs.AI2026-06

让智能机器在合理情况下拒绝用户指令,同时保障安全与责任可追溯。

Towards Responsibly Non-Compliant Machines

论文配图:Towards Responsibly Non-Compliant Machines
图 1 · 摘自论文原文
  • 基于拒绝对话理由设计负责任的不执行机制
  • 提供绕过拒绝的合法路径以避免系统僵局
  • 关注安全风险与责任转移,适合伦理合规研究者

我们探讨如何构建能够负责任地拒绝用户请求的自主智能体。机器的不合规行为存在多种形式,需在任务拒绝的理由、违规行为的纠正路径、以及安全风险和责任转移等方面进行深入研究。本文提出应从正当性解释、可控干预机制和风险追踪三方面构建负责任的非合规能力,以确保智能系统在复杂场景下的安全与可信运行。

原文摘要 · Abstract (English)

We consider the problem of engineering autonomous intelligent agents that are capable to responsibly not comply with user requests. We argue that machine non-compliance comes in many different forms, and sketch the issues we should pursue on the road of accomplishing responsibly non-compliant intelligent machines. We anchor responsible non-compliance in justifications for task refusal, pathways to override the non-compliance, as well as careful tracking of security risks and liability transfers.

智能代理伦理合规拒绝机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。