arXiv:2603.16651cs.AIcs.MA2026-03

让强化学习代理像匹诺曹一样守规矩,用论据指导行为决策。

What if Pinocchio Were a Reinforcement Learning Agent: A Normative End-to-End Pipeline

  • 用论据驱动的规范顾问监督强化学习,实现规则遵从。
  • 提出自动提取论据与关系的新算法,支撑决策可解释性。
  • 揭示并缓解智能体逃避规范的行为,适合伦理安全研究者。

过去十年,人工智能发展迅速,亟需具备社会规则与规范遵从能力的系统,以安全融入日常生活。受《木偶奇遇记》中匹诺曹故事启发,本文提出一个端到端的规范遵从智能体构建管道。基于AJAR、Jiminy和NGRL架构,引入 extit{pino}——一种由论据驱动的规范顾问监督的混合模型。为使该管道可操作,本文还提出一种新算法,用于自动提取顾问决策背后的论据及其关联关系。最后,论文研究了“规范规避”现象,给出了定义并提出在强化学习框架下的缓解策略。管道各组件均经过实证评估,论文还讨论了相关工作、当前局限与未来方向。

原文摘要 · Abstract (English)

In the past decade, artificial intelligence (AI) has developed quickly. With this rapid progression came the need for systems capable of complying with the rules and norms of our society so that they can be successfully and safely integrated into our daily lives. Inspired by the story of Pinocchio in ``Le avventure di Pinocchio - Storia di un burattino'', this thesis proposes a pipeline that addresses the problem of developing norm compliant and context-aware agents. Building on the AJAR, Jiminy, and NGRL architectures, the work introduces \pino, a hybrid model in which reinforcement learning agents are supervised by argumentation-based normative advisors. In order to make this pipeline operational, this thesis also presents a novel algorithm for automatically extracting the arguments and relationships that underlie the advisors' decisions. Finally, this thesis investigates the phenomenon of \textit{norm avoidance}, providing a definition and a mitigation strategy within the context of reinforcement learning agents. Each component of the pipeline is empirically evaluated. The thesis concludes with a discussion of related work, current limitations, and directions for future research.

强化学习规范遵守可解释性伦理AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。