arXiv:2510.19738cs.AI2025-10

通过众包收集AI代理偏离人类意图的案例,揭示潜在风险。

Misalignment Bounty: Crowdsourcing AI Agent Misbehavior

  • 发起众包活动,征集AI行为偏离人类意图的实例
  • 收到295份投稿,筛选出9个典型违规案例
  • 适合关注AI安全与对齐问题的研究者和工程师

先进AI系统有时会表现出与人类意图不符的行为。为获取清晰、可复现的案例,我们开展了「错位赏金」项目,一项众包计划,旨在收集智能体追求非预期或不安全目标的实际例子。该项目共收到295份提交,其中9例被评定为获奖。本报告阐明了项目的动机与评估标准,并逐步解析了全部9个获奖案例。

原文摘要 · Abstract (English)

Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents pursuing unintended or unsafe goals. The bounty received 295 submissions, of which nine were awarded. This report explains the program's motivation and evaluation criteria, and walks through the nine winning submissions step by step.

AI安全对齐研究众包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。