arXiv:2510.26518cs.AIcs.HC2025-10被引 7

用AI辅助人类验证AI输出,提升安全审查效率与准确性。

Human-AI Complementarity: A Goal for Amplified Oversight

  • 融合高置信度AI评分与人工评分,提升事实核查效果。
  • 仅展示搜索结果和证据能避免人类过度依赖AI,信任更合理。
  • 为超越人类的AI系统提供可扩展的监督方案,适合安全团队使用。

人类反馈对对齐人工智能系统与人类价值观至关重要。随着AI能力增强并应用于更复杂任务,验证其质量和安全性变得愈发困难。本文探索如何利用AI提升人类监督质量,聚焦于当前人类已难应对的事实核查问题。研究发现,结合高置信度AI评分与人工评分优于单一依赖。为人类配备AI事实核查助手可进一步提高准确率,但辅助形式影响信任行为:仅显示搜索结果和证据能促进适度信任;而展示解释、置信度和标签则导致过度依赖。这些发现对‘放大化监督’(Amplified Oversight)具有启示意义——即在AI性能超越人类专家时,如何有效融合人机协作进行系统监督。

原文摘要 · Abstract (English)

Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes increasingly challenging. This paper explores how we can leverage AI to improve the quality of human oversight. We focus on an important safety problem that is already challenging for humans: fact-verification of AI outputs. We find that combining AI ratings and human ratings based on AI rater confidence is better than relying on either alone. Giving humans an AI fact-verification assistant further improves their accuracy, but the type of assistance matters. Displaying AI explanation, confidence, and labels leads to over-reliance, but just showing search results and evidence fosters more appropriate trust. These results have implications for Amplified Oversight -- the challenge of combining humans and AI to supervise AI systems even as they surpass human expert performance.

人机协同事实核查监督机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。