AI评估应转向人机协作,而非追求超越人类的独立表现。
AI Evaluation Should Work With Humans

- 从评估单机超人性能转向评测人机协同效果
- 强调AI应作为人类能力的补充而非替代
- 适合关注AI社会影响与人机共生的研究者
本文主张,当前主流的AI评估范式(侧重于超越人类的自主性能,隐含以取代人类为目标)正引导AI发展走向歧途。相反,学术界应转向评估人机团队的表现。我们提出,这种协作式转变将促使AI系统真正成为人类能力的互补者,从而带来远优于现行模式的社会效益。
原文摘要 · Abstract (English)
This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction. Instead, the AI community should pivot to evaluating the performance of human--AI teams. We argue that this collaborative shift will foster AI systems that act as true complements to human capabilities and therefore lead to far better societal outcomes than will the current process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。