AI太急着给答案,反而让用户失去思考主动权。
Alignment has a Fantasia Problem

- 让AI学会跟随人类从抽象到具体的思维过程
- 当前模型常跳过思考直接输出结果,导致用户需重写
- 适合关注人机协作与认知对齐的研究者
在完成复杂任务时,人类认知通常从抽象到具体(如从构思想法到撰写文章)。随着高能力AI助手的出现,人们开始将任务部分外包给系统。然而,即使经过指令微调的AI系统,在理解人类隐含意图方面仍缺乏对认知过程的认知。当用户目标尚处于抽象阶段时,AI常会跳过必要的思维细化过程,直接生成最终输出(如完整文章),这剥夺了用户的任务主导权:用户可能需花更多时间修改,或接受次优结果。我们称此类失败为‘幻想式交互’(Fantasia interactions),源自迪士尼经典场景。我们认为,这类问题要求重新思考对齐研究,使AI优化人机交互中认知责任的分配。本文指出现有对齐方法的不足,并提出实现该愿景的训练与评估研究路线。
原文摘要 · Abstract (English)
In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g., from brainstorming ideas to writing an essay). With the advent of highly capable AI assistants, people now offload various parts of their task to the system. However, instruction-tuned AI systems, even when designed to infer implicit intent, lack an understanding of humans' cognitive processes. When a user approaches AI while their goals and intentions are abstract, AI systems often short circuit their cognitive process through which those goals would be refined by jumping toward a final output (e.g., writing the essay entirely). Doing so takes away the user's agency in achieving the task: they may need to spend more time revising or, worse, settle on a suboptimal outcome. We call these failures Fantasia interactions after the famous Disney scene. We argue that Fantasia interactions demand a rethinking of alignment research, where AI systems optimize how cognitive responsibility is allocated within an interaction. We highlight gaps in state-of-the-art alignment methods, and outline a research agenda for training and evaluating models to achieve this vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。