arXiv:2505.19428cs.CL2025-05ACL被引 13

让AI在协作中主动提醒认知偏差,提升动态对话的对齐能力

Frictional Agent Alignment Framework: Slow Down and Don't Break Things

  • 设计双策略框架,通过可解释的'摩擦'提示重新审视观点
  • 在三个基准上实现更简洁、可解释的摩擦生成和跨域泛化
  • 适合需要动态互动的智能助手、决策支持等场景

AI在协作互动中需调解对话者信念间的潜在偏差。常见偏好对齐方法如DPO在静态场景表现良好,但在动态协作任务中因对话者信念信号稀疏且偏斜而失效。本文提出摩擦式代理对齐框架(FAAF),生成精准、上下文感知的‘摩擦’,促使反思现有证据。其双玩家目标解耦数据偏斜:摩擦状态策略识别信念不一致,干预策略生成合作者偏好的回应。我们推导出该目标的解析解,使单策略可通过简单监督损失训练。在三个基准上的实验表明,FAAF在生成简洁、可解释摩擦及分布外泛化方面优于现有方法。通过让大模型充当自适应‘思考伙伴’而非被动响应者,推动可扩展的动态人机协作。代码与数据见:https://github.com/csu-signal/FAAF_ACL。

原文摘要 · Abstract (English)

AI support of collaborative interactions entails mediating potential misalignment between interlocutor beliefs. Common preference alignment methods like DPO excel in static settings, but struggle in dynamic collaborative tasks where the explicit signals of interlocutor beliefs are sparse and skewed. We propose the Frictional Agent Alignment Framework (FAAF), to generate precise, context-aware "friction" that prompts for deliberation and re-examination of existing evidence. FAAF's two-player objective decouples from data skew: a frictive-state policy identifies belief misalignments, while an intervention policy crafts collaborator-preferred responses. We derive an analytical solution to this objective, enabling training a single policy via a simple supervised loss. Experiments on three benchmarks show FAAF outperforms competitors in producing concise, interpretable friction and in OOD generalization. By aligning LLMs to act as adaptive "thought partners" -- not passive responders -- FAAF advances scalable, dynamic human-AI collaboration. Our code and data can be found at https://github.com/csu-signal/FAAF_ACL.

人机协作对齐框架思维伙伴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。