arXiv:2502.12436cs.CL2025-02ACL被引 5

用反事实强化学习检测谈判中的欺骗行为,提升识别精度。

Should I Trust You? Detecting Deception in Negotiations using Counterfactual RL

  • 基于反事实强化学习分析玩家沟通逻辑与收益预期
  • 在《外交》游戏中实现高精度欺骗检测,优于大模型方法
  • 适合开发人机交互防骗工具,提升决策可靠性

日益普遍的社会技术问题是人们容易被看似‘好得难以置信’的提议所误导,其中说服力和信任感影响决策。本文研究如何利用人工智能帮助识别此类欺骗场景。我们分析人类在《外交》(Diplomacy)这一需自然语言交流与策略推理的棋类游戏中如何相互欺骗。通过提取玩家沟通中提议的逻辑形式,并结合代理的价值函数计算提议的相对收益,再融合文本特征,可有效提升欺骗检测能力。相较于将大量真实信息误判为欺骗的大语言模型方法,本方法在检测人类欺骗时展现出更高精度。未来的人机交互工具可借鉴此方法,对可疑提议触发‘摩擦机制’,给予用户质疑机会。

原文摘要 · Abstract (English)

An increasingly common socio-technical problem is people being taken in by offers that sound ``too good to be true'', where persuasion and trust shape decision-making. This paper investigates how \abr{ai} can help detect these deceptive scenarios. We analyze how humans strategically deceive each other in \textit{Diplomacy}, a board game that requires both natural language communication and strategic reasoning. This requires extracting logical forms of proposed agreements in player communications and computing the relative rewards of the proposal using agents' value functions. Combined with text-based features, this can improve our deception detection. Our method detects human deception with a high precision when compared to a Large Language Model approach that flags many true messages as deceptive. Future human-\abr{ai} interaction tools can build on our methods for deception detection by triggering \textit{friction} to give users a chance of interrogating suspicious proposals.

欺骗检测强化学习人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。