测试大模型能否理解善意谎言背后的社交动机。
TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?
- 通过人机协作生成包含信息差的对话,模拟真实善意谎言场景。
- 顶尖模型表现远低于人类,暴露其对社交心理理解的不足。
- 适合研究大模型社会认知能力的学者或伦理对齐开发者。
尽管已有研究探讨大语言模型(LLMs)在心智理论(ToM)推理任务中的表现,但针对更复杂社交情境(如善意谎言)的研究仍有限。本文提出TactfulToM,一个新型英文基准,用于评估LLMs在真实对话中理解善意谎言的能力,以及推断其背后利他动机——例如为维护他人感受与社会和谐。该基准通过多阶段人机协同流程生成:由人工设计初始故事,再由LLMs扩展成保持参与者间信息不对称的对话,以确保谎言的真实性。实验表明,当前最先进模型在该任务上表现显著低于人类,揭示其在实现真正善意谎言理解所需的深层心智理论推理方面仍存在明显缺陷。
原文摘要 · Abstract (English)
While recent studies explore Large Language Models' (LLMs) performance on Theory of Mind (ToM) reasoning tasks, research on ToM abilities that require more nuanced social context is limited, such as white lies. We introduce TactfulToM, a novel English benchmark designed to evaluate LLMs' ability to understand white lies within real-life conversations and reason about prosocial motivations behind them, particularly when they are used to spare others' feelings and maintain social harmony. Our benchmark is generated through a multi-stage human-in-the-loop pipeline where LLMs expand manually designed seed stories into conversations to maintain the information asymmetry between participants necessary for authentic white lies. We show that TactfulToM is challenging for state-of-the-art models, which perform substantially below humans, revealing shortcomings in their ability to fully comprehend the ToM reasoning that enables true understanding of white lies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。