arXiv:2502.13270cs.CL2025-02被引 31

真实对话数据集揭示模型在情感与记忆上的短板

REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation

  • 基于21天真实聊天记录构建数据集,分析情绪与人格一致性
  • 模型仅靠对话历史难以模拟用户人格,需针对性微调提升表现
  • 适合研究长时对话、情感智能和个性化生成的学者使用

长期、开放域对话能力对具备回忆过往互动和情感智能(EI)的聊天机器人至关重要。然而,现有研究多依赖大语言模型生成的合成数据,缺乏对真实世界对话模式的了解。为弥补这一空白,我们提出REALTALK,一个包含21天真实消息应用对话的语料库,为真实人类交互提供直接基准。我们首先开展数据集分析,聚焦情绪属性与人格一致性,揭示真实对话的独特挑战。与大语言模型生成对话对比发现,真实对话具有更丰富的感情表达和更不稳定的个性特征,而合成数据常无法捕捉这些特性。基于此,我们设计两项基准任务:(1) 人格模拟——模型在给定历史对话的基础上代入特定用户继续对话;(2) 记忆探测——模型回答需要长期记忆支持的问题。结果表明,模型仅凭对话历史难以有效模拟用户人格,但在针对特定用户进行微调后可改善人格模仿效果;此外,现有模型在真实对话中仍面临显著的长期上下文记忆与利用难题。

原文摘要 · Abstract (English)

Long-term, open-domain dialogue capabilities are essential for chatbots aiming to recall past interactions and demonstrate emotional intelligence (EI). Yet, most existing research relies on synthetic, LLM-generated data, leaving open questions about real-world conversational patterns. To address this gap, we introduce REALTALK, a 21-day corpus of authentic messaging app dialogues, providing a direct benchmark against genuine human interactions. We first conduct a dataset analysis, focusing on EI attributes and persona consistency to understand the unique challenges posed by real-world dialogues. By comparing with LLM-generated conversations, we highlight key differences, including diverse emotional expressions and variations in persona stability that synthetic dialogues often fail to capture. Building on these insights, we introduce two benchmark tasks: (1) persona simulation where a model continues a conversation on behalf of a specific user given prior dialogue context; and (2) memory probing where a model answers targeted questions requiring long-term memory of past interactions. Our findings reveal that models struggle to simulate a user solely from dialogue history, while fine-tuning on specific user chats improves persona emulation. Additionally, existing models face significant challenges in recalling and leveraging long-term context within real-world conversations.

长时对话情感智能真实数据记忆建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。