用新闻访谈数据评估大模型的对话策略短板
NewsInterview: a Dataset and a Playground to Evaluate LLMs' Ground Gap via Informational Interviews
- 构建4万条新闻访谈数据集,模拟真实对话场景
- 大模型难识别回答、不会转问高阶问题,信息获取效率低
- 适合研究对话智能、人机交互与模型评估的学者
大型语言模型在生成连贯文本方面表现优异,但在语言接地和策略性对话上仍存不足。为此,我们聚焦新闻访谈这一富含信息接地与策略互动的领域,从美国国家公共电台(NPR)和有线电视新闻网(CNN)收集了4万组两人访谈数据。分析发现,相比人类采访者,大模型更少使用确认性回应,也更难转向更高层次的问题。为解决多轮对话规划与长期策略思维缺陷,我们构建了一个包含角色设定与说服元素的仿真环境,以支持具备长程奖励机制的智能体开发。实验表明,尽管源模型能模仿人类的信息分享行为,但作为采访者的模型在判断问题是否被回答及有效说服对方方面表现不佳,导致信息提取效果随模型规模与能力提升仍不理想。研究强调需加强大模型的战略对话能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated impressive capabilities in generating coherent text but often struggle with grounding language and strategic dialogue. To address this gap, we focus on journalistic interviews, a domain rich in grounding communication and abundant in data. We curate a dataset of 40,000 two-person informational interviews from NPR and CNN, and reveal that LLMs are significantly less likely than human interviewers to use acknowledgements and to pivot to higher-level questions. Realizing that a fundamental deficit exists in multi-turn planning and strategic thinking, we develop a realistic simulated environment, incorporating source personas and persuasive elements, in order to facilitate the development of agents with longer-horizon rewards. Our experiments show that while source LLMs mimic human behavior in information sharing, interviewer LLMs struggle with recognizing when questions are answered and engaging persuasively, leading to suboptimal information extraction across model size and capability. These findings underscore the need for enhancing LLMs' strategic dialogue capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。