构建首个信息挖掘对话数据集,让大模型学会主动提问。
YIELD: A Large-Scale Dataset and Evaluation Framework for Information Elicitation Agents
- 将信息挖掘设计为有限时长远见决策问题,建模更真实
- 2600万词对话数据,涵盖2281场伦理合规的人类对话
- 适合研究司法、调查等需要主动追问的场景
大多数对话系统以用户需求为导向,但在学术访谈、司法程序和新闻调查等现实场景中,需要能主动从用户处获取信息的智能体。本文提出信息挖掘代理(IEAs),其目标是通过对话支持机构或任务目标。为此,我们构建了YIELD数据集,包含2,281场伦理合规的人类间对话,总计2600万词。我们形式化信息挖掘为有限时长远见部分可观测马尔可夫决策过程(POMDP),并提出针对IEAs的新评估指标。在多个基础大模型上的初步实验表明,基于YIELD训练可显著提升模型与真实挖掘行为的一致性,且人类评估结果验证了该效果。数据集、代码、评估工具和微调适配器已开源。
原文摘要 · Abstract (English)
Most conversational agents (CAs) are designed to satisfy user needs through user-driven interactions. However, many real-world settings, such as academic interviewing, judicial proceedings, and journalistic investigations, involve broader institutional decision-making processes and require agents that can elicit information from users. In this paper, we introduce Information Elicitation Agents (IEAs) in which the agent's goal is to elicit information from users to support the agent's institutional or task-oriented objectives. To enable systematic research on this setting, we present YIELD, a 26M-token dataset of 2,281 ethically sourced, human-to-human dialogues. Moreover, we formalize information elicitation as a finite-horizon POMDP and propose novel metrics tailored to IEAs. Pilot experiments on multiple foundation LLMs show that training on YIELD improves their alignment with real elicitation behavior and findings are corroborated by human evaluation. We release YIELD under CC BY 4.0. The dataset, project code, evaluation tools, and fine-tuned model adapters are available at: https://github.com/infosenselab/yield.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。