发现用户与AI对话中,隐含事实判断比显式提问更常见。
WildClaims: Information Access Conversations in the Wild(Chat)
- 分析真实对话数据,发现系统生成的陈述常被用户当作事实核查对象
- 提取12万条事实主张,其中18%~76%的对话含值得验证的陈述
- 提出新数据集WildClaims,推动对隐式信息获取的研究
大型语言模型的快速发展使对话系统成为亿万用户的实用工具。然而,现实对话中信息获取的形态与必要性仍缺乏研究,现有工作多集中于传统的显式信息查询。核心问题是:真实世界的信息获取对话是什么样?为此,我们对大规模用户-ChatGPT对话数据集WildChat进行了观察研究,发现即使对话主要目的非信息性(如创意写作),用户也会对系统生成的、具有潜在事实性的陈述进行隐式信息核实。为系统研究此现象,我们发布了WildClaims数据集,包含从3,000次对话中的7,587条语句提取的121,905条事实主张,并标注其可核查性。初步分析显示,保守估计18%至51%的对话含有可核查主张,更激进估计可达76%。这一高频率凸显了超越传统显式信息访问理解的重要性,需关注真实用户-系统对话中涌现的隐式信息获取行为。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has transformed conversational systems into practical tools used by millions. However, the nature and necessity of information retrieval in real-world conversations remain largely unexplored, as research has focused predominantly on traditional, explicit information access conversations. The central question is: What do real-world information access conversations look like? To this end, we first conduct an observational study on the WildChat dataset, large-scale user-ChatGPT conversations, finding that users' access to information occurs implicitly as check-worthy factual assertions made by the system, even when the conversation's primary intent is non-informational, such as creative writing. To enable the systematic study of this phenomenon, we release the WildClaims dataset, a novel resource consisting of 121,905 extracted factual claims from 7,587 utterances in 3,000 WildChat conversations, each annotated for check-worthiness. Our preliminary analysis of this resource reveals that conservatively 18% to 51% of conversations contain check-worthy assertions, depending on the methods employed, and less conservatively, as many as 76% may contain such assertions. This high prevalence underscores the importance of moving beyond the traditional understanding of explicit information access, to address the implicit information access that arises in real-world user-system conversations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。