arXiv:2608.30683cs.CLcs.CY2026-08中稿 · EMNLP

构建真实用户查询数据集,评估大模型在高风险信息请求中的表现。

WildSEEK: Evaluating Language Models for Information-Seeking

论文配图:WildSEEK: Evaluating Language Models for Information-Seeking
图 1 · 摘自论文原文
  • 基于3000条真实用户查询构建手动标注数据集
  • 超三分之一查询属高风险,且多为分析型问题
  • 发现模型在安全、公平性上存在四大缺陷

语言模型正越来越多地作为用户获取信息的中介,亟需系统评估其响应质量以保障信息生态的公正与可靠。现有评估多局限于特定主题或合成数据,难以捕捉真实世界信息查询的复杂性及模型响应中的风险。为此,我们提出WildSEEK——一个包含3000条真实用户交互信息查询的手动标注数据集,以及一套针对大模型生成响应的评估框架。该数据集涵盖健康、金融等敏感领域,并区分事实型与分析型查询。我们在其上训练分类器,分析超过180万条真实用户查询。结果表明,超过三分之一的信息查询属于高风险,且更常为分析型。模型在四种表现上失败率较高:谄媚行为、过度依赖、默认美国中心视角、对弱势群体处理不当,且分析型查询中失败率更高。本研究为监控大模型在信息获取中的可靠性、安全性和公平性提供了实证基础。

原文摘要 · Abstract (English)

Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of "in the wild" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our findings show that LLM responses fail more often in four criteria: sycophantic behavior, overreliance, a default US-centric perspective, and poor handling of vulnerable populations -- with failure rates being mostly higher for analytical queries. By providing methods to monitor the reliability, safety, and fairness of LLM behavior, our dataset and evaluation framework offer an empirical foundation for the broader question of how these systems should behave as they take on a growing role in information access.

大模型评估信息检索安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。