arXiv:2608.25245cs.IRcs.CL2026-08

提出概念溯源框架,精准识别大模型查询中隐含答案知识的违规信息。

The "Curse of Knowledge" in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion

论文配图:The "Curse of Knowledge" in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion
图 1 · 摘自论文原文
  • 通过概念归属四区法区分查询中的知识来源,突破传统评估指标盲区。
  • 7.4%非通用概念源于答案侧,97个主题中均有出现,主题解释67%方差变化。
  • 可诊断检索边界违规,适合评测大模型生成查询的可信性与合规性。

大模型生成的搜索查询广泛用于信息检索评估,但可能包含预设答案文档知识的概念,违背了预搜索用户的知识边界。现有重叠、多样性与有效性等验证指标无法区分人类尾部变异与候选答案侧入侵。本文提出概念溯源框架,将查询概念划分为背景支持、人类中心、人类尾部和候选答案侧四个区域,实现了仅靠检索指标无法检测的边界判定。在100个UQV100主题、8个大模型及5种提示条件下,对77,004条查询应用两种提取流程,跨流程的词级HCIR斯皮尔曼等级相关系数达1.0(五条件均值)。候选答案侧概念占非通用概念的7.40%,出现在97个主题中,主题因素解释约67%的方差。人工验证显示68.2%宽松准确率,揭示两类机制:知识入侵占比45.5%,部署入侵占比45.0%。诊断探针显示局部删除效应更显著(效应量d = -0.47),高于随机删除(d = -0.34),但这些概念仅解释不足2%的整体评估方差。因此,概念溯源是边界合规诊断工具,而非评估偏移预测器。在测试条件下,无提示方式能消除入侵;后处理概念溯源筛选可实现99%的入侵消除。

原文摘要 · Abstract (English)

LLM-generated search queries are widely used to augment IR evaluation, yet they may contain concepts that presuppose answer-side document knowledge, violating the information-access boundary of pre-search users. Existing validation metrics, including overlap, diversity, and effectiveness, cannot distinguish rare human-tail variation from candidate answer-side intrusion. We introduce concept provenance, a framework that assigns query concepts to backstory-supported, human-central, human-tail, and candidate answer-side zones, operationalizing a boundary that retrieval metrics alone cannot detect. Applying concept provenance to 77,004 queries across 100 UQV100 topics, 8 LLMs, and 5 prompt conditions with two extraction pipelines, we obtain a cross-pipeline token-HCIR Spearman rho of 1.0 over five condition means. Candidate answer-side concepts constitute 7.40 percent of non-generic concepts and appear in 97 of 100 topics, with topic explaining approximately 67 percent of variance. Human validation yields 68.2 percent relaxed precision, revealing two mechanisms: knowledge intrusion at 45.5 percent and deployment intrusion at 45.0 percent. Diagnostic probes show disproportionate localized retrieval effects, with deletion effect size d = -0.47 compared with d = -0.34 for random deletion, but these concepts explain less than 2 percent of aggregate evaluation variance. Concept provenance therefore serves as a boundary-compliance diagnostic rather than an evaluation-shift predictor. Under the tested conditions, no prompt condition eliminates intrusion; post-generation concept-provenance selection achieves 99 percent elimination.

大模型评估信息检索知识泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。