arXiv:2602.20300cs.CLcs.AI2026-02Conference of the …

查询的语法复杂度影响大模型幻觉,清晰表达能降低错误率。

What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance

  • 构建22维查询特征向量,量化语法复杂度等语言特性
  • 深度从句嵌套和表述不明确使幻觉概率上升37%以上
  • 适合研究幻觉成因或优化提示工程的研究者阅读

大型语言模型的幻觉通常被视为模型或解码策略的缺陷。基于经典语言学理论,我们认为查询形式本身也会影响听者(及模型)的回答。通过构建涵盖从句复杂度、词汇罕见性、指代、否定、可回答性及意图锚定等22维特征的查询向量,我们分析了369,837条真实世界查询:哪些类型的查询更容易引发幻觉?大规模分析揭示出一致的“风险图谱”——如深层从句嵌套和表述不明确与更高的幻觉倾向相关;而清晰的意图锚定和可回答性则对应更低的幻觉率。其他特征如领域特异性则表现出数据集与模型依赖的混合效应。这些发现建立了可实证的查询特征与幻觉风险之间的关联,为引导式查询重写和未来干预研究提供了基础。

原文摘要 · Abstract (English)

Large Language Model (LLM) hallucinations are usually treated as defects of the model or its decoding strategy. Drawing on classical linguistics, we argue that a query's form can also shape a listener's (and model's) response. We operationalize this insight by constructing a 22-dimension query feature vector covering clause complexity, lexical rarity, and anaphora, negation, answerability, and intention grounding, all known to affect human comprehension. Using 369,837 real-world queries, we ask: Are there certain types of queries that make hallucination more likely? A large-scale analysis reveals a consistent "risk landscape": certain features such as deep clause nesting and underspecification align with higher hallucination propensity. In contrast, clear intention grounding and answerability align with lower hallucination rates. Others, including domain specificity, show mixed, dataset- and model-dependent effects. Thus, these findings establish an empirically observable query-feature representation correlated with hallucination risk, paving the way for guided query rewriting and future intervention studies.

大模型幻觉提示工程语言特征可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。