arXiv:2605.15635cs.CL2026-05

构建首个中文歧义数据集,揭示大模型在理解中文歧义时的三大失败模式。

Evaluating Chinese Ambiguity Understanding in Large Language Models

论文配图:Evaluating Chinese Ambiguity Understanding in Large Language Models
图 1 · 摘自论文原文
  • 基于潜在歧义理论设计半自动构建流程,生成5712句中文歧义语料
  • 大模型歧义检测准确率低,思维链提示可提升性能,但仍有盲点
  • 发现模型偏好主流解释,且指令微调导致过度自信

语言歧义对大模型的鲁棒性至关重要,但现有研究主要聚焦英文,中文相关研究匮乏。现有中文歧义数据集(如CHAmbi)可扩展性差。本文基于潜在歧义(PA)理论,设计半自动流程构建CHA-Gen,这是首个基于PA理论的中文歧义数据集,包含5,712句(2,414句歧义,3,298句非歧义),覆盖18种潜在歧义结构。通过直接查询与机器翻译评估Gemma 3、Qwen 2.5/3系列等模型,发现大模型在歧义检测上表现不佳,思维链提示可改善结果。分析Qwen3-32B的思维链推理发现三种典型错误:歧义盲视、误归因和过早消解。采用语义熵度量不确定性,显示歧义句子不确定性更高。此外,指令微调导致模型过度自信,而基础模型更善于捕捉语义多样性。模型还表现出对主导解释的偏好。本研究提供可扩展的中文歧义语料构建方法,并为提升中文大模型歧义理解能力奠定基础。

原文摘要 · Abstract (English)

Linguistic ambiguity is critical to the robustness of Large Language Models (LLMs), yet existing research focuses mostly on English, with limited attention devoted to Chinese. Existing Chinese ambiguity datasets (e.g., CHAmbi) suffer from poor scalability. Guided by Potential Ambiguity (PA) Theory, we design a semi-automatic pipeline to construct CHA-Gen. It is the first PA Theory-grounded Chinese ambiguity dataset, which comprises 5,712 sentences (2,414 ambiguous, 3,298 unambiguous) across 18 potential ambiguous structures. Evaluating LLMs (e.g. Gemma 3, Qwen 2.5/3 series) via direct querying and machine translation, we find that LLMs struggle with ambiguity detection (improved by CoT prompting). Analysis of Qwen3-32B's CoT rationales reveals three common failure modes: ambiguity blindness, misattribution, and premature resolution. Uncertainty quantification with semantic entropy metric shows higher uncertainty for ambiguous sentences. Moreover, instruction tuning induces overconfidence, whereas Base models better capture semantic diversity. We further observe that models exhibit a bias toward dominant interpretations. Our work provides a scalable approach for Chinese ambiguity corpus and insights into LLMs' ambiguity handling, laying a foundation for enhancing Chinese ambiguity research in LLMs.

中文歧义大模型评测语义熵思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。