arXiv:2509.16107cs.CL2025-09中稿 · EMNLP被引 5

大模型在简短对话中难靠常识解歧,简化指令更让问题恶化。

It Depends: Resolving Referential Ambiguity in Minimal Contexts with Commonsense Knowledge

  • 用常识和上下文判断指代对象,而非盲目选一个或全列出来
  • 简化指令使模型减少常识推理,正确率大幅下降
  • 对小模型微调后,解歧能力显著提升,适合对话系统优化

指代模糊需依赖共享语境和常识来解决。本文系统研究大语言模型(LLMs)在多轮对话中利用常识处理指代歧义的能力,并分析当歧义持续存在时的应对行为。进一步考察简化语言请求对这一能力的影响。基于新构建的多语言评估数据集,测试了DeepSeek v3、GPT-4o、Qwen3-32B、GPT-4o-mini和Llama-3.1-8B,采用大模型自评与人工标注相结合的方式。结果表明:当前模型难以有效解决歧义,常固执于单一解释或罗列所有可能,而非选择性澄清或保留余地。此局限在简化提示下尤为严重,导致常识推理与策略多样性急剧下降。对Llama-3.1-8B使用直接偏好优化微调后,各类请求下的解歧表现均显著改善。研究强调需通过高级微调提升模型对歧义的处理能力,以保障不同沟通风格下的鲁棒性。

原文摘要 · Abstract (English)

Ambiguous words or underspecified references require interlocutors to resolve them, often by relying on shared context and commonsense knowledge. Therefore, we systematically investigate whether Large Language Models (LLMs) can leverage commonsense to resolve referential ambiguity in multi-turn conversations and analyze their behavior when ambiguity persists. Further, we study how requests for simplified language affect this capacity. Using a novel multilingual evaluation dataset, we test DeepSeek v3, GPT-4o, Qwen3-32B, GPT-4o-mini, and Llama-3.1-8B via LLM-as-Judge and human annotations. Our findings indicate that current LLMs struggle to resolve ambiguity effectively: they tend to commit to a single interpretation or cover all possible references, rather than hedging or seeking clarification. This limitation becomes more pronounced under simplification prompts, which drastically reduce the use of commonsense reasoning and diverse response strategies. Fine-tuning Llama-3.1-8B with Direct Preference Optimization substantially improves ambiguity resolution across all request types. These results underscore the need for advanced fine-tuning to improve LLMs' handling of ambiguity and to ensure robust performance across diverse communication styles.

指代消解常识推理对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。