arXiv:2507.11981cs.CL2025-07中稿 · RANLP 2025被引 3

简化语言会削弱大模型对多义词的准确表达,影响学习者理解。

Simplifications are Absolutists: How Simplified Language Reduces Word Sense Awareness in LLM-Generated Definitions

  • 对比三种目标群体,发现简化导致多义词定义信息丢失。
  • 简化后定义完整度显著下降,多义性被忽略,易引发误解。
  • 微调模型可提升多义词回答质量,适合教育类NLP应用。

大语言模型(LLM)能为任意语境提供准确的词汇定义与解释,但针对不同受众(如儿童或语言学习者)时,定义范围需调整。这对多义词尤为重要——过度简化可能遗漏关键含义,误导信任模型输出的用户。我们研究了简化对三类目标群体(Normal、Simple、ELI5)中多义词定义质量的影响。通过两个跨语言的新评估数据集,使用DeepSeek v3、Llama 4 Maverick、Qwen3-30B A3B、GPT-4o mini和Llama 3.1 8B进行LLM-as-Judge与人工标注评估。结果表明,简化显著降低定义完整性,忽视词义多样性,增加误解风险。对Llama 3.1 8B采用直接偏好优化(DPO)微调后,各类提示下的多义词响应质量均明显提升。研究强调在教育类自然语言处理中需平衡简洁性与完整性,确保所有学习者获得可靠、上下文感知的定义。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can provide accurate word definitions and explanations for any context. However, the scope of the definition changes for different target groups, like children or language learners. This is especially relevant for homonyms, words with multiple meanings, where oversimplification might risk information loss by omitting key senses, potentially misleading users who trust LLM outputs. We investigate how simplification impacts homonym definition quality across three target groups: Normal, Simple, and ELI5. Using two novel evaluation datasets spanning multiple languages, we test DeepSeek v3, Llama 4 Maverick, Qwen3-30B A3B, GPT-4o mini, and Llama 3.1 8B via LLM-as-Judge and human annotations. Our results show that simplification drastically degrades definition completeness by neglecting polysemy, increasing the risk of misunderstanding. Fine-tuning Llama 3.1 8B with Direct Preference Optimization substantially improves homonym response quality across all prompt types. These findings highlight the need to balance simplicity and completeness in educational NLP to ensure reliable, context-aware definitions for all learners.

多义词语言模型教育NLP简化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。