arXiv:2510.03762cs.CL2025-10中稿 · GlobalNLP 2025: Wo…被引 2

少样本提示中的数据不平衡会误导多语言词义消歧,影响模型表现。

Prompt Balance Matters: Understanding How Imbalanced Few-Shot Learning Affects Multilingual Sense Disambiguation in LLMs

  • 用不平衡样本进行少样本提示,导致多语言词义消歧出错。
  • 英语无此问题,但德、西、法、意四种语言均出现错误率上升。
  • 适用于关注多语言模型公平性与提示设计的研究者。

大型语言模型(LLMs)的进展显著改变了自然语言处理的格局。其中,少样本提示因其实用性和有效性受到广泛关注。本研究探讨了少样本提示策略对词义消歧(WSD)任务的影响,特别关注不平衡样本分布带来的偏差。我们采用先进的英文词义消歧提示方法 GLOSSGPT,测试其在五种语言(英语、德语、西班牙语、法语、意大利语)中的表现。结果显示,不平衡的少样本示例会导致多语言词义消歧出现错误预测,但在英语中未观察到此现象。我们评估了 GPT-4o 与 LLaMA-3.1-70B 模型,结果表明多语言词义消歧对少样本设置中的样本分布极为敏感,强调了采用平衡且具代表性的提示策略的重要性。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) have significantly reshaped the landscape of Natural Language Processing (NLP). Among the various prompting techniques, few-shot prompting has gained considerable attention for its practicality and effectiveness. This study investigates how few-shot prompting strategies impact the Word Sense Disambiguation (WSD) task, particularly focusing on the biases introduced by imbalanced sample distributions. We use the GLOSSGPT prompting method, an advanced approach for English WSD, to test its effectiveness across five languages: English, German, Spanish, French, and Italian. Our results show that imbalanced few-shot examples can cause incorrect sense predictions in multilingual languages, but this issue does not appear in English. To assess model behavior, we evaluate both the GPT-4o and LLaMA-3.1-70B models and the results highlight the sensitivity of multilingual WSD to sample distribution in few-shot settings, emphasizing the need for balanced and representative prompting strategies.

词义消歧多语言提示工程大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。