测试大模型在17种语言中的概念理解能力,发现低资源语言表现更差。
XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
- 构建跨17种语言的概念最小对数据集,评估模型理解力。
- 低资源语言概念理解差,复杂语言需更深层推理。
- 指令微调提升表现但不增强内在能力,知识蒸馏可改善低资源语言
本文提出XCOMPS,一个涵盖17种语言的概念最小对多语言基准数据集。通过元语言提示、直接概率测量和神经语言探针,评估大语言模型的多语言概念理解能力。对比基础模型、指令微调模型和知识蒸馏模型发现:1)低资源语言的概念理解能力较弱,相同概念集在不同语言间准确率差异明显;2)模型能有效区分明显不同的概念-属性对,但在语义细微相似的负样本对上性能显著下降;3)指令微调提升任务表现但未增强内部概念理解能力;知识蒸馏可提升低资源语言的内在理解能力,但显式任务性能提升有限;4)形态复杂的语言概念理解得分较低,且需要更深层网络进行概念推理。
原文摘要 · Abstract (English)
We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct probability measurement, and neurolinguistic probing. By comparing base, instruction-tuned, and knowledge-distilled models, we find that: 1) LLMs exhibit weaker conceptual understanding for low-resource languages, and accuracy varies across languages despite being tested on the same concept sets. 2) LLMs excel at distinguishing concept-property pairs that are visibly different but exhibit a marked performance drop when negative pairs share subtle semantic similarities. 3) Instruction tuning improves performance in concept understanding but does not enhance internal competence; knowledge distillation can enhance internal competence in conceptual understanding for low-resource languages with limited gains in explicit task performance. 4) More morphologically complex languages yield lower concept understanding scores and require deeper layers for conceptual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。