评测大模型对拉脱维亚和立陶宛短答案匹配的细微差异识别能力。
The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching
- 构建了包含502组拉脱维亚和690组立陶宛问答对的新数据集。
- 72b和70b级大模型几乎完美区分细微文本差异,小模型表现波动大。
- 小型模型如Mistral 7b在少量示例下表现较弱,适合多语言细节检测研究者。
本文针对拉脱维亚和立陶宛语言的短答案匹配任务,提出新数据集,包含502个拉脱维亚和690个立陶宛问题-答案对。每对通过特定修改规则生成匹配与非匹配答案,以测试大模型对细微语义差异的识别能力。部分数据经人工验证。结果显示,大型模型如QWEN2.5 72b和LLaMa3.1 70b在区分匹配与非匹配答案上接近完美;而小模型如LLaMa3.1 8b和EuroLLM 9b在少量示例下表现提升,但Mistral Nemo 12b在立陶宛语中仍表现不佳,即使提供额外示例。值得注意的是,QWEN2.5 7b和Mistral 7b在零样本和少样本实验中表现强劲,可媲美70b级模型,但后者在少样本设置下性能下降。
原文摘要 · Abstract (English)
In this work, we address the challenge of evaluating large language models (LLMs) on the short answer matching task for Latvian and Lithuanian languages. We introduce novel datasets consisting of 502 Latvian and 690 Lithuanian question-answer pairs. For each question-answer pair, we generated matched and non-matched answers using a set of alteration rules specifically designed to introduce small but meaningful changes in the text. These generated answers serve as test cases to assess the ability of LLMs to detect subtle differences in matching of the original answers. A subset of the datasets was manually verified for quality and accuracy. Our results show that while larger LLMs, such as QWEN2.5 72b and LLaMa3.1 70b, demonstrate near-perfect performance in distinguishing matched and non-matched answers, smaller models show more variance. For instance, LLaMa3.1 8b and EuroLLM 9b benefited from few-shot examples, while Mistral Nemo 12b underperformed on detection of subtle text alteration, particularly in Lithuanian, even with additional examples. QWEN2.5 7b and Mistral 7b were able to obtain a strong and comparable performance to the larger 70b models in zero and few shot experiments. Moreover, the performance of Mistral 7b was weaker in few shot experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。