测试大模型对日语比较句的推理能力,发现提示设计影响巨大。
Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
- 构建日语比较句专用NLI数据集,评估大模型零样本与少样本表现
- 模型对提示格式敏感,且在处理日语特有语言现象时表现不佳
- 加入逻辑语义提示可提升模型在复杂推理任务中的准确率
大语言模型在自然语言推理(NLI)任务中表现优异,但涉及数值和逻辑表达的推理仍具挑战性。比较句是此类推理的关键语言现象,然而大模型在非主流训练语言如日语中的鲁棒性尚未充分研究。为此,我们构建了一个聚焦于比较句的日语NLI数据集,并在零样本与少样本设置下评估多种大语言模型。结果表明,模型性能在零样本设置中对提示格式敏感,少样本示例中的正确标签也会影响结果。此外,模型难以处理日语特有的语言现象。值得注意的是,包含逻辑语义表示的提示能帮助模型解决即使在少样本情况下也难以应对的推理问题。
原文摘要 · Abstract (English)
Large Language Models (LLMs) perform remarkably well in Natural Language Inference (NLI). However, NLI involving numerical and logical expressions remains challenging. Comparatives are a key linguistic phenomenon related to such inference, but the robustness of LLMs in handling them, especially in languages that are not dominant in the models' training data, such as Japanese, has not been sufficiently explored. To address this gap, we construct a Japanese NLI dataset that focuses on comparatives and evaluate various LLMs in zero-shot and few-shot settings. Our results show that the performance of the models is sensitive to the prompt formats in the zero-shot setting and influenced by the gold labels in the few-shot examples. The LLMs also struggle to handle linguistic phenomena unique to Japanese. Furthermore, we observe that prompts containing logical semantic representations help the models predict the correct labels for inference problems that they struggle to solve even with few-shot examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。