arXiv:2502.10995cs.CL2025-02被引 2

评测大模型理解韩语间接言辞能力,发现现有模型仍有明显短板。

Evaluating Large language models on Understanding Korean indirect Speech acts

  • 基于对话上下文评估大模型对间接言辞的理解能力
  • Claude3-Opus表现最佳,多选题正确率71.94%,开放题65%
  • 多数模型在间接言辞上远低于人类水平,适合关注对话理解的研究者

准确理解话语意图在对话交流中至关重要。随着对话型人工智能模型在各领域快速应用,评估大语言模型(LLM)理解用户话语意图的能力变得尤为重要。本研究聚焦于在对话上下文中识别间接言辞——即实际意图与字面意思不一致的情况。实验结果表明,Claude3-Opus在多选题(MCQ)中达到71.94%正确率,在开放题(OEQ)中达65%,显著优于其他模型。总体而言,商用模型性能普遍高于开源模型,但所有模型均未达到人类水平。大多数模型在间接言辞理解上表现远差于直接言辞(意图明确表达时)。本研究不仅通过分析开放题回答模式对各模型的语用能力进行整体评估,更强调需进一步提升模型对间接言辞的理解,以实现更自然的人机对话。

原文摘要 · Abstract (English)

To accurately understand the intention of an utterance is crucial in conversational communication. As conversational artificial intelligence models are rapidly being developed and applied in various fields, it is important to evaluate the LLMs' capabilities of understanding the intentions of user's utterance. This study evaluates whether current LLMs can understand the intention of an utterance by considering the given conversational context, particularly in cases where the actual intention differs from the surface-leveled, literal intention of the sentence, i.e. indirect speech acts. Our findings reveal that Claude3-Opus outperformed the other competing models, with 71.94% in MCQ and 65% in OEQ, showing a clear advantage. In general, proprietary models exhibited relatively higher performance compared to open-source models. Nevertheless, no LLMs reached the level of human performance. Most LLMs, except for Claude3-Opus, demonstrated significantly lower performance in understanding indirect speech acts compared to direct speech acts, where the intention is explicitly revealed through the utterance. This study not only performs an overall pragmatic evaluation of each LLM's language use through the analysis of OEQ response patterns, but also emphasizes the necessity for further research to improve LLMs' understanding of indirect speech acts for more natural communication with humans.

语言理解间接言辞大模型评测韩语处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。