评测大模型对马来口语语气词的处理能力,发现需结构化引导才能提升表现。
Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay

- 构建马来语口语语气词评测基准MalayPrag,系统评估大模型能力。
- 十款现成大模型在识别语气词语用功能时准确率普遍偏低。
- 提出五维语用属性框架,显著提升模型对语气词的理解能力。
语气词如'well'和'kind of'是使大语言模型更像人类交流的关键成分,用于表达情感、意图和人际态度。然而,现有研究尚未建立对大模型处理语气词能力的全面理解,且多数聚焦于英语等高资源语言,忽视东南亚语言。本文提出MalayPrag基准,系统评估大模型在口语马来语中处理语气词的能力;引入五个理论基础一致的语用属性,构建统一解释框架。通过提示十款现成大模型完成三项预测任务,实验显示当前模型在连接语气词与其语用功能方面存在显著挑战。而本研究设计的五维属性显著改善了这一连接,凸显了为模型提供结构化支架对其语用能力的重要性。
原文摘要 · Abstract (English)
Discourse particles, such as well and kind of, are crucial components that enable LLMs to "speak" more like humans. They are used to convey emotions, intentions, and interpersonal attitudes. However, existing studies have not yet built a comprehensive understanding of LLMs' capabilities in handling discourse particles. Moreover, the limited number of research focuses primarily on high-resource languages such as English, with little attention paid to Southeast Asian languages. In this paper, we (1) propose MalayPrag, a benchmark designed to systematically evaluate and analyze LLMs' capabilities in handling discourse particles in colloquial Malay; (2) introduce five attributes that provide a theoretically grounded, unified framework for interpreting pragmatic functions of discourse particles. Applying these two, we prompt ten off-the-shelf LLMs to perform three prediction tasks. The experimental results reveal substantial challenges for current LLMs to accurately connect discourse particles and their pragmatic functions in Malay. The provision of the five attributes designed in this study is found to significantly improve the connections, highlighting the need for structured scaffolding for models' pragmatic competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。