梳理大模型时代语用评估资源,揭示评测短板与方向
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges
- 按语用现象分类整理评测数据集,涵盖隐含意义与指代等
- 发现现有评测在真实场景适配性上存在明显不足
- 适合关注大模型语境理解能力的研究者参考
理解语用——即语言在语境中的使用——对于构建能够解析语言细微差别的自然语言系统至关重要。尽管近年来语言技术(包括大语言模型)取得进展,但评估模型处理隐含意义、指代等语用现象的能力仍具挑战性。本综述全面回顾了用于评估NLP模型语用能力的资源,按其所针对的语用现象进行分类。分析了任务设计、数据收集方法、评估方式及其对现实应用的相关性。通过结合现代语言模型的背景,揭示了当前评测的趋势、挑战与空白。旨在厘清语用评估现状,指导更全面、有针对性的基准建设,推动发展更具语境感知能力的自然语言模型。
原文摘要 · Abstract (English)
Understanding pragmatics-the use of language in context-is crucial for developing NLP systems capable of interpreting nuanced language use. Despite recent advances in language technologies, including large language models, evaluating their ability to handle pragmatic phenomena such as implicatures and references remains challenging. To advance pragmatic abilities in models, it is essential to understand current evaluation trends and identify existing limitations. In this survey, we provide a comprehensive review of resources designed for evaluating pragmatic capabilities in NLP, categorizing datasets by the pragmatic phenomena they address. We analyze task designs, data collection methods, evaluation approaches, and their relevance to real-world applications. By examining these resources in the context of modern language models, we highlight emerging trends, challenges, and gaps in existing benchmarks. Our survey aims to clarify the landscape of pragmatic evaluation and guide the development of more comprehensive and targeted benchmarks, ultimately contributing to more nuanced and context-aware NLP models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。