arXiv:2608.30270cs.CL2026-09ACL

构建多模态基准READI,评估视觉语境中的间接言语行为理解能力

Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts

论文配图:Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts
图 1 · 摘自论文原文
  • 基于语用学理论构建分级间接性模型,融合视觉与对话上下文
  • 现有顶尖多模态模型在高间接性任务上性能显著下降
  • 支持中英文跨语言评估,推动高语境语言的语用理解研究

间接言语行为(ISAs)需要基于上下文进行语用推理,其指令意图无法仅从表面形式推断。以往文本研究及现有多模态基准大多忽略这一需求,侧重显式上下文或感知识别,因而未能充分探索依赖上下文的语用理解,尤其在韩语等高语境语言中表现不足。本文提出READI,一个通过整合视觉上下文与对话进行联合推理的多模态基准,用于评估ISA理解。READI基于语用学理论建模间接性的等级,并将任务定义为基于视觉的语用问答(V-PQA),支持英语与韩语的跨语言评估。实验表明,即使最先进的多模态模型在视觉引导的间接言语行为理解上仍表现不佳,且随着间接程度增加,性能持续下降,凸显了专门针对上下文语用推理的评测基准的重要性。

原文摘要 · Abstract (English)

Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this requirement, focusing instead on explicitly encoded context or perceptual recognition, and thus underex- plore context-dependent pragmatic understand- ing, particularly in high-context languages such as Korean. We introduce READI, a multimodal benchmark for evaluating ISA understanding through integrated reasoning over visual con- text and dialogue. READI models graded in- directness grounded in pragmatic theory and formulates the task as vision-based pragmatic question answering (V-PQA), supporting cross- lingual evaluation in English and Korean. Ex- periments show that even state-of-the-art multi- modal models struggle with visually grounded indirect speech acts, with performance declin- ing as indirectness increases, underscoring the need for benchmarks that explicitly target con- textual pragmatic reasoning.

多模态语用理解间接言语视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。