arXiv:2605.28346cs.CL2026-05

检验视觉语言模型如何组织信息以符合对话语境。

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs

论文配图:When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs
图 1 · 摘自论文原文
  • 用匈牙利语语法显化信息结构,检测模型对旧话题与新焦点的区分能力。
  • 模型过度规整化表达,仅依赖少数固定句式,缺乏人类的灵活策略。
  • 研究提示评估模型需关注信息组织方式,而不仅是内容正确性。

视觉语言模型(VLMs)常被评估是否识别正确的视觉内容,但对其内容表达是否符合对话语境知之甚少。本文通过信息结构(IS)研究这一空白,测试VLMs在视觉问答任务中能否区分话语中已知的‘话题’(Topic)与新提及的‘焦点’(Focus)。利用匈牙利语中话题与焦点对应特定句法位置的特点,使信息结构选择可观察。对比六种VLMs与人类参与者发现,模型虽能生成符合信息结构的表达,但过度规整化,表现出类似模式坍缩(mode collapse)的现象。在话语状态、语法角色(偏好主语作话题)和限定性(偏好不定冠词作焦点)的多重压力下,人类采用多样的实现策略,而模型则局限于少数固定响应模板。结果表明,评估VLMs不应仅关注内容准确性,还需考察其内容在对话中的组织方式。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly evaluated for whether they identify the right visual content, but little is known about whether they express such content in a discourse-appropriate form. We address this research gap using information structure (IS), testing whether VLMs distinguish discourse-old Topics from discourse-new Foci in visually grounded question answering. We exploit Hungarian, a language in which Topic and Focus map onto dedicated syntactic positions, making IS choices observable in text. Comparing six VLMs with human participants, we find that models produce IS-relevant constructions, but over-regularise this sensitivity. Under the interacting pressures of discourse status, grammatical role (preference for subject Topics) and definiteness (preference for indefinite Foci), humans choose variable strategies for IS realisation. VLMs, by contrast, collapse onto narrow response templates, resembling mode collapse (Kirk et al., 2024). Our findings suggest that VLM evaluation should look beyond content accuracy to how content is packaged for the discourse.

视觉语言模型信息结构对话语境语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。