arXiv:2602.14653cs.CL2026-02Conference of the …被引 4

研究发现,有视觉和语境支撑的语言信息分布更均匀。

Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?

  • 用多模态模型分析30种语言的图文数据,计算信息突变值。
  • 视觉和语境双重支撑使信息分布更平滑,全局与局部均匀性提升。
  • 适合关注语言认知、多模态交互的研究者阅读。

统一信息密度(UID)假说认为,说话者会受交际压力影响,在话语中均匀分配信息,以最小化意外程度的方差。尽管该假说已有实证检验,但以往研究仅限于纯文本输入,忽略了话语产生的感知背景。本文首次在视觉接地情境下进行UID的计算研究,利用多语言视觉-语言模型对30种语言的图文数据及13种语言的视觉叙事数据(涵盖11个语言家族)进行突变值估计。结果表明,感知接地能持续平滑信息分布,在多种语言类型中均提升了全局与局部均匀性,优于纯文本场景。在视觉叙事中,图像与语篇上下文的双重接地产生额外效应,突变值下降最显著发生在语篇单元起始处。本研究首次探索生态合理、多模态语言使用中的信息流时序动态,发现接地语言具有更高信息均匀性,支持一种情境敏感的UID形式。

原文摘要 · Abstract (English)

The Uniform Information Density (UID) hypothesis posits that speakers are subject to a communicative pressure to distribute information evenly within utterances, minimising surprisal variance. While this hypothesis has been tested empirically, prior studies are limited exclusively to text-only inputs, abstracting away from the perceptual context in which utterances are produced. In this work, we present the first computational study of UID in visually grounded settings. We estimate surprisal using multilingual vision-and-language models over image-caption data in 30 languages and visual storytelling data in 13 languages, together spanning 11 families. We find that grounding on perception consistently smooths the distribution of information, increasing both global and local uniformity across typologically diverse languages compared to text-only settings. In visual narratives, grounding in both image and discourse contexts has additional effects, with the strongest surprisal reductions occurring at the onset of discourse units. Overall, this study takes a first step towards modelling the temporal dynamics of information flow in ecologically plausible, multimodal language use, and finds that grounded language exhibits greater information uniformity, supporting a context-sensitive formulation of UID.

信息密度多模态语言语言认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。