arXiv:2607.05222cs.CV2026-07

提出五类图表与图像协同推理模式,帮读者理解科学图示的深层逻辑。

A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication

论文配图:A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication
图 1 · 摘自论文原文
  • 构建从R1到R5的五类推理缺口分类体系,刻画图文协作机制。
  • 在79篇脑损伤论文中验证,该分类能精准预测专家与非专家理解差异。
  • 适用于提升科研图表设计,弥合专业与普通读者的认知鸿沟。

科学论文中的图表与图像常共同出现,但现有计算研究未系统分析其一致性。本文提出一个基于沟通接地理论的多模态推理类型学,将图表、图像与标题构成的多模态单元所要求的推断工作系统化分类为R1至R5五类:包括重复数据、量化图像结构、投影图像内容至外部变量、审计图像主张,以及联合构建单一画面无法独立成立的论证框架。该类型学由神经科学专家从79篇创伤性脑损伤论文中提取的32组图表-图像对,通过自下而上的归纳方法建立。实验表明,在25组配对的视觉-语言模型描述评估中,该分类能有效预测领域专家与三名非专家判断的一致性与分歧点,明确指出哪些部分依赖上下文知识而非图像本身维持连贯性。该框架为图示设计者提供了系统工具,平衡文字与图文组合的关系,促进科学发现的有效传播。

原文摘要 · Abstract (English)

Charts and images appear together throughout scientific publications, yet most computational work does not characterize their coherence. We argue that a chart, its accompanying image, and the caption that links them form a multimodal unit, and that the inferential work required to read it varies systematically. To capture this variation, we develop a typology of reasoning gaps, R1 through R5, that characterizes how chart, image, and text jointly convey a scientific claim, and the interpretive work this demands of the reader. Some pairs restate the same data, while in other pairs, charts are used to quantify a structure the image localizes, project image content onto an external variable, audit an image-based claim, or jointly construct a frame that neither panel can establish alone. The typology is anchored in the grounding theory of communication and was derived bottom-up, with a neuroscience expert, from a corpus of 79 traumatic brain injury papers and 32 chart-image pairs. Crucially, the levels provide a systematic mechanism for identifying where grounding succeeds or breaks down, rather than leaving it to subjective inference. We show this in a study in which a domain expert and three non-experts judge vision-language model (VLM) descriptions of 25 pairs: the level predicts where their judgments align and where they diverge, isolating the points at which contextual knowledge, not the figure, carries coherence. This typology thus offers figure designers a systematic way to balance text against chart-image pairs, bridging the expert-to-non-expert divide in reading a scientific takeaway.

多模态推理科学可视化图文一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。