用漫画结构提升多模态推理效率与准确性
Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling
- 以漫画为媒介,融合时间序列与文本信息,构建高效视觉推理路径
- 在多步时序与因果推理任务中超越静态图像,且推理成本远低于视频
- 不同漫画叙事风格影响表现,适合作为视觉推理的中间表示范式
思维链推理已推动大语言模型从文本推理拓展至图像与视频。然而,静态图像难以表达时间结构,视频则带来显著冗余与计算开销。本文提出“思考漫画”(Thinking with Comics)的视觉推理范式,利用漫画这一信息密度高、介于图像与视频之间的媒介,保留时间结构、嵌入文本与叙事连贯性,同时大幅降低推理成本。我们系统研究了两种基于漫画的推理路径,并在多个推理任务和长上下文理解任务上进行评估。实验表明,相较于图像推理,在多步时序与因果推理任务中,漫画推理表现更优;而相比视频推理,其效率显著更高。进一步分析显示,不同漫画叙事结构与风格对任务表现具有稳定影响,证明漫画是提升多模态推理的有效中间视觉表征。
原文摘要 · Abstract (English)
Chain-of-Thought reasoning has driven large language models to extend from thinking with text to thinking with images and videos. However, different modalities still have clear limitations: static images struggle to represent temporal structure, while videos introduce substantial redundancy and computational cost. In this work, we propose Thinking with Comics, a visual reasoning paradigm that uses comics as a high information-density medium positioned between images and videos. Comics preserve temporal structure, embedded text, and narrative coherence while requiring significantly lower reasoning cost. We systematically study two reasoning paths based on comics and evaluate them on a range of reasoning tasks and long-context understanding tasks. Experimental results show that Thinking with Comics outperforms Thinking with Images on multi-step temporal and causal reasoning tasks, while remaining substantially more efficient than Thinking with Video. Further analysis indicates that different comic narrative structures and styles consistently affect performance across tasks, suggesting that comics serve as an effective intermediate visual representation for improving multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。