arXiv:2601.02991cs.CVcs.AI2026-01

让小模型更准确理解漫画,通过多线索并行推理提升逻辑一致性。

Towards Faithful Reasoning in Comics for Small MLLMs

  • 构建模块化监督框架MoCoT,保留多线索联合推理过程
  • 在≤4B参数量下超越多个7B基线,平均提升12.1%
  • 适合追求推理真实性与效率的小模型应用

漫画理解对多模态大模型(MLLMs)构成重大挑战,因漫画的真正含义往往源于视觉、文本与社会线索的联合解读。这自然催生了思维链(CoT)提示方法,因其显式中间推理有望整合异构信号。然而现有CoT方法与漫画结构不匹配:它们倾向于在多线索联合分析前强制进入单一推理路径,导致性能下降,尤其影响小模型表现。本文提出两阶段框架以实现小模型的忠实漫画推理。首先引入MoCoT——一种模块化监督构建框架,保持多线索解释能力,并转化为更真实的监督信号;其次提出VERA,一种结构化奖励机制,使优化目标同时对齐推理忠实性与答案正确性。在五个涵盖漫画理解及幽默、抽象视觉推理任务的基准上实验表明,该框架在≤4B参数规模下表现优异,超越多个7B基线,使四款小模型平均提升12.1%,显著增强推理忠实性且保持推理效率。

原文摘要 · Abstract (English)

Comic understanding presents a significant challenge for Multimodal Large Language Models (MLLMs), as the intended meaning of a comic often emerges from the joint interpretation of visual, textual, and social cues. This naturally motivates Chain-of-Thought (CoT) prompting, since explicit intermediate reasoning appears promising for integrating such heterogeneous signals. However, existing CoT methods are poorly matched to this structure: they tend to force interpretation into a single reasoning path before multiple cues have been jointly considered, often degrading performance, especially for small MLLMs. Our key idea is to explicitly preserve multi-cue interpretation during supervision construction, rather than collapsing comic understanding into a single reasoning chain. To this end, we propose a two-stage framework for faithful comic reasoning in small MLLMs. First, we introduce MoCoT, a modular supervision construction framework that preserves multi-cue interpretation and turns it into more faithful supervision. Second, we propose VERA, a structured reward mechanism that turns such supervision into faithful reasoning behavior by aligning optimization with both reasoning faithfulness and answer correctness. Extensive experiments on five benchmarks spanning comic understanding and broader humor-centric and abstract visual reasoning tasks demonstrate that our framework achieves strong results in the $\leq$ 4B regime, surpasses several 7B baselines, improves four small MLLMs by an average of $\mathbf{12.1%}$ as a plug-in, and consistently enhances reasoning faithfulness while preserving inference efficiency.

漫画理解小模型推理忠实性CoT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。