arXiv:2511.06490cs.CVcs.AI2025-11被引 1

用区域感知强化学习提升视觉语言模型对漫画的细粒度理解

Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models

  • 设计区域感知强化学习,让模型动态聚焦漫画关键区域
  • 在漫画角色识别和情节排序上,性能显著优于传统方法
  • 适合研究漫画理解、多模态推理或视觉语言模型优化的人群

漫画等复杂视觉叙事对视觉语言模型(VLMs)构成重大挑战。尽管在自然图像上表现优异,现有VLMs在风格化线稿、拟声词和密集分镜布局上仍存在明显缺陷。为此,我们提出AI4VA-FG,首个面向VLM漫画理解的细粒度综合基准,涵盖从基础识别到角色推理与叙事构建的任务,并提供角色、姿态和深度的密集标注。评估GPT-4o、Gemini-2.5及Qwen2.5-VL等主流模型发现其在核心任务上存在显著性能差距,表明漫画理解仍是未解难题。我们系统研究了后训练策略,包括基于答案的监督微调(SFT-S)、基于推理轨迹的监督微调(SFT-R)及强化学习(RL)。受“以图思考”范式启发,提出区域感知强化学习(RARL),训练模型通过缩放操作动态关注相关区域。实验显示,将该方法应用于Qwen2.5-VL时,低层实体识别与高层情节排序性能均获得显著提升。

原文摘要 · Abstract (English)

Complex visual narratives, such as comics, present a significant challenge to Vision-Language Models (VLMs). Despite excelling on natural images, VLMs often struggle with stylized line art, onomatopoeia, and densely packed multi-panel layouts. To address this gap, we introduce AI4VA-FG, the first fine-grained and comprehensive benchmark for VLM-based comic understanding. It spans tasks from foundational recognition and detection to high-level character reasoning and narrative construction, supported by dense annotations for characters, poses, and depth. Beyond that, we evaluate state-of-the-art proprietary models, including GPT-4o and Gemini-2.5, and open-source models such as Qwen2.5-VL, revealing substantial performance deficits across core tasks of our benchmarks and underscoring that comic understanding remains an unsolved challenge. To enhance VLMs' capabilities in this domain, we systematically investigate post-training strategies, including supervised fine-tuning on solutions (SFT-S), supervised fine-tuning on reasoning trajectories (SFT-R), and reinforcement learning (RL). Beyond that, inspired by the emerging "Thinking with Images" paradigm, we propose Region-Aware Reinforcement Learning (RARL) for VLMs, which trains models to dynamically attend to relevant regions through zoom-in operations. We observe that when applied to the Qwen2.5-VL model, RL and RARL yield significant gains in low-level entity recognition and high-level storyline ordering, paving the way for more accurate and efficient VLM applications in the comics domain.

漫画理解强化学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。