arXiv:2504.07089cs.CVcs.CL2025-04被引 16

一个模型搞定各类图像的细粒度描述生成

OmniCaptioner: One Captioner to Rule Them All

  • 统一框架处理自然图像、图文内容和结构化图表
  • 长上下文提示提升大模型多模态推理能力
  • 适合需要跨视觉类型生成描述的研究与应用

我们提出 OmniCaptioner,一个通用的视觉描述生成框架,可对多种视觉领域生成细粒度文本描述。与仅限于特定图像类型(如自然图像或几何图形)的现有方法不同,该框架统一支持自然图像、视觉文本(如海报、UI界面、教科书)以及结构化视觉内容(如文档、表格、图表)。通过将低层像素信息转化为语义丰富的文本表示,该框架弥合了视觉与文本模态之间的鸿沟。实验表明其具有三大优势:(i) 增强的大模型视觉推理能力,长上下文描述使 DeepSeek-R1 系列等大模型在多模态场景中表现更优;(ii) 提升图像生成质量,详细描述有助于文本到图像生成与图像变换任务;(iii) 高效监督微调(SFT),可在较少数据下实现更快收敛。我们认为 OmniCaptioner 的多功能性与适应性为连接语言与视觉模态提供了新视角。

原文摘要 · Abstract (English)

We propose OmniCaptioner, a versatile visual captioning framework for generating fine-grained textual descriptions across a wide variety of visual domains. Unlike prior methods limited to specific image types (e.g., natural images or geometric visuals), our framework provides a unified solution for captioning natural images, visual text (e.g., posters, UIs, textbooks), and structured visuals (e.g., documents, tables, charts). By converting low-level pixel information into semantically rich textual representations, our framework bridges the gap between visual and textual modalities. Our results highlight three key advantages: (i) Enhanced Visual Reasoning with LLMs, where long-context captions of visual modalities empower LLMs, particularly the DeepSeek-R1 series, to reason effectively in multimodal scenarios; (ii) Improved Image Generation, where detailed captions improve tasks like text-to-image generation and image transformation; and (iii) Efficient Supervised Fine-Tuning (SFT), which enables faster convergence with less data. We believe the versatility and adaptability of OmniCaptioner can offer a new perspective for bridging the gap between language and visual modalities.

视觉描述多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。