用开源模型实现媲美GPT-4的通用视觉描述,成本降89.5%
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- 设计多智能体协作流程CapFlow,整合开源模型生成高质量描述
- 在多个视觉领域实现与GPT-4相当的生成质量,成本降低89.5%
- 产出可大规模应用的通用视觉描述模型MetaCaptioner,适合研究者使用
通用视觉描述不仅要求外观描述,还需整合多种视觉线索并处理多样视觉域。当前开源模型性能远低于商用模型,限制了数据合成等应用。本文提出CapFlow多智能体协作流程,首次证明利用开源模型可实现与GPT-4.1相当的生成质量,同时成本降低89.5%。基于CapFlow构建大规模图像与视频领域高质量视觉描述数据集,并通过微调得到通用视觉描述模型MetaCaptioner。大量实验表明,MetaCaptioner不仅达到商用模型水平,还在开源社区中处于顶尖多模态性能。我们希望CapFlow与MetaCaptioner能为未来多模态研究提供高效可靠的视觉描述解决方案。
原文摘要 · Abstract (English)
Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains. In this task, current open-source models present a large performance gap with commercial ones, which limits various applications such as data synthesis. To bridge the gap, this paper proposes CapFlow, a novel multi-agent collaboration workflow. CapFlow demonstrates for the first time that, by capitalizing on open-source models, it is possible to achieve caption quality on par with GPT-4.1 in various domains with an 89.5% reduction in costs. By leveraging CapFlow as the data synthesizer, we produce high-quality visual captions from image and video domains at scale, and obtain a generalist visual captioner via fine-tuning, namely MetaCaptioner. Through extensive experiments, we show that MetaCaptioner not only achieves comparable captioning capabilities with commercial models but also reaches top-tier multimodal performance in the open-source community. We hope CapFlow and MetaCaptioner can benefit future multimodal research by providing a strong and cost-effective visual captioning solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。