提升多模态模型性能关键在知识密度,而非任务多样性。
Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling

- 用结构化图文描述增强数据知识密度,效果优于增加任务类型。
- 从图文描述中可重构视觉问答,性能损失微乎其微。
- 适合追求高效训练和通用能力的多模态研究者参考。
多模态大语言模型虽进展迅速,但其缩放规律仍不明确且难以预测。增大模型规模和任务多样性常导致收益递减。本文认为,多模态缩放的主要瓶颈并非任务形式,而是训练数据的知识密度。我们首先证明,视觉问答(VQA)等任务特定监督提供的语义信息极少,远超图像描述:仅靠图文描述即可重构VQA信号,性能损失可忽略。随后,通过结构化描述增强与跨模态知识注入提升知识密度,模型在多模态及下游基准上均实现持续提升。控制实验表明,性能与语义覆盖度相关性高于任务多样性。结果表明当前多模态模型难以扩展,主要因训练数据知识覆盖不足。我们倡导以知识为中心的多模态训练,作为可扩展模型的坚实基础。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved rapid progress, yet their scaling behavior remains less clearly characterized and often less predictable than that of text-only LLMs. Increasing model size and task diversity often yields diminishing returns. In this work, we argue that the primary bottleneck in multimodal scaling is not task format, but knowledge density in training data. We first show that task-specific supervision such as Visual Question Answering (VQA) contributes little incremental semantic information beyond image captions: VQA signals can be reconstructed from captions with negligible performance loss. We then demonstrate that increasing knowledge density -- through structured caption enrichment and cross-modal knowledge injection -- leads to consistent performance improvements across multimodal and downstream benchmarks. Across controlled experiments, performance correlates more strongly with semantic coverage than with task diversity. These findings suggest that current MLLMs fail to scale primarily because training data lacks sufficient knowledge coverage. We advocate for knowledge-centric multimodal training as a principled foundation for scalable multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。