1.7B参数小模型也能生成细节丰富的图像描述
VisionPangu: A Compact and Fine-Grained Multimodal Assistant with 1.7B Parameters
- 用轻量级投影融合视觉与语言主干,提升多模态对齐效率
- 基于DOCCI数据集的密集人工描述训练,显著提升描述丰富度
- 适合资源有限但需高质量图文生成的场景
大型多模态模型在视觉-语言理解上表现优异,但多数依赖大规模架构和粗粒度监督,限制了细节图像描述能力。本文提出VisionPangu,一个仅1.7B参数的紧凑型多模态模型,通过高效多模态对齐和高质量监督,提升细节图像描述质量。模型采用基于InternVL的视觉编码器与OpenPangu-Embedded语言主干,通过轻量级MLP投影连接,并借鉴LLaVA的指令微调流程。引入来自DOCCI数据集的密集人工描述进行训练,在不依赖激进模型扩展的前提下,显著提升语义连贯性与描述丰富度。实验表明,紧凑模型仍可达到竞争性性能,生成更结构化、更详细的图像描述。代码与模型权重将公开于https://www.modelscope.cn/models/asdfgh007/visionpangu。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have achieved strong performance in vision-language understanding, yet many existing approaches rely on large-scale architectures and coarse supervision, which limits their ability to generate detailed image captions. In this work, we present VisionPangu, a compact 1.7B-parameter multimodal model designed to improve detailed image captioning through efficient multimodal alignment and high-quality supervision. Our model combines an InternVL-derived vision encoder with the OpenPangu-Embedded language backbone via a lightweight MLP projector and adopts an instruction-tuning pipeline inspired by LLaVA. By incorporating dense human-authored descriptions from the DOCCI dataset, VisionPangu improves semantic coherence and descriptive richness without relying on aggressive model scaling. Experimental results demonstrate that compact multimodal models can achieve competitive performance while producing more structured and detailed captions. The code and model weights will be publicly available at https://www.modelscope.cn/models/asdfgh007/visionpangu.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。