将视觉与代码模型融合,实现多模态代码生成新突破
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- 用任务向量技术融合视觉与代码大模型,保留双项能力
- 在59.8万样本数据集上训练,性能逼近闭源模型GPT-4o
- 专为真实编程场景设计挑战性评测基准InfiBench-V
多模态大语言模型虽在视觉与文本理解方面进展显著,但其从多模态输入生成代码的能力仍受限。本文提出VisCodex,一种统一框架,通过任务向量驱动的模型融合技术,将先进代码大模型无缝整合至强大的视觉-语言主干网络中,同时保持出色的视觉理解与高级编码能力。为支持训练与评估,我们构建了大规模多样化的多模态编码数据集(MCD),包含59.8万条样本,涵盖高质量HTML代码、图表-代码配对、图像增强的StackOverflow问答及算法问题。此外,提出InfiBench-V这一新型挑战性基准,专门评估模型在富含视觉信息的真实编程问题上的表现,需综合理解文本与视觉上下文。大量实验表明,VisCodex在开源多模态模型中达到最先进水平,接近闭源模型GPT-4o的表现,验证了所提融合策略与数据集的有效性。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have significantly advanced the integration of visual and textual understanding. However, their ability to generate code from multimodal inputs remains limited. In this work, we introduce VisCodex, a unified framework that seamlessly merges vision and coding language models to empower MLLMs with strong multimodal code generation abilities. Leveraging a task vector-based model merging technique, we integrate a state-of-the-art coding LLM into a strong vision-language backbone, while preserving both visual comprehension and advanced coding skills. To support training and evaluation, we introduce the Multimodal Coding Dataset (MCD), a large-scale and diverse collection of 598k samples, including high-quality HTML code, chart image-code pairs, image-augmented StackOverflow QA, and algorithmic problems. Furthermore, we propose InfiBench-V, a novel and challenging benchmark specifically designed to assess models on visually-rich, real-world programming questions that demand a nuanced understanding of both textual and visual contexts. Extensive experiments show that VisCodex achieves state-of-the-art performance among open-source MLLMs and approaches proprietary models like GPT-4o, highlighting the effectiveness of our model merging strategy and new datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。