用高质量数据训练出性能媲美半开源模型的全开源多模态大模型。
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
- 构建1500万条高质量问答对,融合长短链思维增强数据
- 训练的Bee-8B模型超越多数全开源模型,接近半开源水平
- 开源全流程工具链,支持透明可复现的数据清洗与训练
全开源多模态大模型目前落后于闭源模型,主要因监督微调数据质量不足。现有开源数据集普遍存在噪声多、复杂推理数据缺失的问题,如思维链(CoT)数据稀缺。为此,本文提出三项贡献:首先,构建Honey-Data-15M数据集,包含约1500万条问答对,经多重清洗并采用新颖的双层级(短/长)思维链增强策略;其次,推出HoneyPipe数据清洗流水线及其底层框架DataStudio,提供可透明、可复用的数据处理方法;最后,基于该数据集训练Bee-8B模型。实验表明,Bee-8B在全开源模型中达到新SOTA,性能可比肩甚至超过部分半开源模型(如InternVL3.5-8B)。本工作向社区开放了包括Honey-Data-15M、HoneyPipe与DataStudio在内的全套资源、训练配方、评估工具及模型权重,证明数据质量是实现高性能全开源多模态大模型的关键路径。
原文摘要 · Abstract (English)
Fully open multimodal large language models (MLLMs) currently lag behind proprietary counterparts, primarily due to a significant gap in data quality for supervised fine-tuning (SFT). Existing open-source datasets are often plagued by widespread noise and a critical deficit in complex reasoning data, such as Chain-of-Thought (CoT), which hinders the development of advanced model capabilities. Addressing these challenges, our work makes three primary contributions. First, we introduce Honey-Data-15M, a new SFT dataset comprising approximately 15 million QA pairs, processed through multiple cleaning techniques and enhanced with a novel dual-level (short and long) CoT enrichment strategy. Second, we introduce HoneyPipe, the data curation pipeline, and its underlying framework DataStudio, providing the community with a transparent and adaptable methodology for data curation that moves beyond static dataset releases. Finally, to validate our dataset and pipeline, we train Bee-8B, an 8B model on Honey-Data-15M. Experiments show that Bee-8B establishes a new state-of-the-art (SOTA) for fully open MLLMs, achieving performance that is competitive with, and in some cases surpasses, recent semi-open models such as InternVL3.5-8B. Our work delivers to the community a suite of foundational resources, including: the Honey-Data-15M corpus; the full-stack suite comprising HoneyPipe and DataStudio; training recipes; an evaluation harness; and the model weights. This effort demonstrates that a principled focus on data quality is a key pathway to developing fully open MLLMs that are highly competitive with their semi-open counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。