构建2400万样本的纯净视觉语言数据集,提升模型性能
FineVision: Open Data Is All You Need
- 整合200+来源,通过人机协作流程统一格式并验证数据质量
- 在66个基准上去污染,去重后模型表现优于现有开源数据集
- 适合追求高质量训练数据的研究者和数据工程团队
视觉语言模型的发展受限于公开数据集碎片化、不一致且含污染的问题。我们提出FineVision,一个经过精心收集、清洗和统一的2400万样本语料库,是目前最大规模的开放资源。通过半自动化、人机协作的流水线,将超过200个来源整合为185个子集:自动化完成批量摄入与模式映射,人工审核映射结果、抽样检查输出,确保标注忠实性、格式正确性、多样性及安全性;发现问题则触发针对性修复并重新运行。流程还对源内和跨源数据进行严格去重,并在66个公开基准上进行去污染处理。FineVision还包含基于代理/图形界面的任务,采用统一动作空间,人工验证模式并抽查轨迹以确认可执行性。在广泛评估中,使用FineVision训练的模型始终优于基于现有开源混合数据集训练的模型,凸显了规模、数据洁净度以及自动化与人工监督平衡的价值。我们公开该语料库及清洗工具,以加速以数据为中心的视觉语言模型研究。
原文摘要 · Abstract (English)
The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public datasets. We introduce FineVision, a meticulously collected, curated, and unified corpus of 24 million samples - the largest open resource of its kind. We unify more than 200 sources into 185 subsets via a semi-automated, human-in-the-loop pipeline: automation performs bulk ingestion and schema mapping, while reviewers audit mappings and spot-check outputs to verify faithful consumption of annotations, appropriate formatting and diversity, and safety; issues trigger targeted fixes and re-runs. The workflow further applies rigorous de-duplication within and across sources and decontamination against 66 public benchmarks. FineVision also encompasses agentic/GUI tasks with a unified action space; reviewers validate schemas and inspect a sample of trajectories to confirm executable fidelity. Models trained on FineVision consistently outperform those trained on existing open mixtures across a broad evaluation suite, underscoring the benefits of scale, data hygiene, and balanced automation with human oversight. We release the corpus and curation tools to accelerate data-centric VLM research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。