构建百万级细粒度视频指令数据集,提升模型对复杂音视频内容的理解能力
Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
- 构建结构化多属性视频指令数据集,支持细粒度标注与多维度监督
- 通过自动验证与修正提升标注一致性,减少幻觉并增强指令遵循能力
- 开源模型在7个基准上表现领先,适合需要精准视频理解的场景
通用视频理解需在多样化现实场景中建模细粒度的视听信息。然而现有模型受限于将复杂音视频内容简化为单一不完整描述的视频-指令数据,缺乏细粒度组织与可靠标注。为此,我们提出:(i) ASID-1M,一个包含一百万条结构化、细粒度音视频指令标注的开源数据集,支持单属性与多属性监督;(ii) ASID-Verify,一种可扩展的数据清洗流程,具备自动验证与优化能力,确保描述与音视频内容在语义和时间上的连贯性;(iii) ASID-Captioner,基于ASID-1M通过监督微调训练的视频理解模型。在涵盖音视频描述、属性化描述、基于描述的问题回答与时间定位的七个基准测试中,ASID-Captioner在提升细粒度描述质量的同时,有效降低幻觉并增强指令遵循能力,其开源模型性能达到当前最优水平,且媲美Gemini-3-Pro。
原文摘要 · Abstract (English)
Universal video understanding requires modeling fine-grained visual and audio information over time in diverse real-world scenarios. However, the performance of existing models is primarily constrained by video-instruction data that represents complex audiovisual content as single, incomplete descriptions, lacking fine-grained organization and reliable annotation. To address this, we introduce: (i) ASID-1M, an open-source collection of one million structured, fine-grained audiovisual instruction annotations with single- and multi-attribute supervision; (ii) ASID-Verify, a scalable data curation pipeline for annotation, with automatic verification and refinement that enforces semantic and temporal consistency between descriptions and the corresponding audiovisual content; and (iii) ASID-Captioner, a video understanding model trained via Supervised Fine-Tuning (SFT) on the ASID-1M. Experiments across seven benchmarks covering audiovisual captioning, attribute-wise captioning, caption-based QA, and caption-based temporal grounding show that ASID-Captioner improves fine-grained caption quality while reducing hallucinations and improving instruction following. It achieves state-of-the-art performance among open-source models and is competitive with Gemini-3-Pro.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。