arXiv:2506.21011cs.CV2025-06中稿 · CVPR

用自动评分生成32万条视频质量指令,提升AI理解与解释能力。

Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring

  • 基于评分自动构建多维度视频质量指令数据
  • 生成超32万条指令对,支持模型高效训练
  • 适合研究视频理解、AI评估与指令微调的学者

传统视频质量评估仅输出单一数值,难以描述复杂质量维度。我们提出Score-based Instruction Generation(SIG)流水线,先对未标注视频的多个质量维度进行评分,并将分数映射为文本级别描述;再通过分层思维链建模各维度与整体质量的关系,模拟人类视觉系统推理过程。该自动化流程摆脱了人工撰写和专有系统的依赖,实现数据规模化与高效生成。由此构建的Score2Instruct数据集包含超过32万条多样化的指令-响应对,支撑视频大模型的指令微调。为进一步提升模型在质量评分与理由说明方面的能力,我们设计渐进式微调策略,并建立名为S2I-Bench的新基准,包含400个开放式问题,用于更全面评估视频大模型的质量解释能力。实验结果表明,该方法在多个视频大模型上均显著提升了质量评分与理由生成性能。

原文摘要 · Abstract (English)

Classical video quality assessment methods generate a numerical score to judge a video's perceived visual fidelity and clarity. Yet, a score fails to describe the video's complex quality dimensions, restricting its applicability. Benefiting from the human-friendly linguistic output, adapting video large multimodal models to VQA via instruction tuning has the potential to address this issue. The core of the approach lies in the video quality-centric instruction data. Previous explorations mainly focus on the image domain, and their data generation processes heavily rely on human quality annotations and proprietary systems, limiting data scalability and effectiveness. To address these challenges, we propose the Score-based Instruction Generation pipeline. Specifically, SIG first scores multiple quality dimensions of an unlabeled video and maps scores to text-defined levels. It then explicitly incorporates a hierarchical Chain-of-Thought to model the correlation between specific dimensions and overall quality, mimicking the human visual system's reasoning process. The automated pipeline eliminates the reliance on expert-written quality descriptions and proprietary systems, ensuring data scalability and generation efficiency. To this end, the resulting Score2Instruct dataset contains over 320K diverse instruction-response pairs, laying the basis for instruction tuning. Moreover, to advance video LMMs' quality scoring and justification abilities simultaneously, we devise a progressive tuning strategy to fully unleash the power of S2I. Built upon SIG, we further curate a benchmark termed S2I-Bench with 400 open-ended questions to better evaluate the quality justification capacity of video LMMs. Experimental results on the S2I-Bench and existing benchmarks indicate that our method consistently improves quality scoring and justification capabilities across multiple video LMMs.

视频质量指令生成多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。