让视频描述模型自我反思进化,无需人工标注
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
- 用自我评分轨迹构建数据集,引导模型自省优化
- 在VDC和DREAM-1K上达顶尖水平,细节与忠实度双优
- 内部化反思能力,推理快且泛化强,适合实用部署
现有视频详细描述方法严重依赖昂贵的人工标注或从大型专有模型蒸馏,造成对外部监督的依赖。本文提出VDC-Agent,一种自主自演化的框架,使单个多模态大语言模型通过原则指导的自我反思生成并优化高质量描述。为解决迭代优化带来的推理延迟,我们进一步将反思能力内化至模型中。具体地,构建了基于代理自评分轨迹的VDC-Agent-19K偏好数据集,并引入课程式直接偏好优化(Curriculum DPO)策略,利用生成候选间的质量差距,从易到难逐步对齐模型。大量实验表明,VDC-Agent在VDC和DREAM-1K基准上达到当前最优性能,生成的描述具备更丰富的细节与更高的忠实度。关键的是,我们的内化策略在保持基础模型推理效率的同时,显著提升了其泛化能力,经量化指标与人工评估双重验证。
原文摘要 · Abstract (English)
Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary models, creating a dependency on external supervision. In this paper, we propose VDC-Agent, an autonomous self-evolving framework that empowers a single Multimodal Large Language Model (MLLM) to generate and refine high-quality captions through principle-guided self-reflection. To overcome the inference latency inherent in iterative refinement, we further propose to internalize this reflective capability into the model. Specifically, we construct VDC-Agent-19K, a preference dataset derived from the agent's self-scored trajectories, and introduce a Curriculum Direct Preference Optimization (DPO) strategy. This strategy leverages the quality gap between generated candidates to progressively align the model from easy to hard samples. Extensive experiments demonstrate that VDC-Agent achieves state-of-the-art performance on VDC and DREAM-1K benchmarks, generating captions with superior detail and faithfulness. Crucially, our internalization strategy retains the inference efficiency of the base model while significantly enhancing its generalization capabilities, as validated by both quantitative metrics and human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。