arXiv:2608.06930cs.CV2026-08

构建首个细粒度音视频联合描述数据集与评估体系

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

论文配图:AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
图 1 · 摘自论文原文
  • 提出细粒度奖励机制,提升音视频联合生成精度
  • 在10万条对齐数据上实现领先性能,超越部分闭源模型
  • 专为原子级细节评估设计,适合多模态生成研究者

细粒度音视频联合描述对多模态视频理解与生成至关重要。现有方法受限于三大问题:(1) 缺乏高质量、具备精细音视频联合标注的公开数据集;(2) 基于粗粒度奖励信号的强化学习方法;(3) 缺少针对原子级别细节的评估基准与指标。为此,我们提出:(1) AVCap-100K,一个包含10万条时间对齐、细节丰富的音视频联合描述数据集;(2) AVCap模型,通过细粒度感知的GRPO(Da-GRPO)优化,在开源模型中达到最先进水平,并在多项评测中匹配或超越闭源模型表现;(3) AVCap-Bench与AVCap-Score,专门用于评估音视频描述中原子级别细节的基准与指标。代码、模型与数据集已公开于https://huggingface.co/collections/Apryle/avcap。

原文摘要 · Abstract (English)

Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main limitations: (1) the scarcity of high-quality public datasets with fine-grained audio-visual joint captions; (2) reinforcement-learning methods that rely on coarse reward signals; and (3) the lack of a benchmark and metric for evaluating detailed audiovisual captions at the atomic level. To address these challenges, we propose: (1) AVCap-100K, a high-quality dataset of 100K temporally aligned, detail-rich audio-video captions; (2) AVCap, a model optimized via Detail-Aware GRPO (Da-GRPO) that achieves state-of-the-art performance among open-source models and matches or surpasses proprietary models on several evaluations; and (3) AVCap-Bench and AVCap-Score, a specialized benchmark and metric for evaluating atomic-level details in audiovisual captions. Our code, models, and datasets are available at https://huggingface.co/collections/Apryle/avcap.

音视频生成强化学习评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。