构建首个带线索标注的音视频计数基准,提升多模态模型计数能力。
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
- 基于长视频和人工标注线索,设计新计数评估基准
- 提出AV-Reasoner模型,在多个基准上达到最佳性能
- 适合研究多模态推理与计数任务的学者使用
尽管视频理解取得进展,当前多模态大模型在计数任务上仍表现不佳。现有基准存在视频过短、查询封闭、缺乏线索标注和多模态覆盖不足等问题。本文提出CG-AV-Counting,一个手工标注的线索接地计数基准,包含1,027个跨模态问题和5,845个标注线索,覆盖497段长视频。该基准支持黑盒与白盒评估,可全面测试端到端及基于推理的计数方法。为提升模型计数能力,我们提出AV-Reasoner,采用GRPO与课程学习训练,从相关任务中泛化计数能力。AV-Reasoner在多个基准上实现领先效果,验证强化学习的有效性。然而实验显示,在域外基准上,语言空间中的推理无法带来性能提升。代码与数据集已开源。
原文摘要 · Abstract (English)
Despite progress in video understanding, current MLLMs struggle with counting tasks. Existing benchmarks are limited by short videos, close-set queries, lack of clue annotations, and weak multimodal coverage. In this paper, we introduce CG-AV-Counting, a manually-annotated clue-grounded counting benchmark with 1,027 multimodal questions and 5,845 annotated clues over 497 long videos. It supports both black-box and white-box evaluation, serving as a comprehensive testbed for both end-to-end and reasoning-based counting. To explore ways to improve model's counting capability, we propose AV-Reasoner, a model trained with GRPO and curriculum learning to generalize counting ability from related tasks. AV-Reasoner achieves state-of-the-art results across multiple benchmarks, demonstrating the effectiveness of reinforcement learning. However, experiments show that on out-of-domain benchmarks, reasoning in the language space fails to bring performance gains. The code and benchmark have been released on https://av-reasoner.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。