arXiv:2604.05117cs.CVcs.AI2026-04被引 2

清理视频理解数据中的文本陷阱,提升模型真视觉理解能力。

Watch Before You Answer: Learning from Visually Grounded Post-Training

论文配图:Watch Before You Answer: Learning from Visually Grounded Post-Training
图 1 · 摘自论文原文
  • 用仅含视觉关联问题的数据做后训练,剔除语言暗示
  • 性能最高提升6.2分,仅用原数据69.1%即可达成
  • 强调数据质量比复杂算法更重要,适合想改进视觉推理的研究者

视觉语言模型(VLM)需全面理解视觉、时间与文本线索。然而,尽管多模态建模进展迅速,视频理解性能仍落后于文本推理。本文发现:主流长视频理解基准中40%-60%的问题仅靠文本即可回答,且此类问题在广泛使用的后训练数据集中也普遍存在,可能削弱后训练对视频理解的提升效果。为此,我们提出VidGround:仅使用真正需要视觉支撑的问题进行后训练。结合基于强化学习的后训练方法,该方案相比使用全量数据提升最高达6.2分,且仅需原数据的69.1%。此外,简单数据筛选的效果优于多种复杂后训练技术,凸显数据质量是提升视频理解的关键瓶颈。结果表明,必须构建真正要求视觉接地的后训练数据与评估基准,才能推动更强大VLM的发展。项目页:http://vidground.etuagi.com。

原文摘要 · Abstract (English)

It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video understanding performance still lags behind text-based reasoning. In this work, we find that progress is even worse than previously assumed: commonly reported long video understanding benchmarks contain 40-60% of questions that can be answered using text cues alone. Furthermore, we find that these issues are also pervasive in widely used post-training datasets, potentially undercutting the ability of post-training to improve VLM video understanding performance. Guided by this observation, we introduce VidGround as a simple yet effective solution: using only the actual visually grounded questions without any linguistic biases for post-training. When used in tandem with RL-based post-training algorithms, this simple technique improves performance by up to 6.2 points relative to using the full dataset, while using only 69.1% of the original post-training data. Moreover, we show that data curation with a simple post-training algorithm outperforms several more complex post-training techniques, highlighting that data quality is a major bottleneck for improving video understanding in VLMs. These results underscore the importance of curating post-training data and evaluation benchmarks that truly require visual grounding to advance the development of more capable VLMs. Project page: http://vidground.etuagi.com.

视觉语言模型视频理解数据清洗后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。