构建合成数据集评估视觉语言时间对齐能力
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
- 用可控生成方法构建平衡的合成视频-语言数据
- 揭示现有模型在时间对齐上的分布偏移敏感性
- 适合研究多模态时序理解与公平评估的学者
视觉-语言时间对齐是真实场景中人类动态识别与认知的关键能力。现有研究虽关注视觉-语言关联,但受限于时间分布偏差、标注不精确和组合多样性不足。为实现公平评估与全面探索,本文聚焦模型在时间维度上将视觉场景与语言上下文同步的能力。首先分析现有基准的统计特性,揭示问题根源。随后提出SVLTA——基于模拟环境的合成视觉-语言时间对齐数据集,通过融合常识知识、可操控动作与约束过滤,生成合理、多样且均衡的数据分布,支持诊断性评估。实验在时间问答、分布偏移敏感性及时间对齐适应性方面揭示关键洞见。
原文摘要 · Abstract (English)
Vision-language temporal alignment is a crucial capability for human dynamic recognition and cognition in real-world scenarios. While existing research focuses on capturing vision-language relevance, it faces limitations due to biased temporal distributions, imprecise annotations, and insufficient compositionally. To achieve fair evaluation and comprehensive exploration, our objective is to investigate and evaluate the ability of models to achieve alignment from a temporal perspective, specifically focusing on their capacity to synchronize visual scenarios with linguistic context in a temporally coherent manner. As a preliminary step, we present the statistical analysis of existing benchmarks and reveal the existing challenges from a decomposed perspective. To this end, we introduce SVLTA, the Synthetic Vision-Language Temporal Alignment derived via a well-designed and feasible control generation method within a simulation environment. The approach considers commonsense knowledge, manipulable action, and constrained filtering, which generates reasonable, diverse, and balanced data distributions for diagnostic evaluations. Our experiments reveal diagnostic insights through the evaluations in temporal question answering, distributional shift sensitiveness, and temporal alignment adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。