构建通用影评理解框架,提升模型对电影语言的全面认知能力
Seeking Universal Shot Language Understanding Solutions
- 设计动态混合数据训练通用模型,兼顾多维度理解
- 在33个任务上实现22%超越商用模型的跨域性能
- 适合影视分析、跨模态理解等研究者使用
影评语言理解(SLU)对电影分析至关重要,但因摄影风格多样及专家主观判断差异而面临挑战。尽管视觉-语言模型(VLMs)在通用视觉理解中表现优异,但其与电影专家在SLU任务上仍存在判断偏差。为此,我们提出SLU-SUITE,一个包含490K人工标注问答对的综合训练与评估套件,覆盖33个任务、六大电影相关维度。基于该套件,我们首次发现:从模型角度看,可诊断关键模块瓶颈;从数据角度看,能量化各任务间的跨维度影响。据此提出两种互补方案:UniShot——通过动态平衡数据混合训练的通用模型;AgentShots——基于提示路由的专家集群,最大化单维度性能。大量实验表明,所提模型在本域任务上优于特定任务集成,在跨域任务上较领先商业VLM提升22%。
原文摘要 · Abstract (English)
Shot language understanding (SLU) is crucial for cinematic analysis but remains challenging due to its diverse cinematographic dimensions and subjective expert judgment. While vision-language models (VLMs) have shown strong ability in general visual understanding, recent studies reveal judgment discrepancies between VLMs and film experts on SLU tasks. To address this gap, we introduce SLU-SUITE, a comprehensive training and evaluation suite containing 490K human-annotated QA pairs across 33 tasks spanning six film-grounded dimensions. Using SLU-SUITE, we originally observe two insights into VLM-based SLU from: the model side, which diagnoses key bottlenecks of modules; the data side, which quantifies cross-dimensional influences among tasks. These findings motivate our universal SLU solutions from two complementary paradigms: UniShot, a balanced one-for-all generalist trained via dynamic-balanced data mixing, and AgentShots, a prompt-routed expert cluster that maximizes peak dimension performance. Extensive experiments show that our models outperform task-specific ensembles on in-domain tasks and surpass leading commercial VLMs by 22% on out-of-domain tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。