首个统一图像与视频对话的多模态专家模型,实现跨模态理解新突破。
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
- 通过专用专家路由图像和视频数据,联合学习空间与时间特征
- 在AVSD和VisDial上零样本与微调均达新SOTA,性能显著提升
- 首次系统研究两类任务间的领域迁移,揭示相互受益潜力
我们提出V²Dial——一种面向图像与视频输入的新型专家模型,旨在同时处理多模态对话任务。现有模型多聚焦于简单任务(如VQA、VideoQA、视频-文本检索),忽视更具挑战性的对话类任务,如视觉对话与视频对话。且这两类任务虽具相似性,却长期独立发展,限制了应用潜力。为此,我们首次通过单一模型统一两类任务,利用专用专家分别处理图像与视频数据,联合学习其空间与时间特征,并通过匹配与对比学习对齐特征。此外,我们系统研究了两类任务间的领域偏移问题,探究其训练数据能否相互促进。在广泛使用的AVSD与VisDial数据集上的大量实验表明,本模型在四个基准上均取得新状态最优结果,涵盖零样本与微调设置。
原文摘要 · Abstract (English)
We present V$^2$Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conversational tasks. Current multimodal models primarily focus on simpler tasks (e.g., VQA, VideoQA, video-text retrieval) and often neglect the more challenging conversational counterparts, such as video and visual/image dialog. Moreover, works on both conversational tasks evolved separately from each other despite their apparent similarities limiting their applicability potential. To this end, we propose to unify both tasks using a single model that for the first time jointly learns the spatial and temporal features of images and videos by routing them through dedicated experts and aligns them using matching and contrastive learning techniques. Furthermore, we systemically study the domain shift between the two tasks by investigating whether and to what extent these seemingly related tasks can mutually benefit from their respective training data. Extensive evaluations on the widely used video and visual dialog datasets of AVSD and VisDial show that our model achieves new state-of-the-art results across four benchmarks both in zero-shot and fine-tuning settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。