arXiv:2504.11232cs.CVcs.MM2025-04

用模态特定标注提升视频理解模型的解释性与性能

Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset

  • 按视觉/文本/音频模态划分标注数据,构建专用训练集
  • 晚融合模型性能逼近早融合,显著优于传统训练
  • 适合追求可解释性的视频分析研究者

我们通过MOByGaze数据集(含人类标注的解释性概念)研究概念引导监督对多模态视频理解模型的影响。提出概念模态特定数据集(CMSDs),将标注概念按视觉、文本或音频模态分类形成子集。在早融合和晚融合方法中,使用CMSD训练的模型均优于传统训练方式。尤其值得注意的是,该方法使晚融合模型性能接近早融合模型。结果表明,模态特定标注对构建鲁棒、自解释的视频模型至关重要,推动复杂视频分析中可解释多模态学习的发展。

原文摘要 · Abstract (English)

We examine the impact of concept-informed supervision on multimodal video interpretation models using MOByGaze, a dataset containing human-annotated explanatory concepts. We introduce Concept Modality Specific Datasets (CMSDs), which consist of data subsets categorized by the modality (visual, textual, or audio) of annotated concepts. Models trained on CMSDs outperform those using traditional legacy training in both early and late fusion approaches. Notably, this approach enables late fusion models to achieve performance close to that of early fusion models. These findings underscore the importance of modality-specific annotations in developing robust, self-explainable video models and contribute to advancing interpretable multimodal learning in complex video analysis.

视频理解多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。