arXiv:2411.19434cs.CVcs.CL2024-11

分离动作与物体特征,实现无需训练的视频问答跨域泛化

Actions and Objects Pathways for Domain Adaptation in Video Question Answering

  • 将预训练特征拆分为动作和物体路径,分途推理增强泛化能力
  • 在未见领域上比传统分类器高5%,在已见领域高4%
  • 仅训练极少参数,却超越需数百万参数的已有方法

本文提出动作与物体路径(AOPath),用于视频问答任务中的跨域泛化。AOPath利用大规模预训练模型的特征,在无需对未见领域显式训练的情况下提升泛化性能。受人类大脑启发,该方法将预训练特征解耦为动作和物体特征,并通过独立的推理路径进行处理。引入一种新模块,可在不增加可训练参数的前提下,将跨域特征转换为领域无关特征。在基于题材划分的TVQA数据集上验证,所提方法在未见领域上较传统分类器提升5%,在已见领域提升4%;同时优于需训练数百万参数的先前方法,而自身仅训练极少量参数。

原文摘要 · Abstract (English)

In this paper, we introduce the Actions and Objects Pathways (AOPath) for out-of-domain generalization in video question answering tasks. AOPath leverages features from a large pretrained model to enhance generalizability without the need for explicit training on the unseen domains. Inspired by human brain, AOPath dissociates the pretrained features into action and object features, and subsequently processes them through separate reasoning pathways. It utilizes a novel module which converts out-of-domain features into domain-agnostic features without introducing any trainable weights. We validate the proposed approach on the TVQA dataset, which is partitioned into multiple subsets based on genre to facilitate the assessment of generalizability. The proposed approach demonstrates 5% and 4% superior performance over conventional classifiers on out-of-domain and in-domain datasets, respectively. It also outperforms prior methods that involve training millions of parameters, whereas the proposed approach trains very few parameters.

视频问答跨域泛化特征解耦少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。