arXiv:2409.09611cs.CVcs.AI2024-09被引 3

融合音效与视觉,提升第一视角动作识别泛化能力

Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition

  • 结合运动、音频与外观特征,增强跨场景识别能力
  • 在ARGO1M数据集上达到当前最优性能
  • 适合关注多模态模型泛化的研究者

第一人称动作识别因可穿戴摄像头普及而快速发展,但面临不同环境间域偏移的挑战,如物体或背景场景差异。本文提出一种多模态框架,通过整合运动、音频与外观特征提升域泛化能力。关键贡献包括分析音频与运动特征对域偏移的鲁棒性,利用音频叙述增强音频-文本对齐,并通过音频与视觉叙述的一致性评分优化训练中音频的影响。该方法在ARGO1M数据集上实现当前最优性能,有效泛化至未见场景与地点。

原文摘要 · Abstract (English)

First-person activity recognition is rapidly growing due to the widespread use of wearable cameras but faces challenges from domain shifts across different environments, such as varying objects or background scenes. We propose a multimodal framework that improves domain generalization by integrating motion, audio, and appearance features. Key contributions include analyzing the resilience of audio and motion features to domain shifts, using audio narrations for enhanced audio-text alignment, and applying consistency ratings between audio and visual narrations to optimize the impact of audio in recognition during training. Our approach achieves state-of-the-art performance on the ARGO1M dataset, effectively generalizing across unseen scenarios and locations.

多模态动作识别域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。