arXiv:2603.16970cs.CVcs.AI2026-03

让视觉和动作传感器协同判断新行为,提升第一人称活动识别的可靠性。

MAND: Modality-Aware Novelty Detection for Open-World Egocentric Activity Recognition

  • 通过自适应调整各模态贡献,更公平利用视觉与惯性数据。
  • 在公开数据集上显著降低误报率95%(FPR95),同时保持旧类别准确率。
  • 适合需要持续学习新动作的智能眼镜、可穿戴设备等开放世界场景。

多模态第一人称活动识别融合视觉与惯性信号以实现鲁棒的行为理解。然而,在开放世界环境中部署此类系统需在持续学习非平稳数据流的同时检测新活动。现有方法仅依赖融合后的分类结果进行新颖性评分,未能充分利用各模态的互补信息。由于这些结果常被RGB主导,其他模态(尤其是IMU)的线索被严重低估,且随着灾难性遗忘积累,这一失衡加剧。为此,我们提出MAND,一种面向多模态第一人称开放世界持续学习的模态感知框架。推理时,模态感知自适应评分(MoAS)基于样本级可靠性动态调整模态权重,并引入偏差与分歧惩罚来优化新颖性评分。训练阶段,模态感知表征稳定训练(MoRST)通过模态专用头和模态级logit蒸馏,保持各模态跨任务的判别能力。在公开多模态第一人称基准上的实验表明,MAND在持续提升新活动检测性能的同时,显著降低FPR95,实现更可靠的开放世界识别。源代码已开源。

原文摘要 · Abstract (English)

Multimodal egocentric activity recognition integrates visual and inertial cues for robust first-person behavior understanding. However, deploying such systems in open-world environments requires detecting novel activities while continuously learning from non-stationary data streams. Existing methods rely on the main fused logits for novelty scoring, without fully exploiting the complementary evidence available from individual modalities. Because these logits are often dominated by RGB, cues from other modalities, particularly IMU, remain underutilized, and this imbalance worsens as catastrophic forgetting accumulates. To address this, we propose MAND, a modality-aware framework for multimodal egocentric open-world continual learning. At inference, Modality-aware Adaptive Scoring (MoAS) adaptively adjusts modality contributions using sample-wise reliability and refines novelty scoring with deviation and disagreement penalties. During training, Modality-aware Representation Stabilization Training (MoRST) preserves the discriminative capacity of each modality across tasks through modality-specific heads and modality-wise logit distillation. Experiments on a public multimodal egocentric benchmark show that MAND consistently improves novel activity detection and known-class accuracy while substantially reducing FPR95, indicating more reliable open-world recognition. The source code is available at \href{https://github.com/HyeJeongIm/MAND}{github.com/HyeJeongIm/MAND}.

活动识别多模态持续学习开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。