arXiv:2512.20409cs.CVcs.AI2025-12

提出DETACH框架,实现非侵入式视频与环境传感器的精准时空对齐。

DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning

  • 分解时空特征,显式保留局部细节
  • 在线聚类生成传感器空间特征,提升语义对齐精度
  • 两阶段对齐:先空间对应,再加权对比学习处理各类负样本

将佩戴式传感器与第一视角视频对齐在人体动作识别中表现良好,但存在用户不适、隐私问题和扩展性差等实际限制。本文探索以环境传感器与第三人称视频作为无侵入、可扩展的替代方案。以往第一视角-佩戴式方法多采用全局对齐,即把完整序列编码为统一表征,但在第三人称-环境设置下因两大问题失效:(P1) 难以捕捉细微运动等局部细节;(P2) 过度依赖模态不变的时间模式,导致语义上下文不同但时间模式相似的动作发生错配。为此,我们提出DETACH——一种分解式时空对齐框架。该框架通过显式分解保留局部细节,同时利用在线聚类发现的新型传感器空间特征提供语义支撑,实现上下文感知对齐。对分解特征的对齐采用两阶段策略:首先通过互监督建立空间对应关系,再通过时空加权对比损失完成时间对齐,自适应处理易负样本、难负样本与假负样本。在Opportunity++和HWU-USP数据集上的下游任务实验表明,该方法显著优于适配后的第一视角-佩戴式基线模型。

原文摘要 · Abstract (English)

Aligning egocentric video with wearable sensors have shown promise for human action recognition, but face practical limitations in user discomfort, privacy concerns, and scalability. We explore exocentric video with ambient sensors as a non-intrusive, scalable alternative. While prior egocentric-wearable works predominantly adopt Global Alignment by encoding entire sequences into unified representations, this approach fails in exocentric-ambient settings due to two problems: (P1) inability to capture local details such as subtle motions, and (P2) over-reliance on modality-invariant temporal patterns, causing misalignment between actions sharing similar temporal patterns with different spatio-semantic contexts. To resolve these problems, we propose DETACH, a decomposed spatio-temporal framework. This explicit decomposition preserves local details, while our novel sensor-spatial features discovered via online clustering provide semantic grounding for context-aware alignment. To align the decomposed features, our two-stage approach establishes spatial correspondence through mutual supervision, then performs temporal alignment via a spatial-temporal weighted contrastive loss that adaptively handles easy negatives, hard negatives, and false negatives. Comprehensive experiments with downstream tasks on Opportunity++ and HWU-USP datasets demonstrate substantial improvements over adapted egocentric-wearable baselines.

视频对齐多模态动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。