arXiv:2604.26227cs.CV2026-04IJCAI被引 11

利用人-物交互动态调整模型,区分相似动作

HOI-aware Adaptive Network for Weakly-supervised Action Segmentation

论文配图:HOI-aware Adaptive Network for Weakly-supervised Action Segmentation
图 1 · 摘自论文原文
  • 基于视频级人-物交互信息动态调整网络参数
  • 在Breakfast和50Salads数据集上提升分割准确率
  • 适合处理动作相似、依赖上下文的弱监督场景

本文提出一种名为AdaAct的HOI感知自适应网络,用于弱监督动作分割。现有方法通常使用固定网络预测每帧动作,但在区分相似动作(如倒果汁与倒咖啡)时易产生歧义。为此,我们利用时空局部但时间全局的人-物交互(HOI)作为视频级先验知识,通过长时序HOI序列提供关键上下文信息,使网络在测试时能根据具体视频的HOI动态自适应调整。首先设计视频级HOI编码器,提取、筛选并整合视频中最具代表性的HOI;随后提出双分支HyperNetwork,学习自适应时间编码器,可依据不同视频的HOI信息实时调整参数。在Breakfast和50Salads两个常用数据集上的大量实验表明,该方法在多种评估指标下均有效。

原文摘要 · Abstract (English)

In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed network to predict the action of each frame with the neighboring frames. However, this would result in ambiguity when estimating similar actions, such as pouring juice and pouring coffee. To address this, we aim to exploit temporally global but spatially local human-object interactions (HOI) as video-level prior knowledge for action segmentation. The long-term HOI sequence provides crucial contextual information to distinguish ambiguous actions, where our network dynamically adapts to the given HOI sequence at test time. More specifically, we first design a video HOI encoder that extracts, selects, and integrates the most representative HOI throughout the video. Then, we propose a two-branch HyperNetwork to learn an adaptive temporal encoder, which automatically adjusts the parameters based on the HOI information of various videos on the fly. Extensive experiments on two widely-used datasets including Breakfast and 50Salads demonstrate the effectiveness of our method under different evaluation metrics.

动作分割弱监督人-物交互自适应网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。