提出首个多人场景下文本指代动作分割方法,解决目标人物定位难题
HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios
- 设计跨输入门控xLSTM增强整体-局部长程推理能力
- 引入傅里叶条件实现细粒度动作生成控制,提升分割精度
- 构建133部电影的大型数据集RHAS133,支持多角色动作识别
动作分割是高层视频理解的核心挑战,旨在将未剪辑视频划分为片段并为每个片段分配预定义的动作标签。现有方法主要针对单人固定动作序列,忽视多人场景。本文首次提出基于文本指代的多人动作分割任务,通过文本描述指定目标人物。我们构建首个该任务的数据集RHAS133,包含133部电影、33小时视频与137种细粒度动作标注。基准测试显示,现有方法在该数据集上性能有限,且难以聚合目标人物视觉线索。为此,我们提出霍帕扩散模型(HopaDIFF),采用新型交叉输入门控xLSTM增强整体-局部长程推理,引入傅里叶条件实现更精细的动作生成控制。HopaDIFF在多种评估设置下均达到当前最优表现。代码与数据集已开源。
原文摘要 · Abstract (English)
Action segmentation is a core challenge in high-level video understanding, aiming to partition untrimmed videos into segments and assign each a label from a predefined action set. Existing methods primarily address single-person activities with fixed action sequences, overlooking multi-person scenarios. In this work, we pioneer textual reference-guided human action segmentation in multi-person settings, where a textual description specifies the target person for segmentation. We introduce the first dataset for Referring Human Action Segmentation, i.e., RHAS133, built from 133 movies and annotated with 137 fine-grained actions with 33h video data, together with textual descriptions for this new task. Benchmarking existing action segmentation methods on RHAS133 using VLM-based feature extractors reveals limited performance and poor aggregation of visual cues for the target person. To address this, we propose a holistic-partial aware Fourier-conditioned diffusion framework, i.e., HopaDIFF, leveraging a novel cross-input gate attentional xLSTM to enhance holistic-partial long-range reasoning and a novel Fourier condition to introduce more fine-grained control to improve the action segmentation generation. HopaDIFF achieves state-of-the-art results on RHAS133 in diverse evaluation settings. The dataset and code are available at https://github.com/KPeng9510/HopaDIFF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。