用双几何扩散模型提升动作分割的层次化理解能力
Learning Action Hierarchies via Hybrid Geometric Diffusion
- 融合欧氏与双曲几何,引导分层去噪过程
- 在三个数据集上达到当前最优性能
- 适合需要精细动作分析的研究者
时序动作分割是视频理解中的关键任务,目标是为视频中每一帧分配动作标签。尽管近期方法采用迭代精炼策略,但未能显式利用人类动作的层次特性。本文提出 HybridTAS 框架,将欧氏与双曲几何结合于扩散模型的去噪过程,以挖掘动作的层次结构。双曲几何天然支持嵌入间的树状关系,使动作标签去噪可沿粗到细层级进行:高扩散步受抽象高层动作类别(根节点)影响,低步则由细粒度动作类别(叶节点)细化。在 GTEA、50Salads 与 Breakfast 三个基准数据集上的实验表明,该方法实现最先进性能,验证了双曲引导去噪在时序动作分割中的有效性。
原文摘要 · Abstract (English)
Temporal action segmentation is a critical task in video understanding, where the goal is to assign action labels to each frame in a video. While recent advances leverage iterative refinement-based strategies, they fail to explicitly utilize the hierarchical nature of human actions. In this work, we propose HybridTAS - a novel framework that incorporates a hybrid of Euclidean and hyperbolic geometries into the denoising process of diffusion models to exploit the hierarchical structure of actions. Hyperbolic geometry naturally provides tree-like relationships between embeddings, enabling us to guide the action label denoising process in a coarse-to-fine manner: higher diffusion timesteps are influenced by abstract, high-level action categories (root nodes), while lower timesteps are refined using fine-grained action classes (leaf nodes). Extensive experiments on three benchmark datasets, GTEA, 50Salads, and Breakfast, demonstrate that our method achieves state-of-the-art performance, validating the effectiveness of hyperbolic-guided denoising for the temporal action segmentation task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。