arXiv:2504.03948cs.CV2025-04ICCV被引 3

用概率跳跃扩散方法高效识别未知动作,提升视觉-语言模型在开放世界中的表现。

ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition

  • 基于概率跳跃扩散构建语义连贯搜索空间,融合常识先验与视觉语言模型优化
  • 在多数据集上达最优性能,且适应从封闭到完全开放的复杂度变化
  • 适合研究开放世界动作识别与智能体行为理解的学者参考

开放世界第一人称动作识别因环境无约束而面临根本挑战,要求模型从庞大且部分可观测的搜索空间中推断未见动作。我们提出ProbRes,一种基于跳跃扩散的概率残差搜索框架,通过平衡先验引导探索与似然驱动利用,高效导航该空间。方法结合结构化常识先验构建语义一致的搜索空间,利用视觉语言模型(VLMs)自适应精炼预测,并采用随机搜索机制定位高似然动作标签,同时避免穷举。我们在多个开放程度等级(L0-L3)下系统评估了ProbRes,证明其对搜索空间复杂度增强的适应性。不仅在基准数据集(GTEA Gaze、GTEA Gaze+、EPIC-Kitchens 和 Charades-Ego)上取得领先性能,还建立开放世界识别的清晰分类体系,明确当前挑战与方法演进路径。结果强调结构化搜索策略的重要性,为可扩展、高效的开放世界动作识别铺平道路。

原文摘要 · Abstract (English)

Open-world egocentric activity recognition poses a fundamental challenge due to its unconstrained nature, requiring models to infer unseen activities from an expansive, partially observed search space. We introduce ProbRes, a Probabilistic Residual search framework based on jump-diffusion that efficiently navigates this space by balancing prior-guided exploration with likelihood-driven exploitation. Our approach integrates structured commonsense priors to construct a semantically coherent search space, adaptively refines predictions using Vision-Language Models (VLMs) and employs a stochastic search mechanism to locate high-likelihood activity labels while minimizing exhaustive enumeration efficiently. We systematically evaluate ProbRes across multiple openness levels (L0-L3), demonstrating its adaptability to increasing search space complexity. In addition to achieving state-of-the-art performance on benchmark datasets (GTEA Gaze, GTEA Gaze+, EPIC-Kitchens, and Charades-Ego), we establish a clear taxonomy for open-world recognition, delineating the challenges and methodological advancements necessary for egocentric activity understanding. Our results highlight the importance of structured search strategies, paving the way for scalable and efficient open-world activity recognition.

动作识别开放世界视觉语言模型概率搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。