arXiv:2508.09428cs.CVcs.AI2025-08AAAI被引 1

同时定位动作与身体接触区域,提升视觉理解精度

What-Meets-Where: Unified Learning of Action and Contact Localization in Images

  • 设计新框架PaIR-Net,联合建模动作与接触位置
  • 在13,979张图像上实现654种动作与17个身体部位的精准识别
  • 适合研究人机交互、场景理解的视觉算法开发者

人们通过肢体与环境建立接触来完成动作。为全面理解多样化视觉场景中的行为,必须同时考虑动作类型(what)和发生位置(where)。现有方法通常未能协同建模动作语义与空间上下文。为此,本文提出一个新视觉任务:同时预测高层动作语义与细粒度身体部位接触区域。提出的PaIR-Net框架包含三个核心组件:接触先验感知模块(CPAM)用于识别相关身体部位,先验引导拼接分割器(PGCS)实现像素级接触分割,交互推理模块(IIM)整合全局交互关系。为支持该任务,我们构建了PaIR数据集,包含13,979张图像,涵盖654种动作、80类物体和17个身体部位。实验表明,PaIR-Net显著优于基线模型,消融实验证实各组件有效性。代码与数据集将在发表后公开。

原文摘要 · Abstract (English)

People control their bodies to establish contact with the environment. To comprehensively understand actions across diverse visual contexts, it is essential to simultaneously consider \textbf{what} action is occurring and \textbf{where} it is happening. Current methodologies, however, often inadequately capture this duality, typically failing to jointly model both action semantics and their spatial contextualization within scenes. To bridge this gap, we introduce a novel vision task that simultaneously predicts high-level action semantics and fine-grained body-part contact regions. Our proposed framework, PaIR-Net, comprises three key components: the Contact Prior Aware Module (CPAM) for identifying contact-relevant body parts, the Prior-Guided Concat Segmenter (PGCS) for pixel-wise contact segmentation, and the Interaction Inference Module (IIM) responsible for integrating global interaction relationships. To facilitate this task, we present PaIR (Part-aware Interaction Representation), a comprehensive dataset containing 13,979 images that encompass 654 actions, 80 object categories, and 17 body parts. Experimental evaluation demonstrates that PaIR-Net significantly outperforms baseline approaches, while ablation studies confirm the efficacy of each architectural component. The code and dataset will be released upon publication.

动作识别接触检测多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。