用3D生成模型和扩散流自动学习人物交互的接触、朝向与空间占据关系。
H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows
- 通过点云扩散过程学习密集3D流表示,捕捉复杂交互模式。
- 在真实物体上泛化表现优于依赖人工标注的方法。
- 无需标注即可建模接触、朝向、空间占用三类3D交互属性,适合机器人感知研究者。
理解人类如何与周围环境互动,尤其是物体交互与功能感知,是计算机视觉、机器人学和人工智能中的关键挑战。现有方法多依赖耗时费力的手工标注数据集来捕捉真实或模拟的人物交互(HOI)任务,且多数3D功能理解方法仅限于接触分析,忽略朝向(如人倾向于面对电视)和空间占据(如人更可能位于微波炉前而非后)等重要方面。为此,我们提出H2OFlow,一种仅使用3D生成模型合成数据的新框架,全面学习包含接触、朝向与空间占据的3D HOI功能。H2OFlow采用基于点云的密集3D流表示,通过密集扩散过程学习,无需人工标注即可发现丰富的3D功能。大量定量与定性评估表明,H2OFlow在真实物体上具有良好的泛化能力,优于依赖人工标注或网格表示的现有方法。
原文摘要 · Abstract (English)
Understanding how humans interact with the surrounding environment, and specifically reasoning about object interactions and affordances, is a critical challenge in computer vision, robotics, and AI. Current approaches often depend on labor-intensive, hand-labeled datasets capturing real-world or simulated human-object interaction (HOI) tasks, which are costly and time-consuming to produce. Furthermore, most existing methods for 3D affordance understanding are limited to contact-based analysis, neglecting other essential aspects of human-object interactions, such as orientation (\eg, humans might have a preferential orientation with respect certain objects, such as a TV) and spatial occupancy (\eg, humans are more likely to occupy certain regions around an object, like the front of a microwave rather than its back). To address these limitations, we introduce \emph{H2OFlow}, a novel framework that comprehensively learns 3D HOI affordances -- encompassing contact, orientation, and spatial occupancy -- using only synthetic data generated from 3D generative models. H2OFlow employs a dense 3D-flow-based representation, learned through a dense diffusion process operating on point clouds. This learned flow enables the discovery of rich 3D affordances without the need for human annotations. Through extensive quantitative and qualitative evaluations, we demonstrate that H2OFlow generalizes effectively to real-world objects and surpasses prior methods that rely on manual annotations or mesh-based representations in modeling 3D affordance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。