用文字描述提升单图3D人物与物体重建的准确性与语义一致性。
TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures
- 引入文本描述增强3D重建的语义对齐,支持非接触交互建模。
- 融合外观特征实现全局上下文感知,提升视觉合理性。
- 适用于需理解复杂人物-物体关系的数字内容生成场景。
从单张图像联合重建3D人体与物体是机器人和数字内容创作中的重要研究方向。尽管近期取得进展,现有方法仍存在两大根本局限:一是依赖物理接触信息,难以建模非接触交互(如凝视、指向);二是主要基于局部几何邻近性,忽略人体与物体外观提供的全局上下文信息。为此,我们提出TeHOR框架,核心设计包括:首先,超越接触信息,利用人-物交互的文本描述强制3D重建与文本线索的语义对齐,实现更广泛交互类型的推理;其次,将3D人体与物体的外观特征融入对齐过程,捕捉整体上下文信息,确保重建结果在视觉上合理。实验表明,该框架实现了当前最优性能,生成了准确且语义一致的3D重建。
原文摘要 · Abstract (English)
Joint reconstruction of 3D human and object from a single image is an active research area, with pivotal applications in robotics and digital content creation. Despite recent advances, existing approaches suffer from two fundamental limitations. First, their reconstructions rely heavily on physical contact information, which inherently cannot capture non-contact human-object interactions, such as gazing at or pointing toward an object. Second, the reconstruction process is primarily driven by local geometric proximity, neglecting the human and object appearances that provide global context crucial for understanding holistic interactions. To address these issues, we introduce TeHOR, a framework built upon two core designs. First, beyond contact information, our framework leverages text descriptions of human-object interactions to enforce semantic alignment between the 3D reconstruction and its textual cues, enabling reasoning over a wider spectrum of interactions, including non-contact cases. Second, we incorporate appearance cues of the 3D human and object into the alignment process to capture holistic contextual information, thereby ensuring visually plausible reconstructions. As a result, our framework produces accurate and semantically coherent reconstructions, achieving state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。