arXiv:2507.08137cs.CVcs.AI2025-07被引 1

解决单目视频中人物与物体交互时的遮挡与时间不一致问题。

Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction

  • 通过时序一致性增强,动态补全被遮挡的3D结构。
  • 在复杂遮挡场景下,重建精度显著优于现有方法。
  • 无需预设模板,适用于多样动态交互场景。

我们提出一种新框架,从单目视频中重建动态人-物交互,克服遮挡和时间不一致带来的挑战。传统3D重建方法通常假设物体静止或主体完全可见,当这些假设不成立(尤其在相互遮挡场景)时性能下降。为此,本框架利用无模态补全(amodal completion)推断部分遮挡区域的完整结构。不同于仅处理单帧的传统方法,我们的方法引入时序上下文,通过跨视频序列的连贯性逐步优化并稳定重建结果。该无模板策略可自适应不同条件,无需依赖预定义模型,在动态场景中显著提升细节恢复能力。我们在具有挑战性的单目视频上使用3D Gaussian Splatting验证方法,结果表明其在处理遮挡和保持时间稳定性方面均优于现有技术。

原文摘要 · Abstract (English)

We introduce a novel framework for reconstructing dynamic human-object interactions from monocular video that overcomes challenges associated with occlusions and temporal inconsistencies. Traditional 3D reconstruction methods typically assume static objects or full visibility of dynamic subjects, leading to degraded performance when these assumptions are violated-particularly in scenarios where mutual occlusions occur. To address this, our framework leverages amodal completion to infer the complete structure of partially obscured regions. Unlike conventional approaches that operate on individual frames, our method integrates temporal context, enforcing coherence across video sequences to incrementally refine and stabilize reconstructions. This template-free strategy adapts to varying conditions without relying on predefined models, significantly enhancing the recovery of intricate details in dynamic scenes. We validate our approach using 3D Gaussian Splatting on challenging monocular videos, demonstrating superior precision in handling occlusions and maintaining temporal stability compared to existing techniques.

3D重建遮挡处理时序一致性人物交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。