arXiv:2506.17746cs.CV2025-06被引 1

仅用一张图实现物理交互,手机端实时渲染不卡顿

PhysID: Physics-based Interactive Dynamics from a Single-view Image

  • 用大模型从单图生成3D网格和物理属性
  • 手机端运行,内存低且支持实时交互
  • 适合做AR/VR互动,无需专业建模经验

将静态图像转化为可交互体验仍是计算机视觉中的挑战。现有方法需多视角图像或预录视频输入。本文提出PhysID,通过大型生成模型实现从单张图像生成3D网格与物理属性,显著降低3D建模和属性校准等工程门槛,支持最小人工干预下的规模化应用。系统集成设备端物理引擎,实现用户交互下的物理合理实时渲染。实验评估了多种多模态大语言模型(MLLMs)在零样本任务上的表现及3D重建模型性能,结果表明端到端框架中各模块协同良好,整体有效。该方法推动移动端交互动态发展,支持实时、非确定性交互与用户个性化,且内存消耗高效。

原文摘要 · Abstract (English)

Transforming static images into interactive experiences remains a challenging task in computer vision. Tackling this challenge holds the potential to elevate mobile user experiences, notably through interactive and AR/VR applications. Current approaches aim to achieve this either using pre-recorded video responses or requiring multi-view images as input. In this paper, we present PhysID, that streamlines the creation of physics-based interactive dynamics from a single-view image by leveraging large generative models for 3D mesh generation and physical property prediction. This significantly reduces the expertise required for engineering-intensive tasks like 3D modeling and intrinsic property calibration, enabling the process to be scaled with minimal manual intervention. We integrate an on-device physics-based engine for physically plausible real-time rendering with user interactions. PhysID represents a leap forward in mobile-based interactive dynamics, offering real-time, non-deterministic interactions and user-personalization with efficient on-device memory consumption. Experiments evaluate the zero-shot capabilities of various Multimodal Large Language Models (MLLMs) on diverse tasks and the performance of 3D reconstruction models. These results demonstrate the cohesive functioning of all modules within the end-to-end framework, contributing to its effectiveness.

交互生成单图建模物理引擎移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。