从视频中同时识别多个物体的材质参数,突破了传统方法的局限。
MOSIV: Multi-Object System Identification from Videos
- 用可微分物理模拟器结合几何目标,直接优化每个物体的连续材质参数。
- 在复杂多物体交互的合成数据集上,定位精度和长时间模拟保真度显著提升。
- 适合做多物体物理属性建模、视频理解与仿真任务的研究者参考。
我们提出了一个新挑战:从视频中进行多物体系统辨识,现有方法因聚焦单物体场景或固定材料原型分类而不适用。为此,我们提出MOSIV框架,通过由视频导出的几何目标引导的可微分模拟器,直接优化每个物体的连续材质参数。我们还构建了一个包含丰富接触交互的合成基准测试集以支持评估。在该基准上,MOSIV相比适配基线显著提升了定位准确率和长时序模拟保真度,确立了此任务的强基线。分析表明,物体级细粒度监督和几何对齐目标对复杂多物体场景下的稳定优化至关重要。代码与数据集将公开。
原文摘要 · Abstract (English)
We introduce the challenging problem of multi-object system identification from videos, for which prior methods are ill-suited due to their focus on single-object scenes or discrete material classification with a fixed set of material prototypes. To address this, we propose MOSIV, a new framework that directly optimizes for continuous, per-object material parameters using a differentiable simulator guided by geometric objectives derived from video. We also present a new synthetic benchmark with contact-rich, multi-object interactions to facilitate evaluation. On this benchmark, MOSIV substantially improves grounding accuracy and long-horizon simulation fidelity over adapted baselines, establishing it as a strong baseline for this new task. Our analysis shows that object-level fine-grained supervision and geometry-aligned objectives are critical for stable optimization in these complex, multi-object settings. The source code and dataset will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。