arXiv:2605.24047cs.CV2026-05中稿 · CVPR

从视频音频等多模态数据中直接恢复系统物理参数,无需传感器或标注。

EMMA: Extracting Multiple physical parameters from Multimodal Data

论文配图:EMMA: Extracting Multiple physical parameters from Multimodal Data
图 1 · 摘自论文原文
  • 融合视频、音频与图像时间序列,统一建模动态系统
  • 在75个真实视频和100+场景中实现多参数准确恢复
  • 适合无标签、有遮挡或隐藏输入的复杂系统分析

我们提出EMMA,一种基于物理规律的多模态框架,可直接从原始视频、音频和图像的时间序列中恢复系统的所有可识别动力学参数。与以往仅依赖视频的方法不同,EMMA能应对遮挡状态、隐藏控制输入以及已知初始条件或坐标系的假设限制,在统一的连续时间模型中联合推断显式参数、隐式动力分量与校准不变量。它采用液态时间常数(LTC)网络从异构模态中学习潜在动力学,并通过物理约束损失确保与控制微分方程的一致性。统一特征管道实现了视频轨迹、声学特征与图表测量间的精准对齐,使EMMA可在强制、隐式及多变量动力学下估计参数,且无需分割掩码、可微渲染或专用传感器。在超过100种场景中验证,包括五个标准动力学基准(75个Delfy视频)、真实探测车与四旋翼系统(含隐藏输入),以及涵盖生物与混沌系统的仿真-图表案例,结果表明其显著优于现有单模态与方程发现基线。该工作确立了从偶然获取的多模态数据中提取物理一致模型的通用可扩展方案。代码与数据见:https://github.com/ImpactLabASU/EMMA-CVPR2026

原文摘要 · Abstract (English)

We introduce EMMA, a physics-informed multimodal framework that recovers all identifiable dynamical parameters of a system directly from raw video, audio, and image-based time-series observations. Unlike prior video-only approaches that struggle with occluded states, hidden actuation inputs, or assumptions about known initial conditions and coordinate frames, EMMA performs joint inference of explicit parameters, implicit dynamical components, and calibration invariants within a unified continuous-time model. EMMA leverages a Liquid Time-Constant (LTC) network to learn latent dynamics from heterogeneous modalities while a physics-constrained loss enforces consistency with the governing differential equations. A unified feature pipeline enables consistent alignment across video trajectories, acoustic signatures, and chart-derived measurements, allowing EMMA to estimate parameters under forced, implicit, and multivariate dynamics without requiring segmentation masks, differentiable rendering, or specialized sensors. Across 100+ scenarios including five standard dynamical benchmarks (75 Delfys videos), real-world rover and quadrotor systems with hidden inputs, and simulation-chart case studies spanning biological and chaotic systems, EMMA delivers robust multi-parameter recovery and significantly outperforms existing single-modality and equation-discovery baselines. Our results establish EMMA as a general, scalable solution for physics-consistent model extraction from opportunistic multimodal data. Code and data are available at: https://github.com/ImpactLabASU/EMMA-CVPR2026

多模态感知物理建模参数估计视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。