arXiv:2512.10778cs.SDcs.MM2025-12被引 3

用手机打造可编辑的音画数字孪生,让空间还原更真实。

Building Audio-Visual Digital Twins with Smartphones

  • 用手机采集房间声学响应,结合视觉信息重建声场。
  • 通过可微渲染恢复表面材质,支持修改后自动更新音画效果。
  • 适合虚拟现实、智能建筑等需真实声景的应用场景。

当前的数字孪生系统几乎完全依赖视觉,忽视了声音这一空间真实感与交互的核心要素。我们提出AV-Twin,首个仅使用普通智能手机即可构建可编辑音画数字孪生的实用系统。AV-Twin结合移动端的混响脉冲响应(RIR)采集与视觉辅助的声场建模,高效重构房间声学特性。进一步通过可微声学渲染恢复各表面材料属性,使用户在修改材质、几何形状或布局时,能自动同步更新音效与视觉内容。该系统为真实世界环境的全可控音画数字孪生提供了可行路径。

原文摘要 · Abstract (English)

Digital twins today are almost entirely visual, overlooking acoustics-a core component of spatial realism and interaction. We introduce AV-Twin, the first practical system that constructs editable audio-visual digital twins using only commodity smartphones. AV-Twin combines mobile RIR capture and a visual-assisted acoustic field model to efficiently reconstruct room acoustics. It further recovers per-surface material properties through differentiable acoustic rendering, enabling users to modify materials, geometry, and layout while automatically updating both audio and visuals. Together, these capabilities establish a practical path toward fully modifiable audio-visual digital twins for real-world environments.

数字孪生音画同步手机建模可微渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。