让3D场景能‘听’到手触物的声音,生成效果逼真。
Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes
- 用真实手部动作-声音配对数据训练模型,从手姿轨迹预测声音。
- 生成声音能准确反映材质特性和动作类型,真人难分辨真伪。
- 适合做交互式虚拟现实、数字孪生的音效生成。
我们研究如何让3D场景重建具备交互性,核心问题是:能否预测人手与场景物理交互时产生的声音?首先,我们录制了人手在3D场景中操作物体的视频,获取动作-声音配对数据。随后,利用这些数据训练一个校正流模型,将3D手部轨迹映射为对应音频。测试时,用户可输入任意手部姿态序列作为查询,模型即可估计其对应的声响。实验表明,生成的声音能准确传达材质属性和动作类型,且多数情况下与真实声音对人类观察者而言难以区分。项目页面:https://www.yimingdou.com/hearing_hands/
原文摘要 · Abstract (English)
We study the problem of making 3D scene reconstructions interactive by asking the following question: can we predict the sounds of human hands physically interacting with a scene? First, we record a video of a human manipulating objects within a 3D scene using their hands. We then use these action-sound pairs to train a rectified flow model to map 3D hand trajectories to their corresponding audio. At test time, a user can query the model for other actions, parameterized as sequences of hand poses, to estimate their corresponding sounds. In our experiments, we find that our generated sounds accurately convey material properties and actions, and that they are often indistinguishable to human observers from real sounds. Project page: https://www.yimingdou.com/hearing_hands/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。