arXiv:2411.13607cs.CV2024-11中稿 · WACV 2025 in Round…被引 1

用音视频联合建模,精准捕捉小提琴演奏的4维动作

VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference

  • 分层融合音视频特征,通过贝叶斯更新提升姿态估计精度
  • 在自建大规模校准数据集上,显著优于现有单目方法
  • 适合音乐动作分析、人机交互与虚拟演奏等场景

演奏者通过精细的身体控制产生音乐,其动作有时过于细微,难以被人眼捕捉。为分析演奏动作与音乐输出的关系,需精确估计4D人体姿态(三维空间随时间变化)。然而,当前最先进的单目视觉姿态估计方法因遮挡、视角限制、人-物交互等问题,难以准确捕捉快速细微的动作,如颤音效果。本文提出VioPose:一种新型多模态网络,通过音视频的层级推理实现动态姿态估计。高层特征向低层特征传递并融入贝叶斯更新机制,显著提升姿态序列准确性。作为本工作的一部分,我们构建了目前最大且最多样化的校准小提琴演奏数据集,包含视频、音频与3D动作捕捉姿态。代码与数据集可在项目页面获取。

原文摘要 · Abstract (English)

Musicians delicately control their bodies to generate music. Sometimes, their motions are too subtle to be captured by the human eye. To analyze how they move to produce the music, we need to estimate precise 4D human pose (3D pose over time). However, current state-of-the-art (SoTA) visual pose estimation algorithms struggle to produce accurate monocular 4D poses because of occlusions, partial views, and human-object interactions. They are limited by the viewing angle, pixel density, and sampling rate of the cameras and fail to estimate fast and subtle movements, such as in the musical effect of vibrato. We leverage the direct causal relationship between the music produced and the human motions creating them to address these challenges. We propose VioPose: a novel multimodal network that hierarchically estimates dynamics. High-level features are cascaded to low-level features and integrated into Bayesian updates. Our architecture is shown to produce accurate pose sequences, facilitating precise motion analysis, and outperforms SoTA. As part of this work, we collected the largest and the most diverse calibrated violin-playing dataset, including video, sound, and 3D motion capture poses. Code and dataset can be found in our project page \url{https://sj-yoo.info/viopose/}.

姿态估计音视频融合小提琴演奏4D动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。