arXiv:2605.01809cs.SDcs.AI2026-05被引 2

新基准TMD-Bench评估音乐舞蹈生成的节奏对齐能力

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

论文配图:TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation
图 1 · 摘自论文原文
  • 构建多层级评估体系,涵盖音舞同步、指令遵循和单模质量
  • 发现主流模型虽画质佳,但节奏对齐仍不一致,存在提升空间
  • 提出新基线RhyJAM,在节奏对齐上表现优异,适合研究音舞协同

统一音视频生成在虚拟制作与互动媒体中日益重要,但音乐舞蹈协同生成更具挑战性:需在细粒度时间分辨率下让音乐节奏、乐句与重音驱动动作。现有评估方法依赖单模指标或通用音视频一致性评分,无法捕捉这种节奏耦合。本文提出TMD-Bench,一个文本驱动的音乐舞蹈协同生成评估基准,从单模生成质量、指令遵循度、跨模态节奏对齐三方面评估系统。该基准结合可计算物理指标与感知多模态判断,基于精心对齐节奏的音乐舞蹈数据集,并引入细粒度音乐描述器(Music Captioner)以结构化表达音乐语义。实验表明:(i) 当前商业音视频模型如Veo 3和Sora 2虽生成高质量音乐与视频,但节奏对齐优化不足;(ii) 新提出的基线模型RhyJAM在节奏对齐数据上训练,实现了媲美的节拍级同步性能,同时保持良好的单模保真度。这为构建下一代显式优化节奏与运动一致性的音乐舞蹈生成模型提供了可能。

原文摘要 · Abstract (English)

Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-dance co-generation, the task becomes substantially harder: musical rhythm, phrasing, and accents must drive choreographic motion at fine temporal resolution, and such rhythmic coupling is not captured by unimodal metrics or generic audiovisual consistency scores used in current evaluation practice. We introduce TMD-Bench, a benchmark for text-driven music-dance co-generation that assesses systems across unimodal generation quality, instruction adherence, and cross-modal rhythmic alignment. The benchmark integrates computable physical metrics with perceptual multimodal judgments, and is supported by a curated rhythm-aligned music-dance dataset and a fine-grained Music Captioner for structured music semantics. TMD-Bench further reveals that (i) modern commercial audio-visual models, such as Veo 3 and Sora 2, produce high-quality music and video, while rhythmic coupling remains less consistently optimized and leaves room for improvement, and (ii) our unified baseline RhyJAM trained on rhythm-aligned data achieves competitive beat-level synchronization while maintaining competitive unimodal fidelity. This presents prospects for building next-generation music-dance models that explicitly optimize rhythmic and kinetic coherence.

音乐舞蹈生成节奏对齐多模态评估生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。