arXiv:2504.20630eess.AScs.MM2025-04中稿 · ACM Multimedia 202…被引 14

首个通过多模态提示生成沉浸式空间戏剧的模型,支持动态音效与情感表达。

ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting

  • 用对比学习融合多模态信息,捕捉移动说话人带来的多普勒效应。
  • 基于流的Mamba-Transformer架构,结合专家选择机制提升戏剧表现力。
  • 构建首个空间戏剧数据集MRSDrama,适合虚拟现实与增强现实应用。

多模态沉浸式空间戏剧生成旨在基于多模态提示创建连续多说话人双耳语音,并带有戏剧性语调,潜在应用于增强现实(AR)、虚拟现实(VR)等场景。该任务需同时建模空间信息与戏剧性语调,但数据采集成本高。据我们所知,这是首次尝试解决此问题的工作。我们构建了首个多模态录制的空间戏剧数据集MRSDrama,包含双耳戏剧音频、剧本、视频、几何姿态和文本提示。随后提出ISDrama,首个通过多模态提示生成沉浸式空间戏剧的模型。ISDrama包含:1)多模态姿态编码器,基于对比学习,考虑移动说话人引起的多普勒效应,从多模态提示中提取统一姿态信息;2)沉浸式戏剧Transformer,一种基于流的Mamba-Transformer模型,通过引入Drama-MOE选择合适专家以增强语调与姿态控制,并设计上下文一致的无分类器引导策略,实现完整戏剧连贯生成。实验结果表明,ISDrama在客观与主观指标上均优于基线模型。演示视频见https://aaronz345.github.io/ISDramaDemo。数据集与评估代码已公开于https://huggingface.co/datasets/AaronZ345/MRSDrama 和 https://github.com/AaronZ345/ISDrama。

原文摘要 · Abstract (English)

Multimodal immersive spatial drama generation focuses on creating continuous multi-speaker binaural speech with dramatic prosody based on multimodal prompts, with potential applications in AR, VR, and others. This task requires simultaneous modeling of spatial information and dramatic prosody based on multimodal inputs, with high data collection costs. To the best of our knowledge, our work is the first attempt to address these challenges. We construct MRSDrama, the first multimodal recorded spatial drama dataset, containing binaural drama audios, scripts, videos, geometric poses, and textual prompts. Then, we propose ISDrama, the first immersive spatial drama generation model through multimodal prompting. ISDrama comprises these primary components: 1) Multimodal Pose Encoder, based on contrastive learning, considering the Doppler effect caused by moving speakers to extract unified pose information from multimodal prompts. 2) Immersive Drama Transformer, a flow-based mamba-transformer model that generates high-quality drama, incorporating Drama-MOE to select proper experts for enhanced prosody and pose control. We also design a context-consistent classifier-free guidance strategy to coherently generate complete drama. Experimental results show that ISDrama outperforms baseline models on objective and subjective metrics. The demos are available at https://aaronz345.github.io/ISDramaDemo. We provide the dataset and the evaluation code at https://huggingface.co/datasets/AaronZ345/MRSDrama and https://github.com/AaronZ345/ISDrama.

空间音频戏剧生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。