让语音叙事自动纠错并支持人机互动,提升长篇音频故事的连贯性与表现力。
AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling

- 构建闭环自反思框架,通过角色画像与情绪指令匹配语音表现。
- 实现缺陷音频的自动检测与重生成,提升整体音质与情感表达。
- 支持自然语言反馈交互,适合内容创作者与个性化叙事场景。
尽管文本与视觉生成技术取得进展,创作连贯的长篇音频叙事仍具挑战。现有框架常存在角色设定与语音表现不匹配、缺乏自我修正机制、人机交互有限等问题。为此,我们提出AuDirector,一种自反思闭环多智能体框架。具体包括:身份感知预生产机制,将叙事文本转化为角色档案与逐句情绪指令,用于检索合适语音并指导有表现力的语音合成,促进上下文对齐的语音适配;协同合成与修正模块引入闭环自修正机制,系统性审计并重生成有缺陷的音频组件;人类引导交互优化模块则通过解析自然语言反馈,实现对脚本的交互式精炼。实验表明,与最先进基线相比,AuDirector在结构连贯性、情感表现力和声学保真度方面均取得更优性能。音频样例可访问 https://anonymous-itsh.github.io/。
原文摘要 · Abstract (English)
Despite advances in text and visual generation, creating coherent long-form audio narratives remains challenging. Existing frameworks often exhibit limitations such as mismatched character settings with voice performance, insufficient self-correction mechanisms, and limited human interactivity. To address these challenges, we propose AuDirector, a self-reflective closed-loop multi-agent framework. Specifically, it involves an Identity-Aware Pre-production mechanism that transforms narrative texts into character profiles and utterance-level emotional instructions to retrieve suitable voice candidates and guide expressive speech synthesis, thereby promoting context-aligned voice adaptation. To enhance quality, a Collaborative Synthesis and Correction module introduces a closed-loop self-correction mechanism to systematically audit and regenerate defective audio components. Furthermore, a Human-Guided Interactive Refinement module facilitates user control by interpreting natural language feedback to interactively refine the underlying scripts. Experiments demonstrate that AuDirector achieves superior performance compared to state-of-the-art baselines in structural coherence, emotional expressiveness, and acoustic fidelity. Audio samples can be found at https://anonymous-itsh.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。