用对象化编辑自动把书变成音频剧,改角色声音一键同步
Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books

- 将人物、场景等故事元素作为可编辑对象,实现跨资产联动修改
- 用户研究显示任务负荷显著降低,听众评估质量接近专业工具产出
- 适合音频剧创作者、内容制作者,降低入门门槛
音频剧通过对话、音效和音乐构建沉浸式叙事。创作者常将书籍改编为音频剧,但过程繁琐,需解读文本、编写脚本、生成音频素材并逐轨拼接。由于角色、场景等元素分布在多个相互依赖的素材中,单次修改需手动更新全项目。本文提出Dramarrator,一种基于对象的音频剧创作工具:从书籍中提取故事元素作为可编辑对象,自动生成语音、音效与音乐,并组合成多轨音频剧。对任一对象(如角色声音)的修改可自动传播至所有关联素材。8名专业人士参与的用户研究显示,使用Dramarrator显著降低了任务负担;300名听众参与的评测表明,经创作者优化后的输出质量接近现有专业工具水平;3人探索性研究提示,该方法能降低入门门槛,且可拓展至其他类型音频创作。
原文摘要 · Abstract (English)
Audio dramas weave dialogue, sound effects, and music into immersive stories. Creators often adapt books into audio dramas, but this process remains labor-intensive, requiring them to interpret source material, author scripts, generate audio assets, and assemble them on a timeline. Because story elements like characters and scenes manifest across many interdependent assets, a single change can ripple into manual updates across the entire project. We present Dramarrator, an audio drama authoring tool built around object-based audio editing, where these story elements are represented as editable objects. Dramarrator extracts these objects from a book, generates linked audio assets (speech, sound effects, and music), and composes a multi-track audio drama. Edits to any object (e.g., a character's voice) automatically propagate to all dependent assets. In a user study with professionals (N=8), Dramarrator significantly lowered task load when creating audio dramas. A listener study (N=300) shows that creator-refined output from Dramarrator approaches the quality of productions made with existing professional tools, and an exploratory study (N=3) suggests object-based editing lowers entry barriers and generalizes beyond audio dramas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。