让漫画角色按情绪说专属台词,一键生成带情感的配音。
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
- 用图像处理识别角色、对话和情绪强度。
- 结合剧情上下文分析对话归属与情绪状态。
- 为每个角色定制声音,按情绪动态调整音色。
本文提出一种端到端的漫画语音生成流程,可基于完整漫画卷册生成角色专属、情绪感知的语音。系统输入为整部漫画,输出与角色对话及情绪状态对齐的语音。图像处理模块完成角色检测、文本识别和情绪强度识别;大语言模型融合视觉信息与剧情演进上下文,实现对话归属与情绪分析;语音合成采用文本转语音模型,为每个角色配置独特声线,并根据情绪动态调整。该工作实现了漫画自动配音,推动互动式沉浸阅读体验的发展。
原文摘要 · Abstract (English)
This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional state. An image processing module performs character detection, text recognition, and emotion intensity recognition. A large language model performs dialogue attribution and emotion analysis by integrating visual information with the evolving plot context. Speech is synthesized through a text-to-speech model with distinct voice profiles tailored to each character and emotion. This work enables automated voiceover generation for comics, offering a step toward interactive and immersive comic reading experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。