arXiv:2410.19199cs.SIcs.CL2024-10被引 1

让社交平台更易用:自动识别文本情绪并生成自然语音

Making Social Platforms Accessible: Emotion-Aware Speech Generation with Integrated Text Analysis

  • 端到端系统从文本推断情绪,融合NLP与语音合成技术
  • 支持实时生成,语音自然且情感表达丰富
  • 适合视障、低读写能力用户提升社交可及性

近期研究指出,盲人或视力障碍者、识字能力较弱的人在使用社交媒体时仍面临可及性挑战,尽管已有单调文本转语音(TTS)屏幕阅读器和表情符号音频解说等辅助技术。传统情感语音生成依赖人工输入预期情绪,存在数据简化导致信息丢失、时长不准等问题,难以实现富有表现力的情感呈现。现实中,同一句子因说话人情绪或口音不同,发音时长可变,即“一文多音”问题。为此,我们提出一种端到端上下文感知的文本转语音(TTS)系统,能从文本输入中推断情感,并生成聚焦于情感与说话人特征的自然、有表现力的语音,整合先进自然语言处理(NLP)与语音合成技术,支持实时应用。系统在推理速度上优于当前主流TTS模型,具备实际部署潜力。

原文摘要 · Abstract (English)

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen readers and audio narration of visual elements such as emojis. Emotional speech generation traditionally relies on human input of the expected emotion together with the text to synthesise, with additional challenges around data simplification (causing information loss) and duration inaccuracy, leading to lack of expressive emotional rendering. In real-life communications, the duration of phonemes can vary since the same sentence might be spoken in a variety of ways depending on the speakers' emotional states or accents (referred to as the one-to-many problem of text to speech generation). As a result, an advanced voice synthesis system is required to account for this unpredictability. We propose an end-to-end context-aware Text-to-Speech (TTS) synthesis system that derives the conveyed emotion from text input and synthesises audio that focuses on emotions and speaker features for natural and expressive speech, integrating advanced natural language processing (NLP) and speech synthesis techniques for real-time applications. Our system also showcases competitive inference time performance when benchmarked against the state-of-the-art TTS models, making it suitable for real-time accessibility applications.

语音合成情感识别无障碍设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。