arXiv:2504.10849cs.HCcs.MM2025-04

实时语音识别中按词级调整文字样式,增强语义表达

Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition

  • 基于语音响度动态调节每个词的字号大小
  • 用户反馈显示能更好传达说话人意图
  • 适合听障、语言学习者及自闭症群体使用

富文本字幕对聋哑及听力障碍(DHH)人群、第二语言学习者以及自闭症谱系障碍(ASD)个体的交流至关重要。它们还能在语音转写时保留语气细节,提升剧本、对话或演讲记录的真实感。然而,现有实时字幕系统无法在词级别调整文本属性(如大小写、字号、字体),难以准确传递通过语调或重音表达的说话人意图。例如,'YOU should do this' 通常强调 'You',而 'You should do THIS' 则强调 'This'。本文提出一种实时在词级别改变文本装饰的解决方案。作为原型,我们开发了根据每个词语发音响度动态调整字号的应用。用户反馈表明,该系统有助于更准确传达说话人意图,带来更生动且可访问的字幕体验。

原文摘要 · Abstract (English)

Rich-text captions are essential to help communication for Deaf and hard-of-hearing (DHH) people, second-language learners, and those with autism spectrum disorder (ASD). They also preserve nuances when converting speech to text, enhancing the realism of presentation scripts and conversation or speech logs. However, current real-time captioning systems lack the capability to alter text attributes (ex. capitalization, sizes, and fonts) at the word level, hindering the accurate conveyance of speaker intent that is expressed in the tones or intonations of the speech. For example, ''YOU should do this'' tends to be considered as indicating ''You'' as the focus of the sentence, whereas ''You should do THIS'' tends to be ''This'' as the focus. This paper proposes a solution that changes the text decorations at the word level in real time. As a prototype, we developed an application that adjusts word size based on the loudness of each spoken word. Feedback from users implies that this system helped to convey the speaker's intent, offering a more engaging and accessible captioning experience.

实时字幕语音识别可访问性语义表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。