用语音+上下文+大模型让辅助沟通更自然有个性
Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication
- 融合文本、语音和情感上下文,用大模型生成个性化回应
- 用户表达力显著提升,路径专家认可其自然度与相关性
- 适合需要真实表达的残障人士及语言治疗师参考
本文提出Speak Ease:一种增强型与替代性沟通(AAC)系统,通过整合文本、语音及上下文线索(对话对象与情感语调),结合大语言模型(LLMs),支持用户更个性化、自然且富有表现力的交流。系统融合自动语音识别(ASR)、基于上下文感知的LLM输出与个性化文本转语音技术,实现多模态输入驱动的智能响应。通过与言语语言病理学家(SLPs)开展探索性可行性研究与焦点小组评估,验证了该系统在提升用户表达力方面的潜力。结果揭示了AAC使用者的核心需求,并证实系统能有效增强沟通的个性化与情境适配性。本工作为利用多模态输入与大模型驱动功能改进AAC系统、支持表达多样性提供了重要洞见。
原文摘要 · Abstract (English)
In this paper, we present Speak Ease: an augmentative and alternative communication (AAC) system to support users' expressivity by integrating multimodal input, including text, voice, and contextual cues (conversational partner and emotional tone), with large language models (LLMs). Speak Ease combines automatic speech recognition (ASR), context-aware LLM-based outputs, and personalized text-to-speech technologies to enable more personalized, natural-sounding, and expressive communication. Through an exploratory feasibility study and focus group evaluation with speech and language pathologists (SLPs), we assessed Speak Ease's potential to enable expressivity in AAC. The findings highlight the priorities and needs of AAC users and the system's ability to enhance user expressivity by supporting more personalized and contextually relevant communication. This work provides insights into the use of multimodal inputs and LLM-driven features to improve AAC systems and support expressivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。