arXiv:2602.01634eess.AScs.AI2026-02被引 4

模仿人类听觉机制,实现多路径语音感知

HuPER: A Human-Inspired Framework for Phonetic Perception

  • 基于声学与语言知识的自适应推理框架
  • 100小时数据训练,5个英语任务达顶尖表现
  • 首次支持多场景下自适应语音感知,适合跨语言研究

我们提出 HuPER,一个受人类启发的语音感知建模框架,将语音感知视为对声学-音位证据与语言知识的自适应推断。仅需100小时训练数据,HuPER在五个英语基准上达到最先进的语音错误率,并在95种未见语言上表现出强劲的零样本迁移能力。HuPER也是首个在多种声学条件下实现自适应、多路径语音感知的框架。所有训练数据、模型和代码均已开源,代码与演示可访问 https://github.com/HuPER29/HuPER。

原文摘要 · Abstract (English)

We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER is also the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions. All training data, models, and code are open-sourced. Code and demo avaliable at https://github.com/HuPER29/HuPER.

语音感知自适应推理多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。