arXiv:2609.02887cs.LGq-bio.NC2026-09被引 1

提出统一评估语音脑机接口通信能力的新指标,解决不同系统难比较的问题。

A Common Measure of Communication for Speech Brain-Computer Interfaces

论文配图:A Common Measure of Communication for Speech Brain-Computer Interfaces
图 1 · 摘自论文原文
  • 基于信息论定义开放词汇互信息(OVMI),衡量系统传达用户意图信息的能力
  • 发现传统准确率等指标可能虚高,而OVMI能真实反映通信效率
  • 可指导词汇设计,使系统在三类语音场景中准确率提升最高达16.3%

语音脑机接口(speech BCI)将神经活动转化为语言,为瘫痪患者恢复言语提供可能,并推动自然人机交互发展。然而,该领域缺乏通用的进展度量标准,因各系统使用不同数据集、记录方法、语音类型和词汇表,导致报告性能难以直接比较。核心挑战在于两个未解问题:(i) 语音BCI应支持用户表达何种词分布?(ii) 系统能从该分布中传递多少信息?本文提出开放词汇互信息(OVMI),一种信息论量度,用于评估解码器相对于用户可能沟通词语参考分布的信息传递量。此方法使不同词汇条件下的系统能力可在同一通信尺度上进行评估。我们证明,常规报告的准确率、词错误率(WER)等仅在系统支持词汇上计算的指标,可能夸大系统实际通信能力。通过使用OVMI比较现有系统,揭示了支持语言范围与解码准确性之间的权衡关系,且该权衡依赖于预期沟通内容。进一步表明,优化词汇选择以最大化OVMI,可在三个语音领域实现最高达16.3%的相对准确率提升。因此,OVMI为语音BCI领域提供了可比性基础,助力系统对比、词汇优化与研究进展度量。

原文摘要 · Abstract (English)

Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user's intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user's language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.

脑机接口语音生成信息论评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。