A Common Measure of Communication for Speech Brain-Computer Interfaces
作者: Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones
分类: cs.LG, q-bio.NC
发布日期: 2026-09-02
备注: Code and OVMI Explorer available from the project page at https://neural-processing-lab.github.io/OVMI/
💡 一句话要点
提出开放词汇互信息以解决语音脑机接口的测量问题
🎯 匹配领域: 支柱三:空间感知与语义 (Perception & Semantics)
关键词: 语音脑机接口 开放词汇互信息 信息论 自然语言处理 人机交互 准确率提升 系统比较
📋 核心要点
- 现有的语音脑机接口系统缺乏统一的进展衡量标准,导致不同系统之间的比较困难。
- 本文提出开放词汇互信息(OVMI),作为一种信息论量度,解决了如何衡量语音BCI系统的信息传递能力。
- 通过OVMI的应用,研究显示在不同语音领域中,系统的准确率可提高至16.3%,为该领域的比较和进步提供了新方法。
📝 摘要(中文)
语音脑机接口(speech BCI)将神经活动转化为语言,为瘫痪患者恢复言语提供了可能,同时也促进了自然人机交互的新形式。然而,该领域缺乏统一的进展衡量标准,因不同系统使用不同的数据集、记录方法、语音类型和词汇,导致其报告的分数难以比较。本文提出开放词汇互信息(OVMI),作为一种信息论量度,衡量解码器相对于用户可能希望交流的词汇参考分布所传达的信息。这一方法使得在不同条件下测量的能力能够在统一的交流尺度上进行评估。研究表明,传统的准确率和词错误率等指标可能会夸大系统能够传达的用户意图言语的程度。通过OVMI比较现有系统,揭示了系统支持的语言量与解码准确性之间的权衡,并展示了最大化OVMI的词汇选择能够在三个语音领域中提高准确率达16.3%。
🔬 方法详解
问题定义:本文解决了语音脑机接口(BCI)领域中缺乏统一进展衡量标准的问题,现有方法在不同数据集和词汇下的比较难以进行,导致无法准确评估系统的性能和进展。
核心思路:提出开放词汇互信息(OVMI),作为一种信息论量度,旨在衡量解码器相对于用户希望交流的词汇分布所传达的信息量,从而实现不同条件下的能力评估。
技术框架:OVMI的计算涉及对用户可能交流的词汇进行参考分布的定义,并通过解码器输出的结果进行信息量的评估。整体流程包括数据收集、词汇选择、信息量计算和系统比较等主要模块。
关键创新:OVMI的提出是本文的核心创新点,它提供了一种新的视角来评估和比较不同的语音BCI系统,克服了传统方法的局限性。与现有方法相比,OVMI能够更准确地反映系统的实际信息传递能力。
关键设计:在设计OVMI时,关键参数包括用户的参考词汇分布、解码器的输出结果以及信息量的计算方法。损失函数的选择和网络结构的设计也对系统的性能有重要影响。具体的技术细节在实验部分进行了详细描述。
🖼️ 关键图片
📊 实验亮点
实验结果表明,通过使用开放词汇互信息(OVMI),现有语音BCI系统的准确率在三个不同的语音领域中提高了最多16.3%。这一显著提升不仅展示了OVMI的有效性,也揭示了系统支持的语言量与解码准确性之间的权衡关系。
🎯 应用场景
该研究的潜在应用领域包括医疗康复、辅助沟通设备以及人机交互界面等。通过提供一种统一的测量标准,OVMI可以帮助研究人员和开发者更好地设计和优化语音BCI系统,推动该领域的进步与创新。未来,OVMI的应用可能会促进更广泛的自然语言处理和人机交互技术的发展。
📄 摘要(原文)
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user's intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user's language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.