Decoding silent reading from non-invasive EEG
作者: Ingo Marquardt, Anthilia Alchanat, Priyanka Jain
分类: cs.LG, q-bio.NC
发布日期: 2026-08-20
备注: 45 pages (including 12 pages of Supplementary Material), 11 figures (including 2 in Supplementary Material)
💡 一句话要点
通过无创EEG解码静默阅读中的词汇信息
🎯 匹配领域: 支柱三:空间感知与语义 (Perception & Semantics) 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 无创EEG 内心语言解码 静默阅读 对比解码器 脑机接口 认知神经科学 词汇信息提取
📋 核心要点
- 现有的内心语言解码方法面临数据收集困难,无法有效获取脑活动与自发内心独白的配对数据。
- 本文提出将静默阅读作为可扩展的代理任务,通过对比解码器提取EEG中的词汇和语义信息。
- 实验结果表明,解码性能显著高于随机基线,且随着训练数据量的增加,解码效果持续提升。
📝 摘要(中文)
无创解码内心语言面临数据问题:无法收集与个人自发内心独白配对的脑活动数据。现有的代理范式(提示重复和回顾性生成内心语言)获取速度慢、时间锁定差且合规性不可验证。因此,本文将静默阅读视为可扩展的代理任务,探讨对比解码器能从中提取多少词汇和语义信息。研究基于19通道干电极EEG记录了约240,000个单词的呈现,采用卷积EEG编码器和因果变换器,使用CLIP风格的对比目标对短EEG窗口与大语言模型的隐藏状态嵌入进行对齐。解码结果显示,词汇信息可从EEG中可靠恢复,且解码受数据限制而非饱和。
🔬 方法详解
问题定义:本文旨在解决无创解码内心语言中的数据问题,现有方法在收集脑活动与内心独白配对数据方面存在显著困难。
核心思路:将静默阅读视为可扩展的代理任务,利用对比解码器从EEG信号中提取词汇和语义信息,以克服数据收集的限制。
技术框架:研究采用19通道干电极EEG记录,使用卷积EEG编码器和因果变换器,结合CLIP风格的对比目标对短EEG窗口与大语言模型的隐藏状态嵌入进行对齐。
关键创新:本研究的创新在于首次展示了在静默阅读过程中,EEG信号中可恢复开放词汇的信息,且解码性能受数据量限制而非饱和。
关键设计:采用快速串行视觉呈现的方式,随机化排版以部分去相关词汇身份与低级视觉形式,卷积网络后接因果变换器,使用对比损失函数进行训练。
🖼️ 关键图片
📊 实验亮点
实验结果显示,解码性能在词汇分组的前10名检索中显著高于随机基线,且对中频和稀有词的解码能力得到扩展。随着训练数据量的增加,解码效果呈对数线性增长,未见饱和迹象。去除枕叶和后颞电极后,词汇解码能力下降约三分之一,但上下文跟踪能力保持不变。
🎯 应用场景
该研究的潜在应用领域包括脑机接口、认知神经科学和心理学等。通过解码静默阅读中的词汇信息,可以为无创脑信号解码技术的发展提供新的思路,推动人机交互和认知状态监测的进步,具有重要的实际价值和未来影响。
📄 摘要(原文)
Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable. We therefore treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract from it. We report an open-vocabulary analysis of approximately 240,000 word presentations recorded from a single densely-sampled participant across 393 runs (ca. 49 h) of 19-channel dry-electrode EEG. Words from continuous narrative text were presented in rapid serial visual presentation, with typography randomised on every trial to partially decorrelate word identity from low-level visual form. A convolutional EEG encoder, optionally followed by a causal transformer, was trained with a CLIP-style contrastive objective to align short EEG windows with hidden-state embeddings of the presented word taken from a large language model. Decoding, evaluated as word-grouped top-10 retrieval against permutation baselines, was reliably above chance, extended to mid-frequency and rare words, and scaled log-linearly with training-data volume with no sign of saturation. Removing occipital and posterior-temporal electrodes reduced the word-level gain by roughly one third but left context tracking unchanged. Control analyses separate word-level decoding from narrative context tracking and from a non-neural positional prior introduced by the transformer's positional embedding. These results establish that open-vocabulary word-level information is recoverable from EEG during silent reading, and that decoding is data-limited rather than saturated.