用扩散模型从脑电生成文本,提升准确率与稳定性。
DELTA: Language Diffusion-based EEG-to-Text Architecture
- 将脑电信号分层离散化,降低噪声和个体差异影响。
- 在ZuCo数据集上,语义对齐提升5.37点,BLEU-1达21.9。
- 适合小样本脑电-文本生成,推动多模态模型发展。
脑电到文本生成因高维噪声、个体差异及自回归解码误差累积而困难。我们提出DELTA,结合残差向量量化(RVQ)EEG分词器与掩码语言扩散模型(LLaDA)。RVQ将连续脑电信号离散为多层令牌,减少噪声与个体差异;LLaDA通过非序列去噪重建句子。在ZuCo数据集上,DELTA相较自回归基线在语义对齐上最高提升5.37点,达到词级条件下BLEU-1 21.9与ROUGE-1 F 17.2。该方法实现小样本脑电-文本数据集的可靠文本生成,为可扩展的多模态脑电-语言模型提供方向。
原文摘要 · Abstract (English)
Electroencephalogram (EEG)-to-text remains challenging due to high-dimensional noise, subject variability, and error accumulation in autoregressive decoding. We introduce DELTA, which pairs a Residual Vector Quantization (RVQ) EEG tokenizer with a masked language diffusion model (LLaDA). RVQ discretizes continuous EEG into multi-layer tokens to reduce noise and individual differences, while LLaDA reconstructs sentences via non-sequential denoising. On ZuCo, DELTA improves semantic alignment by up to 5.37 points over autoregressive baselines, achieving BLEU-1 21.9 and ROUGE-1 F 17.2 under word-level conditions. These results enable reliable text generation from small EEG-text datasets and point toward scalable multimodal EEG-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。