用真实音频样本引导脑信号转音频,解决生成音不准的问题
RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction

- 用检索匹配的真实音频初始化生成器,避免先验主导
- 在Brain2Music上识别准确率从0.14提至0.43,FAD降为1.25
- 适合需要高保真音频重建的研究者,尤其关注生成质量
脑信号转音频受限于先验主导:当预训练生成器仅由弱神经信号条件化时,会输出逼真但与刺激不匹配的音频。我们提出RAG-Audio,将fMRI解码为语义音频嵌入,检索匹配的真实音频样本,并用该样本初始化冻结生成器的采样轨迹,同时保留解码嵌入作为条件。在Brain2Music数据集上,10类刺激识别准确率从0.14提升至0.43(接近随机水平0.10),接近检索性能;AudioLDM的弗雷歇音频距离(FAD)从13.49降至1.25。RAG-Audio在识别性能上接近最近邻检索,同时保持生成能力;更高的FAD源于检索直接复现真实音频。自回归负控制实验未见类似提升,说明改进源于轨迹初始化。结果表明,检索引导的初始化可缓解脑信号转音频中的先验主导问题。
原文摘要 · Abstract (English)
Brain-to-audio reconstruction is limited by \emph{prior domination}: when a pretrained generator is conditioned on a weak neural signal, it produces realistic but stimulus-inaccurate audio. We introduce RAG-Audio, which decodes fMRI into a semantic audio embedding, retrieves a matching real-audio exemplar, and initializes the frozen generator's sampling trajectory from that exemplar while retaining the decoded embedding as conditioning. On Brain2Music, RAG-Audio improves 10-way stimulus identification from $0.14$--$0.18$ for direct generation, near the $0.10$ chance level, to $0.40$--$0.43$, comparable to retrieval. It also reduces Fréchet Audio Distance by roughly an order of magnitude, from $13.49$ to $1.25$ for AudioLDM. RAG-Audio approaches nearest-neighbor retrieval in identification while remaining generative; its higher FAD is expected because retrieval directly replays real audio. An autoregressive negative control, which lacks an initializable latent trajectory, shows no comparable gain, attributing the improvement to trajectory initialization. These results suggest that retrieval-guided initialization can mitigate prior domination in brain-to-audio generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。