用脑活动数据构建可操控大模型行为的解释性坐标系。
Brain-Grounded Axes for Reading and Steering LLM States
- 用脑电数据建模词级神经响应,提取潜在语义轴。
- 无微调下实现对模型状态的稳定控制,频率相关轴效果显著。
- 适用于多模型、跨层分析,为模型可解释性提供新工具。
大型语言模型(LLM)的可解释性方法通常依赖文本监督,缺乏外部依据。本文提出将人类脑活动作为坐标系,而非训练信号,用于读取和调控模型状态。基于SMN4Lang MEG数据集,构建词级相位锁定值(PLV)模式图谱,通过独立成分分析(ICA)提取潜在轴。使用独立词表与基于命名实体识别的标签验证轴的有效性,且以词频和词性作为合理性检查。训练轻量级适配器将模型隐藏状态映射至这些脑源轴,无需微调模型。沿脑源方向调控可在中等规模的TinyLlama模型中生成稳健的词汇频率轴(在匹配困惑度的控制下仍有效),且脑源轴相比文本轴表现出更大词频变化但更低困惑度。第13轴呈现一致的函数/内容调控,在TinyLlama、Qwen2-0.5B和GPT-2中均有效,且经词频匹配的文本对照验证。层4效应虽大但不一致,视为次要。当图谱重建时去除GPT嵌入特征或改用word2vec嵌入,轴间相关性仍保持在|r|=0.64–0.95之间,降低循环论证风险。探索性fMRI锚定显示嵌入变化与词频可能存在对齐,但结果受血流动力学建模假设影响,仅作群体水平证据。结果表明,神经生理学基底轴可为模型行为提供可解释且可控的新接口。
原文摘要 · Abstract (English)
Interpretability methods for large language models (LLMs) typically derive directions from textual supervision, which can lack external grounding. We propose using human brain activity not as a training signal but as a coordinate system for reading and steering LLM states. Using the SMN4Lang MEG dataset, we construct a word-level brain atlas of phase-locking value (PLV) patterns and extract latent axes via ICA. We validate axes with independent lexica and NER-based labels (POS/log-frequency used as sanity checks), then train lightweight adapters that map LLM hidden states to these brain axes without fine-tuning the LLM. Steering along the resulting brain-derived directions yields a robust lexical (frequency-linked) axis in a mid TinyLlama layer, surviving perplexity-matched controls, and a brain-vs-text probe comparison shows larger log-frequency shifts (relative to the text probe) with lower perplexity for the brain axis. A function/content axis (axis 13) shows consistent steering in TinyLlama, Qwen2-0.5B, and GPT-2, with PPL-matched text-level corroboration. Layer-4 effects in TinyLlama are large but inconsistent, so we treat them as secondary (Appendix). Axis structure is stable when the atlas is rebuilt without GPT embedding-change features or with word2vec embeddings (|r|=0.64-0.95 across matched axes), reducing circularity concerns. Exploratory fMRI anchoring suggests potential alignment for embedding change and log frequency, but effects are sensitive to hemodynamic modeling assumptions and are treated as population-level evidence only. These results support a new interface: neurophysiology-grounded axes provide interpretable and controllable handles for LLM behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。