用脑数据辅助训练模型,能省多少标注数据?
How Much is Brain Data Worth for Machine Learning?

- 构建线性高斯模型,分析脑数据与任务数据的协同效应
- 发现脑数据价值随任务-脑对齐度提升而显著增加
- 在预算有限时,为特定场景提供是否采集脑数据的决策依据
如果人能完成某项任务,测量其大脑活动能否帮助更高效地训练模型?最近的神经AI研究显示,结合脑电数据与任务标签可轻微提升模型性能与鲁棒性。但何时有收益、收益多大尚不明确。本文建立数学模型,基于一个简单可解析的线性高斯框架,推导出联合使用脑数据与任务标签的多模态估计器的性能缩放规律。我们得出脑样本与任务样本之间的相对价值与交换率,量化了脑数据在不同任务-脑对齐度、神经噪声、任务噪声、潜空间维度及脑数据量下的等价任务样本数。还分析了测试分布偏移下,脑正则化学习通过学习不变性带来的鲁棒性增益。最后,在固定采集预算下,识别出脑数据值得收集的条件区间。结果为理解脑数据在机器学习中的潜在价值提供了理论基础。
原文摘要 · Abstract (English)
If a person can solve a task, can measuring their brain make it easier to train a model to solve that task too? Recent NeuroAI work suggests that supplementing task training with neural recordings can modestly improve model performance and robustness. However, it is unclear when there should be a benefit from using neural data and how much benefit to expect. We formulate this question mathematically, and begin to address it theoretically using a simple, analytically tractable linear gaussian model of task targets and neural recordings. For a multimodal estimator trained on both brain data and task labels, we derive scaling laws for how performance scales with the numbers of brain and task samples. From these laws we derive relative value and exchange rates between brain samples and task samples, quantifying how much extra task samples neural data is worth as a function of task-brain alignment, neural and task noise, latent dimension, and brain data sample size. We also analyze test distribution shift, to identify conditions where brain-regularized learning can produce substantial robustness gains through learned invariances. Finally, under a fixed collection budget, we characterize the regimes in which brain data is worth collecting. Our results provide a foundation for understanding how valuable brain data could be for improving machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。