脑基础模型训练数据需建立更强伦理规范。
Training Data Governance for Brain Foundation Models
- 基于脑电、fMRI等神经数据构建通用AI模型。
- 神经数据具高敏感性,传统治理难适应大规模复用。
- 提出隐私、知情同意、利益共享等关键防护建议。
脑基础模型将基础模型范式引入神经科学领域。与语言和图像基础模型类似,它们通过大规模数据预训练,可灵活适配下游任务。但不同于文本与图像模型,其训练数据为脑电图(EEG)、功能性磁共振成像(fMRI)等神经数据,这些数据长期在临床与科研的严格监管下收集。本文指出,以神经数据训练基础模型开启了新的规范领域:神经数据因源自人体且历史受严格管控,保护期待远高于文本或图像。然而,基础模型范式却推动大规模再利用、跨场景拼接与开放应用,且参与者范围扩大至商业开发者,当前治理碎片化且模糊。为此,本文首先梳理脑基础模型的技术基础与数据生态,继而结合人工智能伦理、神经伦理与生物伦理,系统归纳隐私、知情同意、偏见、利益分配与治理等维度的问题,针对每个方面提出核心议题与基础性保障措施,助力该领域健康发展。
原文摘要 · Abstract (English)
Brain foundation models bring the foundation model paradigm to the field of neuroscience. Like language and image foundation models, they are general-purpose AI systems pretrained on large-scale datasets that adapt readily to downstream tasks. Unlike text-and-image based models, however, they train on brain data: large-datasets of EEG, fMRI, and other neural data types historically collected within tightly governed clinical and research settings. This paper contends that training foundation models on neural data opens new normative territory. Neural data carry stronger expectations of, and claims to, protection than text or images, given their body-derived nature and historical governance within clinical and research settings. Yet the foundation model paradigm subjects them to practices of large-scale repurposing, cross-context stitching, and open-ended downstream application. Furthermore, these practices are now accessible to a much broader range of actors, including commercial developers, against a backdrop of fragmented and unclear governance. To map this territory, we first describe brain foundation models' technical foundations and training-data ecosystem. We then draw on AI ethics, neuroethics, and bioethics to organize concerns across privacy, consent, bias, benefit sharing, and governance. For each, we propose both agenda-setting questions and baseline safeguards as the field matures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。