用规则约束AI分析流程,提升神经影像研究的可信赖性
Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis

- 在预设规则下运行,强制执行可接受的分析流程
- 工具选择准确率提升70.2个百分点,证据可验证性从4.6%升至22.0%
- 适合需要严谨科学论证的研究者和同行评审场景
AI代理能执行科学研究,但只有在权衡多种可能性、限制结论范围并基于证据时,分析结果才能成为可信主张。现有代理常出现选择性分析、过早宣称成功或优化不完善标准等问题。我们提出Brain Researcher,一个嵌入神经影像研究计算环境的智能研究框架,内置可接受分析规则、必检项与主张范围限制。在基准测试中,其使七种模型的第一选择工具准确率提升70.2个百分点(无时为23.3%,有时达93.6%),可验证证据支持率从4.6%增至22.0%。在合作主导与自演化研究中,多宇宙分析揭示了分析选择敏感性,科学评审将主张分类为接受、有条件接受、修改、阻止、拒绝或推迟。通过将决策与证据和溯源关联,Brain Researcher将方法论判断嵌入工作流,而非事后补充。
原文摘要 · Abstract (English)
AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher's computational environment under rules for admissible analyses, required checks and claim scope. In benchmarks, Brain Researcher increased first-choice tool-selection accuracy across seven models by 70.2 percentage points (23.3% without it versus 93.6% with it) and verifiable grounding from 4.6% to 22.0%. In collaborator-led and self-evolving studies, multiverse analyses exposed analytic-choice sensitivity, and scientific review classified claims as accepted, qualified, revised, blocked, rejected or deferred. By linking decisions to evidence and provenance, Brain Researcher embeds methodological judgment within the workflow, not after it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。