arXiv:2505.04638cs.AIcs.CL2025-05

构建专家参与的AI科研助手评估框架,提升生物医学AI可靠性

Advancing AI Research Assistants with Expert-Involved Learning

  • 用专家标注任务+统一评估协议测试模型能力
  • 顶尖模型摘要流畅但不完整,视觉推理能力弱
  • 优化后可生成可验证的生物学假设,适合医学研究者使用

大语言模型(LLMs)和多模态模型(LMMs)有望加速生物医学发现,但其可靠性仍不明确。我们提出ARIEL(AI Research Assistant for Expert-in-the-Loop Learning),一个开源的评估与优化框架,结合精选的多模态生物医学语料库和专家审核的任务,用于检测两项核心能力:全文摘要生成与细粒度图像解析。采用统一协议与盲评的博士级评估,结果显示当前先进模型生成的摘要虽流畅但信息不全,而LMM在细节视觉推理上表现不佳。后续发现提示工程与轻量微调显著提升文本覆盖度,计算规模扩展的推理策略改善了视觉问答性能。我们构建了集成文本与视觉线索的ARIEL代理,能够提出可验证的机制假说。ARIEL明确了基础模型的现有优劣势,并为可信生物医学AI的发展提供了可复现的平台。

原文摘要 · Abstract (English)

Large language models (LLMs) and large multimodal models (LMMs) promise to accelerate biomedical discovery, yet their reliability remains unclear. We introduce ARIEL (AI Research Assistant for Expert-in-the-Loop Learning), an open-source evaluation and optimization framework that pairs a curated multimodal biomedical corpus with expert-vetted tasks to probe two capabilities: full-length article summarization and fine-grained figure interpretation. Using uniform protocols and blinded PhD-level evaluation, we find that state-of-the-art models generate fluent but incomplete summaries, whereas LMMs struggle with detailed visual reasoning. We later observe that prompt engineering and lightweight fine-tuning substantially improve textual coverage, and a compute-scaled inference strategy enhances visual question answering. We build an ARIEL agent that integrates textual and visual cues, and we show it can propose testable mechanistic hypotheses. ARIEL delineates current strengths and limitations of foundation models, and provides a reproducible platform for advancing trustworthy AI in biomedicine.

AI助手生物医学多模态专家协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。