arXiv:2503.13399cs.CVcs.AI2025-03CVPR被引 41

构建了用于显微镜科研的多模态推理基准,评估专家级图像理解、假设生成与实验设计能力。

MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research

  • 采用两阶段流程生成无语言捷径的多选题,提升题目质量
  • 顶尖模型在该基准上最高仅达53%准确率,显示多模态推理仍是难点
  • 适合从事生物医学AI、多模态大模型研究者使用

科学研究需要对多模态数据进行复杂推理,这在生物学领域尤为突出。尽管多模态大语言模型(MLLMs)在辅助科研方面取得进展,但现有基准最多仅覆盖大学水平难度,而研究级基准又侧重底层感知,难以衡量科学发现所需的复杂多模态推理能力。为此,我们提出MicroVQA,一个专为显微镜科研设计的视觉问答(VQA)基准,旨在评估三种关键科研能力:专家级图像理解、假设生成和实验提案。MicroVQA包含由生物学专家精选的1,042道多选题,覆盖多种显微技术,确保样本贴近真实科研实践。研究发现,标准多选题生成方法易引入语言捷径,因此我们设计了两阶段新流程:首先用优化后的LLM提示将问答对转化为多选题;再通过基于代理的RefineBot修正题目以消除捷径。对主流MLLM的测试表明,模型最佳表现仅为53%;使用较小语言模型的模型表现仅略低,说明语言推理相对容易,而多模态推理才是核心挑战;且用科研文章微调可提升性能。专家分析思维链发现,感知错误最常见,其次是知识错误和过度泛化错误。这些洞察揭示了多模态科研推理的关键障碍,表明MicroVQA是推动生物医学AI发展的宝贵资源。MicroVQA可在https://huggingface.co/datasets/jmhb/microvqa获取,项目页面见https://jmhb0.github.io/microvqa。

原文摘要 · Abstract (English)

Scientific research demands sophisticated reasoning over multimodal data, a challenge especially prevalent in biology. Despite recent advances in multimodal large language models (MLLMs) for AI-assisted research, existing multimodal reasoning benchmarks only target up to college-level difficulty, while research-level benchmarks emphasize lower-level perception, falling short of the complex multimodal reasoning needed for scientific discovery. To bridge this gap, we introduce MicroVQA, a visual-question answering (VQA) benchmark designed to assess three reasoning capabilities vital in research workflows: expert image understanding, hypothesis generation, and experiment proposal. MicroVQA consists of 1,042 multiple-choice questions (MCQs) curated by biology experts across diverse microscopy modalities, ensuring VQA samples represent real scientific practice. In constructing the benchmark, we find that standard MCQ generation methods induce language shortcuts, motivating a new two-stage pipeline: an optimized LLM prompt structures question-answer pairs into MCQs; then, an agent-based `RefineBot' updates them to remove shortcuts. Benchmarking on state-of-the-art MLLMs reveal a peak performance of 53\%; models with smaller LLMs only slightly underperform top models, suggesting that language-based reasoning is less challenging than multimodal reasoning; and tuning with scientific articles enhances performance. Expert analysis of chain-of-thought responses shows that perception errors are the most frequent, followed by knowledge errors and then overgeneralization errors. These insights highlight the challenges in multimodal scientific reasoning, showing MicroVQA is a valuable resource advancing AI-driven biomedical research. MicroVQA is available at https://huggingface.co/datasets/jmhb/microvqa, and project page at https://jmhb0.github.io/microvqa.

多模态推理显微镜科学智能VQA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。