arXiv:2511.11407cs.CV2025-11被引 1

构建高质量显微图像问答数据集,提升大模型在生物医学图像推理能力。

MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model

  • 基于专家验证的图文对与图结构过滤,构建高质量数据集。
  • 引入多模态一致性的图模型,自动识别并剔除不一致样本。
  • 适合研究生物医学视觉语言模型与数据构建的学者使用。

多模态大语言模型在生物医学成像中应用日益广泛,但显微图像科学推理仍受限于大规模高质量训练数据的缺乏。我们提出 MicroVQA++,一个基于 BIOMEDICA 数据库的三阶段大规模高质显微图像问答语料库。第一阶段从同行评审论文中的专家验证图文对中自举监督信号;第二阶段采用 HiCQA-Graph——一种融合自然语言推理、CLIP 视觉-语言对齐与代理信号的异构图结构,用于识别和过滤不一致样本;第三阶段由多模态大语言模型生成多项选择题,并经人工筛选。最终发布版本包含大规模训练集与人工校验测试集,其布卢姆认知层级难例分布优于 MicroVQA 基准。本工作贡献:(i) 融合专家文献与图结构过滤及人工精修的质量控制数据集;(ii) 首个联合建模 (图像, 文本, 问答) 的跨模态一致性过滤图结构;(iii) 证明精心构造数据可使 40 亿参数规模大模型达到接近 GPT-5 的显微图像推理性能,并在开源模型中取得领先。代码与数据将在审稿结束后公开。

原文摘要 · Abstract (English)

Multimodal Large Language Models are increasingly applied to biomedical imaging, yet scientific reasoning for microscopy remains limited by the scarcity of large-scale, high-quality training data. We introduce MicroVQA++, a three-stage, large-scale and high-quality microscopy VQA corpus derived from the BIOMEDICA archive. Stage one bootstraps supervision from expert-validated figure-caption pairs sourced from peer-reviewed articles. Stage two applies HiCQA-Graph, a novel heterogeneous graph over images, captions, and QAs that fuses NLI-based textual entailment, CLIP-based vision-language alignment, and agent signals to identify and filter inconsistent samples. Stage three uses a MultiModal Large Language Model (MLLM) agent to generate multiple-choice questions (MCQ) followed by human screening. The resulting release comprises a large training split and a human-checked test split whose Bloom's level hard-sample distribution exceeds the MicroVQA benchmark. Our work delivers (i) a quality-controlled dataset that couples expert literature with graph-based filtering and human refinement; (ii) HiCQA-Graph, the first graph that jointly models (image, caption, QA) for cross-modal consistency filtering; (iii) evidence that careful data construction enables 4B-scale MLLMs to reach competitive microscopy reasoning performance (e.g., GPT-5) and achieve state-of-the-art performance among open-source MLLMs. Code and dataset will be released after the review process concludes.

显微图像多模态数据集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。