arXiv:2608.02615cs.CLcs.AI2026-08

构建跨影像病理基因组的癌症多模态问答基准,推动全癌种精准诊断研究。

OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

论文配图:OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
图 1 · 摘自论文原文
  • 基于TCGA数据构建9281例患者多模态问答数据集,覆盖32种癌症。
  • 提出OncoVLM模型,在全模态下多项指标领先现有模型10.7分。
  • 适合医学AI、肿瘤多组学融合与临床决策支持系统研究者使用。

癌症诊断与表征需整合影像学、病理学、基因组学及临床信息。然而,现有医疗大模型与视觉语言模型基准多聚焦单一模态或简单图文任务,缺乏对多证据流患者级评估的测试。我们提出OncoTriad-QA,一个面向泛癌种的患者级影像-病理-基因组问答基准。该数据集包含来自32个癌症队列的9,281例TCGA患者,涵盖86.1万条语义问题,对齐CT/MRI影像、全切片病理图像、体细胞突变、拷贝数变异、DNA甲基化、批量RNA-seq及临床数据。病例标注通过基于来源的LLM辅助流程构建,以标准化报告、分子谱和模态衍生证据为真值源,并经自动化一致性检查与临床医生审核。我们还引入参考多模态模型OncoVLM,通过学习投影器将影像、病理、甲基化与RNA-seq证据映射至大模型接口。实验表明,现有通用及医学大模型在综合泛癌问答中仍受限,尤其当问题需融合影像、形态与分子证据时。在OncoTriad-QA上微调后,OncoVLM在多项选择题准确率与BERTScore-F1上平均优于MedGemma-4B 10.7分,且在仅影像、仅病理及全模态设置下均表现更优,验证了该基准在训练与评估集成癌症问答模型中的价值。

原文摘要 · Abstract (English)

Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated modalities or narrow image-text tasks, leaving patient-level oncology assessment across multiple evidence streams largely untested. We introduce OncoTriad-QA, a patient-level radiology-pathology-genomics benchmark for pan-cancer question answering. OncoTriad-QA contains 86.1k semantic questions across 9,281 TCGA patient cases from 32 cancer cohorts, aligning CT/MRI radiology, whole-slide histopathology, somatic mutations, copy-number alterations, DNA methylation, bulk RNA-seq, and clinical metadata. Case-specific annotations are constructed through a source-grounded LLM-assisted pipeline using curated labels, diagnostic reports, molecular profiles, and modality-derived evidence as primary sources of truth, with automated consistency checks and clinician review. We also introduce OncoVLM, a reference multimodal model that maps modality-native radiology, pathology, DNA methylation, and RNA-seq evidence into an LLM interface through learned projectors. Experiments show that existing general-purpose and medical LLMs remain limited on comprehensive pan-cancer QA, especially when questions require integrating imaging findings, tumor morphology, and molecular evidence. After fine-tuning on OncoTriad-QA, OncoVLM exceeds MedGemma-4B by an average of 10.7 points when using MCQ accuracy and BERTScore-F1, with consistent gains across multiple-choice and open-ended questions under radiology-only, pathology-only, and all-available settings. These results demonstrate the benchmark's value for training and evaluating models for integrated cancer question answering.

多模态癌症诊断大模型基因组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。