构建放射科癌症筛查的多层推理数据集,让AI学会像医生一样思考。
RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology

- 按推理深度分层设计三类问题:感知、单步规则、多步链式推理
- 包含20,362例扫描、43种癌种,每道复杂题配可验证的推理链条
- 适合训练和评估医疗AI的临床推理能力,尤其支持强化学习
癌症筛查是一项推理任务。放射科医生需观察影像、对比既往扫描、结合临床背景并得出诊断结论,最终由病理结果验证。我们提出RadThinking,一个视觉问答(VQA)数据集,显式呈现这一推理过程并支持可训练性。该数据集包含三个难度层级的VQA对:基础级为原子感知问题;单步推理级应用一条临床规则;组合级需多步链式思考以达到如LI-RADS-5等指南分类。每个组合问题均配套其解决所需的底层问题链,且遵循临床报告标准规则。数据集涵盖9,131名患者共20,362例CT扫描,覆盖43种癌症类型,并包括2,077例经验证的健康对照组(随访超1年)。据我们所知,RadThinking是首个按推理深度分级且将组合推理锚定于临床标准的癌症筛查VQA语料库。基础级提供原子感知监督,组合级提供链式推理数据与可验证奖励,适用于DeepSeek-R1、OpenAI o1等强化学习方法。该数据集使AI系统不仅能检测癌症,更能真正实现临床推理的系统性训练与评估。
原文摘要 · Abstract (English)
Cancer screening is a reasoning task. A radiologist observes findings, compares them to prior scans, integrates clinical context, and reaches a diagnostic conclusion confirmed by pathology. We present RadThinking, a Visual Question Answering (VQA) dataset that makes this reasoning explicit and trainable. RadThinking releases VQA pairs at three difficulty tiers. Foundation VQAs are atomic perception questions. Single-step reasoning VQAs apply one clinical rule. Compositional VQAs require multi-step chain-of-thought to reach a guideline category such as LI-RADS-5. For every compositional VQA, we release the chain of foundation VQAs that solves it. The chain follows the rules of the governing clinical reporting standard. The dataset spans 20,362 CT scans from 9,131 patients across 43 cancer groups, plus 2,077 verified healthy controls with >1-year follow-up. To our knowledge, RadThinking is the first cancer-screening VQA corpus that stratifies questions by reasoning depth and grounds compositions in clinical reporting standards. The foundation tier supplies atomic perception supervision. The compositional tier supplies chain-of-thought data and verifiable rewards for reinforcement-learning recipes such as DeepSeek-R1 and OpenAI o1. RadThinking enables systematic training and evaluation of whether AI systems can reason about cancer, not merely detect it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。