医学影像研究应从生成掩码转向结构化解析,让模型不仅能定位病灶,还能描述其特征与关联。
Beyond Masks: The Case for Medical Image Parsing

- 提出医疗图像解析新范式,整合实体、属性和关系的统一结构化输出
- 11个代表性系统均未生成完整可验证的解析结果,属性与关系识别几乎空白
- 强调模型需学会解释而非仅测量,适合临床决策支持与多模态融合研究者
过去十年,医学影像研究在逐体素分割方面取得显著进展,生成的掩码可用于计算大小、体积和位置,已成为临床基础设施的核心。然而,放射科报告中包含的信息几乎无法通过掩码表达。本文主张将医学图像解析作为核心输出:一种包含实体、属性与关系的结构化表示,且三者相互一致。实体指命名的解剖结构或病灶(存在/不存在);属性描述其特征,如边缘规则性、强化模式或严重程度;关系则建立结构间的空间或演化关联。一个理想的解析需满足三个标准:(1)决策正确性(命名当前图像中的关键项),(2)重建能力(内容足够丰富以复现原图),(3)预测能力(能推断患者状态演变)。量化指标由解析内容推导,而非与之并行预测。我们对11个代表性系统进行评估,发现尽管实体识别已基本解决,但属性、关系和闭包性仍近乎空白。未来方向并非新架构,而是对更丰富输出的承诺,以及能奖励此类输出的训练信号。分割教会模型测量,而解析要求模型解释。
原文摘要 · Abstract (English)
Medical imaging research has spent a decade getting very good at one thing: producing per-voxel masks. Masks tell us size, volume, and location, and a decade of clinical infrastructure rests on those outputs. Yet the report a radiologist writes contains almost nothing a mask can express. We argue that medical imaging research should adopt medical image parsing as its central output: a structured representation in which entities, attributes, and relationships are emitted together and mutually consistent. Entities are the named structures and findings, present or absent. Attributes describe those entities, capturing things like margin regularity, enhancement pattern, or severity grade. Relationships connect them, naming where one structure sits relative to another, what abuts what, and what has changed since the prior scan. A good parse satisfies three properties, in order: (1) decision (the parse names the right things in the current image), (2) reconstruction (its content is rich enough to regenerate that image), and (3) prediction (its content is rich enough to forecast how the patient state will evolve). Quantitative measurements are derived from this content; they are not predicted alongside it. To test how close the field is to producing such an output, we audit eleven representative systems against the three parsing primitives plus closure. None emits a well-formed parse. Entities are largely solved. Attributes, relationships, and closure remain near-empty. The path forward is not a new architecture. It is a commitment to a richer output, and to training signals that reward it. Segmentation taught models to measure. Parsing asks them to explain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。