用组学数据做可验证的形态假设,提升乳腺癌多任务预测性能。
HERO: Hypothesis-Driven Evidence Retrieval from Omics for Multi-Task Breast Cancer Analysis

- 将组学数据转化为形态假设向量,驱动精准区域检索。
- 在TCGA-BRCA上多项指标达新纪录,优于多模态融合与VLM基线。
- 结果可逐字追溯,适合需要可解释性的临床研究场景。
配对多组学数据可提升基于全切片图像(WSI)的生物标志物与预后预测,但现有流程多将组学作为并行特征或文本上下文,而非显式检索约束。HERO探究组学是否可作为可检验的形态学假设:稀疏通路到形态先验将DNA甲基化与miRNA映射为16维意图向量m,基于结构化10个描述的TF-IDF检索选出终点相关区域,余弦门c=cos(m,v)在c<τc时触发确定性缺陷修复。该闭环设计限制视觉语言模型调用,降低对嵌入语义匹配依赖,并使每步检索与验证均可逐字审计。在TCGA-BRCA数据集(930张WSI,患者级5折交叉验证)上,HERO在ER、PR、HER2、亚型及风险预测等多项任务中均达到新最优表现,超越多模态融合与VLM基线方法。
原文摘要 · Abstract (English)
Matched multi-omics can improve WSI-based biomarker and prognosis prediction, but most existing pipelines use omics as a paral lel feature stream or textual context rather than as an explicit retrieval constraint. HERO asks whether observed omics can be a testable mor phology hypothesis: a sparse pathway-to-morphology prior maps DNA methylation and miRNA into a K-dimensional intent vector m (K=16), TF-IDF retrieval over structured 10 captions selects endpoint-relevant regions, and a cosine gate c=cos(m,v) triggers deterministic deficit driven repair when c<τc. This closed-loop design bounds VLM calls, reduces reliance on embedding-based semantic matching, and makes every retrieval and verification step lexically auditable. On TCGA-BRCA (930WSIs, patient-level 5-fold CV), HERO sets new state-of-the-art across ER, PR, HER2, subtype, and risk prediction, outperforming both multimodal fusion and VLM-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。