arXiv:2607.17508cs.LGcs.AI2026-07

用自然语言描述任务,零样本生成可解释的医疗预测模型。

Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare

论文配图:Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare
图 1 · 摘自论文原文
  • 从任务描述中检索相似任务,迁移系数空间结构生成新模型
  • 零样本下达73.4%准确率,极少数样本下仍保持73.2%性能
  • 支持不确定性预警与特征级解释,适合需要人工审核的医疗场景

我们提出检索增强型可解释学习(RAIL),一种概率元学习框架,用于零样本生成任务特定且可解释的模型。RAIL通过自然语言任务描述和历史任务预测器记忆,检索相关源任务,将系数空间结构迁移至原始诊断特征空间,生成新预测器,实现零样本与少样本临床操作预测,并提供特征级解释。其概率化设计对检索、模型系数与预测结果均建模不确定性,支持可靠性感知部署:当预测不确定或解释不稳时,可标记以供临床复核,而非自动决策。该方法特别适用于医疗环境,其中任务分布长尾、新目标频繁出现,且模型必须可检视、知风险、兼容人类监督。在长尾临床操作预测任务中,RAIL在无训练数据的零样本设置下达到73.4%准确率,在仅2-4个样本的极端少样本场景下仍保持73.2%准确率,远超监督模型在类似条件下的近随机表现。此外,结合临床先验的任务表示使系统具备检索、不确定性与系数层面的诊断能力,提升模型透明性。结果表明,RAIL为可扩展的临床预测系统提供了兼顾适应性、可解释性与可靠性的路径。

原文摘要 · Abstract (English)

We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-specific interpretable models that synthesizes coefficient-space structure from natural-language task descriptions and a memory of previously learned task-specific predictors. RAIL retrieves related source tasks, transfers structure through coefficient space, and generates a new predictor in the original diagnostic-feature space, enabling zero-shot and few-shot clinical procedure prediction with feature-level explanations. Its probabilistic formulation provides uncertainty over retrieval, model coefficients, and predictions, supporting reliability-aware deployment: uncertain predictions or unstable explanations can be flagged for additional clinical review rather than treated as automatic decisions. This makes RAIL particularly suited for healthcare settings, where prediction tasks are highly long-tailed, new clinical targets arise frequently, and models must remain inspectable, uncertainty-aware, and compatible with human oversight. Across long-tailed clinical procedure prediction tasks, RAIL maintains reliable performance across data-availability regimes: it achieves 73.4% accuracy in the held-out zero-shot settings, where no supervised task-specific model can be trained, and remains near 73.2% accuracy in the extreme few-shot regime with only 2-4 examples, where supervised task-specific models perform close to chance. RAIL further benefits from clinically informed task representations and yields retrieval, uncertainty, and coefficient-level diagnostics that make model behavior more transparent. These results suggest a path toward scalable clinical prediction systems that can adapt to new tasks while preserving interpretability and reliability.

医疗AI零样本可解释性元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。