提出双流校准框架,让模型推理时动态内化临床细节。
From Exposure to Internalization: Dual-Stream Calibration for In-context Clinical Reasoning
- 双流协同校准:语义流降熵稳定生成,结构流迭代优化推理逻辑。
- 在13个临床数据集上超越现有方法,测试时学习效果显著。
- 适合需要精准个性化推理的医疗AI场景,如诊断辅助与病历分析。
上下文临床推理依赖于对复杂异构病历的稳健推断。尽管当前最先进的微调、上下文学习(ICL)和检索增强生成(RAG)能实现知识暴露,却常难以实现真正的上下文内化:即在推理时动态调整模型内部表示以适应个体病例的细微差异。为此,我们提出双流校准(DSC),一种测试时训练框架,突破表面知识暴露,实现在推理过程中的深度内化。DSC通过协同对齐两条校准流来促进输入内化。不同于被动的上下文暴露,语义校准流通过最小化熵强制对核心证据进行深思熟虑,内化语义锚点以稳定生成轨迹;同时,结构校准流通过迭代元学习目标吸收潜在推理依赖。通过在测试时使用专业支持集训练,该流使模型弥合外部证据与内部逻辑之间的差距,将碎片化数据整合为连贯响应。我们的方法将推理范式从被动注意力匹配转变为对潜在推理空间的主动精炼。在十三个临床数据集上验证,DSC在三种不同任务范式中均表现优异,持续超越从依赖训练的模型到测试时学习框架的最先进基线。
原文摘要 · Abstract (English)
Contextual clinical reasoning demands robust inference grounded in complex, heterogeneous clinical records. While state-of-the-art fine-tuning, in-context learning (ICL), and retrieval-augmented generation (RAG) enable knowledge exposure, they often fall short of genuine contextual internalization: dynamically adjusting a model's internal representations to the subtle nuances of individual cases at inference time. To address this, we propose Dual-Stream Calibration (DSC), a test-time training framework that transcends superficial knowledge exposure to achieve deep internalization during inference. DSC facilitates input internalization by synergistically aligning two calibration streams. Unlike passive context exposure, the Semantic Calibration Stream enforces a deliberative reflection on core evidence, internalizing semantic anchors by minimizing entropy to stabilize generative trajectories. Simultaneously, the Structural Calibration Stream assimilates latent inferential dependencies through an iterative meta-learning objective. By training on specialized support sets at test-time, this stream enables the model to bridge the gap between external evidence and internal logic, synthesizing fragmented data into a coherent response. Our approach shifts the reasoning paradigm from passive attention-based matching to an active refinement of the latent inferential space. Validated against thirteen clinical datasets, DSC demonstrates superiority across three distinct task paradigms, consistently outstripping state-of-the-art baselines ranging from training-dependent models to test-time learning frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。