arXiv:2605.17765cs.LG2026-05

让医疗模型的隐藏表示更清晰稳定,避免多个因素混淆。

AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models

论文配图:AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models
图 1 · 摘自论文原文
  • 将医疗数据的语义因素分解到正交子空间,提升可解释性。
  • 在多种临床任务中优于现有方法,且对机构分布变化更鲁棒。
  • 适合关注医疗模型可解释性与泛化能力的研究者。

近期医疗基础模型通过大规模自监督学习取得了优异的预测性能,但其隐含表示常将生理严重程度、干预强度、观测结构和机构工作流程等因素纠缠在同一嵌入方向中。虽然有利于下游预测,但这些表示语义模糊且在上下文变化下不稳定。我们提出AURORA(基于上下文潜在几何的自适应不确定性感知表示),一种新框架,通过分解表示为对应不同上下文因素的正交语义子空间,并在每个子空间内学习关系一致性目标,实现既语义解耦又几何可解释的潜在空间。在多个临床预测与检索任务中,AURORA持续优于重建、对比学习和自蒸馏基线,显著提升上下文解耦度、邻域纯度及机构分布偏移下的鲁棒性。结果表明,潜在几何本身是医疗基础模型设计的重要维度,按上下文语义显式构建表示空间,是超越传统预测压缩目标的补充方向。

原文摘要 · Abstract (English)

Recent healthcare foundation models have achieved strong predictive performance through large scale self supervised learning, yet their latent representations frequently entangle physiologic severity, intervention intensity, observational structure, and institutional workflow into shared embedding directions. While effective for downstream prediction, such representations remain semantically opaque and unstable under contextual shift. We introduce AURORA, Adaptive Uncertainty aware Representations through Orthogonalized Relational Alignment, a new framework for healthcare representation learning based on contextual latent geometry. Rather than optimizing a single unified embedding manifold, AURORA decomposes representations into orthogonal semantic subspaces corresponding to distinct contextual factors and learns relational consistency objectives within each subspace. This induces latent spaces that are both semantically disentangled and geometrically interpretable. Across multiple clinical prediction and retrieval tasks, AURORA consistently outperforms reconstruction, contrastive, and self distillation baselines while substantially improving contextual disentanglement, neighborhood purity, and robustness under institutional distribution shift. Our results suggest that latent geometry itself constitutes an important axis of healthcare foundation model design and that explicitly structuring representation space according to contextual semantics provides a complementary direction beyond conventional predictive compression objectives.

医疗表征解耦学习几何表示基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。