不同健康大模型能通过符号对齐实现跨模型信息迁移。
Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer

- 用线性分解提取模型中的符号成分,再通过线性映射对齐。
- 跨模型符号分类器性能保留超95%,双向一致。
- 适合研究医疗表征学习与多模型协同的学者。
我们发现,即使独立训练,基于约2000万分钟可穿戴传感器数据(来自约17.2万名参与者)的健康基础模型(FMs)也能在事后实现信息迁移。通过对其依赖数据的坐标系统进行对齐,从冻结嵌入中使用线性分解方法提取候选符号组件,并以简单线性映射实现跨模型对齐。对齐后的符号能选择性关联健康状况与生理特征,在不同模态和架构间具有相似关联性。在一个模型上训练的分类器应用于另一模型时,其域内性能保留超过95%,且双向表现相当。总体表明,独立训练的健康基础模型趋向于共享同一生理底层表征。
原文摘要 · Abstract (English)
We show that information can be transferred post-hoc across independently trained health foundation models (FMs), each pretrained on ~20M minutes of wearable sensor data from ~172K participants, by aligning their data-dependent coordinate systems. From frozen embeddings we extract candidate symbol-like components using linear decomposition methods, and align them across models with simple linear maps. Aligned symbols associate selectively with health conditions and physiological attributes, with associations similar across modalities and architectures. A classifier trained on one model's symbols and applied to another retains more than 95% of its in-domain performance, with similar retention in both directions. Overall, our results indicate that independently trained health FMs converge toward a common representation of the same underlying physiology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。