arXiv:2606.13556cs.AIcs.HC2026-06

用基因信息做个人生理基准,区分身体变化是天生的还是环境造成的。

Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation

论文配图:Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation
图 1 · 摘自论文原文
  • 以基因数据为先验,构建个体化生理基准线
  • 动态调整基因与观测数据权重,实现渐进式个性化分析
  • 适合临床研究与精准健康管理,避免误判环境影响

个性化健康人工智能系统面临冷启动难题:生理解读模型需数周行为数据才能区分先天差异与环境影响。本文提出基于因果推断与贝叶斯先验设计的解决方案。个体基因组作为外生遗传锚点——一个出生即确定、不受反向因果干扰的个性化先验,在首次行为观测前即可使用。该锚点初始化个体生理设定值 <strong>G-hat = mu + sum(beta_i * g_i)</strong>,其中 beta_i 为全基因组关联研究(GWAS)得出的效应值,g_i 为风险等位基因数量。每条生理测量值 P 产生非体质偏差 <strong>delta = P - G-hat</strong>,分离出环境与状态相关信号。随着行为数据积累,先验按 <strong>G-hat_t = w(t)*G-hat_genomic + [1-w(t)]*P-bar_t</strong> 动态衰减,从基因主导过渡到经验基线主导。相同心率变异性(HRV)55 ms 在基因预测值为80 ms者中提示抑制假设,而在预测值为30 ms者中提示增强假设——无个性化锚点则无法实现这种反转。本框架覆盖六个生理领域,依据证据强度分级基因先验,明确区分已验证锚点(FTO、FADS1/2、FKBP5)与争议候选基因(SLC6A4、MAOA、DRD2)。阐明关联、孟德尔随机化与个体因果推断的边界,并提出四条部署约束:证据分级先验、动态衰减机制、匹配祖先的效应值、归因而非确定性输出。

原文摘要 · Abstract (English)

Personalized health AI systems face a fundamental cold-start problem: machine learning models for physiological interpretation require weeks of individual behavioral data before they can distinguish constitutional variation from environmentally driven deviation. We propose a solution grounded in causal inference and Bayesian prior design. An individual's genomic profile serves as an exogenous genetic anchor -- a domain-informed, personalized prior that is fixed at conception, immune to reverse causation, and available before a single behavioral observation is collected. The anchor initializes a Bayesian belief state over an individual's physiological set point G-hat = mu + sum(beta_i * g_i), where beta_i are GWAS-derived effect sizes and g_i are risk-allele counts. Each incoming physiological measurement P produces a non-constitutional deviation delta = P - G-hat that separates the signal attributable to environment and state from the constitutionally fixed baseline. As behavioral data accrue, the prior decays according to G-hat_t = w(t)*G-hat_genomic + [1-w(t)]*P-bar_t, transitioning from genome-dominated to empirical-baseline-dominated inference. The same observed HRV of 55 ms generates a suppression hypothesis for a person whose prior predicts 80 ms, and an enhancement hypothesis for a person whose prior predicts 30 ms -- a reversal impossible without a personalized anchor. We develop this architecture across six physiological domains, grading genomic priors by evidence strength, distinguishing robustly replicated anchors (FTO, FADS1/2, FKBP5) from contested candidate genes (SLC6A4, MAOA, DRD2). We address the inference boundary between association, Mendelian randomization, and individual token causation, and define four constraints for deployment: evidence-graded priors, dynamic decay, ancestry-matched effect sizes, and attribution rather than deterministic output.

个性化医疗贝叶斯推断基因-环境交互生理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。