arXiv:2510.15127q-bio.QMcs.LG2025-10

用博弈反演法分析呼吸机数据,揭示不同通气策略的相对后果。

Inferring Relative Consequences of Mechanical Ventilation from Observational Data Using Game-Based Comparisons

  • 通过博弈反演对比不同状态下的临床后果,构建可解释的评估框架。
  • 发现通气效果受患者群体、时间点和比较维度影响,具有强上下文依赖性。
  • 为强化学习优化个性化通气方案提供可落地的因果推断基础。

重症监护中识别机械通气(MV)方案的效果,需分析临床决策环境中异构患者-呼吸机系统产生的数据。这些耦合组件间的多尺度交互导致高维状态空间,尽管数据量大,仍采样稀疏。现有数据分析对理解当前呼吸管理实践及生成可检验假设至关重要。数据规模与复杂性促使采用强化学习(RL)探索数据一致的反事实轨迹。但实际应用中需定义时空依赖的奖励过程,以刻画状态到结果的关系及其上下文依赖性与后果延迟。这些要素无法先验获得,须从数据中通过假设推断。为此,本文通过求解博弈型逆问题,对比分类观测状态的相对后果,识别下游概率与随机方法(如强化学习)所需的比较模型,用于寻求通气优化与个性化。逆博弈推断在合成数据上验证,揭示潜在陷阱后应用于真实ICU数据,暴露了数据生成过程的复杂性。临床数据表明,通气类型后果及其相对排序均具内在上下文与时序依赖性,随患者亚组、时间与比较指标变化,且作用时长各异。讨论提出基于实证数据与博弈推断比较构建状态转移模型,以模拟通气管理行为的影响。

原文摘要 · Abstract (English)

Identifying the effects of mechanical ventilation (MV) protocols in critical care requires analyzing data from heterogeneous patient-ventilator systems in the clinical decision-making environment. Multiscale interactions among these coupled components generate a high-dimensional state space that remains sparsely sampled despite extensive data collection. Analysis of existing data is essential for understanding current respiratory management practices and generating testable hypotheses about improvement. The scale and complexity of available data motivate the use of reinforcement learning (RL) to explore data-consistent counterfactual trajectories. However, formulating RL in practical applications requires a spatiotemporally dependent reward process that defines state-to-consequence relationships, their context dependence, and the delays over which consequences emerge. These poorly understood elements are not known \emph{a priori} and inferred from data via hypotheses. To that end, categorized observed states are contrasted according to their relative consequences by solving a game-based inverse problem that identifies a comparison model required for downstream probabilistic and stochastic methods such as reinforcement learning for seeking MV optimization and personalization. The inverted-game inference is validated on synthetic data to reveal potential caveats before proceeding to real-world ICU data applications that expose complexities of the data-generating process. Clinical data applications revealed that both breath-type consequences and their relative ordering are inherently context- and time-dependent, varying across patient subgroups, time, and comparison quantities, and effect timescale. The discussion includes potential developments toward a state transition model for simulating the effects of MV management actions using empirical data and game-inferred comparisons.

机械通气因果推断强化学习重症监护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。