用顺序结构替代图结构,实现更稳健的未来避险决策
Order-based Rehearsal Learning
- 基于信息论方法学习决策顺序,不依赖具体函数形式和噪声分布
- 通过可微优化提升决策后成功概率,实验优于图/顺序学习方法
- 适合缺乏完整因果图但有观测数据的决策场景,如医疗或金融
当机器学习模型预测到不良事件时,常需采取措施避免其发生,即所谓避免不良未来(AUF)问题。现有回放学习方法多依赖底层图结构,但从观测数据中学习图结构困难且易产生显著估计误差。本文证明,仅使用顺序结构即可满足AUF决策需求,提出首个基于顺序的回放学习方法。尽管顺序信息少于图结构,但仍能从观测数据中识别决策影响,表明完全学习图结构并非必要。为此,我们设计一种无函数形式与噪声类型限制的信息论方法来学习顺序结构;针对AUF决策,构建基于顺序的采样器以近似决策影响,并结合最大化决策后成功率的代理目标,将问题转化为可微优化。实验表明,所提顺序学习方法优于现有方法,其AUF策略不仅超越依赖学习图或学习顺序的方法,甚至达到或超过已知真实图的基准性能。
原文摘要 · Abstract (English)
When a machine learning (ML) model forecasts an undesired event, one often seeks a decision to avoid it, known as the avoiding undesired future (AUF) problem. Many rehearsal learning methods have been proposed for AUF, but they rely on an underlying graph structure; learning such a graph from observational data is challenging and can incur substantial estimation error. In this work, we demonstrate that the order structure can be sufficient for AUF decision-making, and propose the first order-based rehearsal learning method. Although an order is less informative than a graph, it can be sufficient to identify the influence of decisions from observational data, suggesting that learning the entire graph is not always necessary. To learn the order, we develop an information-theoretic method that imposes no restrictions on the form of structural functions or the type of noise distributions. For AUF decision-making, we construct an order-based sampler to approximate the influence of decisions and, combined with a surrogate objective for maximizing the post-decision success probability, reduce the AUF task to a differentiable optimization problem. Experiments show that our order learning method outperforms existing methods, and that our AUF approach not only surpasses methods relying on learned graphs or learned orders, but also matches or even exceeds oracle baselines that are given the true graph.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。