提出无需假设数据生成方式的非参数重演学习方法,解决未来不良结果预测问题。
Non-Parametric Rehearsal Learning via Conditional Mean Embeddings
- 用核方法将目标解耦为可预测性与行为影响两部分,摆脱线性假设限制。
- 在合成与半真实数据上实现更高准确率,验证了对非线性系统的适应性。
- 适合研究复杂决策系统、需避免不良后果的智能体建模者使用。
在机器学习中,一类关键的决策问题是防止预测到不良未来结果,即‘避免不良未来’(AUF)问题。为应对该问题,已有重演学习框架通过建模影响关系来实现有效决策。然而,现有方法依赖于线性系统或加性噪声等严格参数假设,限制了实际应用。本文首次提出一种无需特定数据生成函数形式的非参数重演学习方法。我们利用核工具将AUF目标重构为统一表示,分离可预测性建模与行为引起的分布变化。针对可预测性指标的不连续性,提出平滑的Probit替代函数,并提供近似误差界。同时,通过条件均值嵌入捕捉行为影响,设计基于核岭回归的嵌套估计器,确保目标估计的一致性。该方法天然适用于非线性系统和非加性噪声场景。在合成数据及由真实数据衍生的半真实基准上的实验表明,本方法在有效性与灵活性方面表现优异。
原文摘要 · Abstract (English)
In machine learning, a critical class of decision-related problems concerns preventing predicted undesirable outcomes, referred to as the \textit{avoiding undesired future} (AUF) problem. To address this, the \textit{rehearsal learning} framework has been proposed to model influence relations for effective decisions. However, existing rehearsal methods rely on restrictive parametric assumptions such as linear systems or additive noise, limiting their practical applicability. In this paper, we propose the first non-parametric rehearsal learning approach for AUF without assuming specific functional forms of data generation processes. Specifically, we use kernel machinery to reformulate the AUF objective into a unified representation that disentangles desirability modeling from action-induced distributional changes. To handle the discontinuity of desirability indicator, we present a smooth Probit surrogate and provide an approximation error bound. Meanwhile, we capture the action-induced changes via conditional mean embeddings, and develop a kernel ridge regression based nested estimator for AUF objective with consistency guarantees. Such a formulation naturally accommodates nonlinear systems and non-additive noise, and empirical results on synthetic and real-data-derived semi-synthetic benchmarks demonstrate the effectiveness and flexibility of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。