arXiv:2606.14650cs.LG2026-06

提出新型图结构半赌盘算法,有效处理非线性奖励关系,提升决策效率。

Graph Structured Combinatorial Semi-Bandit with Nonlinear Reward Associations through Separable Signals

论文配图:Graph Structured Combinatorial Semi-Bandit with Nonlinear Reward Associations through Separable Signals
图 1 · 摘自论文原文
  • 基于图因果建模与核方法,构建可适配的非线性奖励分析框架
  • 理论证明收益随时间呈次线性、数据量呈线性增长,性能稳定
  • 适用于交通网络等复杂系统优化,尤其适合缺乏先验知识的场景

在海量互联数据中识别最优结构需要大量采样与计算。学习并利用潜在信号依赖关系可显著提升效率与预测能力,但普遍存在的非线性统计关系增加了挑战。本文提出新颖的通用自适应策略,包含基于图的因果奖励建模、解析再生核方法及函数过程的泰勒近似。理论上证明了收益随时间呈次线性、随数据量呈线性增长。分析涵盖噪声干扰、渐进模型收敛和解空间不匹配等多种不确定性下的鲁棒性。该框架仅需极少假设或先验估计,具备广泛适用性;多种变体可应对特定或扩展场景。通过基准合成数据与真实交通数据集的数值实验,验证了其实际有效性。

原文摘要 · Abstract (English)

The identification of optimal structures within vast arrays of interconnected data necessitates significant sampling- and computational effort. Learning and leveraging underlying signal dependencies can improve efficiency and predictive capabilities considerably, but the ubiquity of nonlinear statistical relations amplifies the complexity of such undertakings. In this paper, we develop novel generic and adaptive strategies equipped with routines for graph-based causal reward modeling, analytic reproducing kernel methods, and Taylor approximation of functional processes. We establish theoretical performance guarantees sublinear in time and linear in data volume over time. Our analyses cover robustness to a multitude of uncertainties arising from noise interference, gradual model convergence, and solution space mismatch. The framework's general appeal is substantiated by a minimalistic set of conditions or reliance on prior estimates, while various outlined modifications address specific or extended settings. To demonstrate practical effectiveness, we conduct numerical experiments using both benchmarked synthetic and real-world transportation datasets.

图学习强化学习非线性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。