arXiv:2504.09192cs.LG2025-04

提升强化学习与博弈算法在不确定环境下的效率、鲁棒性与泛化能力。

Towards More Efficient, Robust, Instance-adaptive, and Generalizable Sequential Decision making

  • 设计适应具体任务的自适应算法,兼顾理论保证与实际表现。
  • 在推荐系统、机器人、大模型微调等场景中实现更优性能。
  • 适合关注算法可靠性与真实应用落地的研究者和工程师。

本博士研究旨在开发数据驱动的序列决策方法,在不确定性下具备可证明的高效性与实用性。研究聚焦于强化学习(RL)、多臂老虎机及其在推荐系统、计算机网络、视频分析和大语言模型(LLMs)中的应用。序列决策方法如老虎机与强化学习已取得显著成功,从超越人类玩家的复杂游戏(如Atari、Go)到推动机器人、推荐系统及大模型微调的发展。然而,许多现有算法依赖理想化模型,在模型误设或对抗扰动下可能失效,尤其在缺乏准确先验知识或存在恶意用户动态系统的场景中尤为明显。此类挑战在现实应用中普遍存在,要求算法具备鲁棒性与自适应能力。此外,最坏情况下的理论保证往往无法反映实例相关性能,导致实际效率不足;同时,泛化至未见环境的能力对部署于动态不可预测系统至关重要。为此,本研究致力于构建更高效、鲁棒、实例自适应且泛化能力强的强化学习与老虎机算法。

原文摘要 · Abstract (English)

The primary goal of my Ph.D. study is to develop provably efficient and practical algorithms for data-driven sequential decision-making under uncertainty. My work focuses on reinforcement learning (RL), multi-armed bandits, and their applications, including recommendation systems, computer networks, video analytics, and large language models (LLMs). Sequential decision-making methods, such as bandits and RL, have demonstrated remarkable success - ranging from outperforming human players in complex games like Atari and Go to advancing robotics, recommendation systems, and fine-tuning LLMs. Despite these successes, many established algorithms rely on idealized models that can fail under model misspecifications or adversarial perturbations, particularly in settings where accurate prior knowledge of the underlying model class is unavailable or where malicious users operate within dynamic systems. These challenges are pervasive in real-world applications, where robust and adaptive solutions are critical. Furthermore, while worst-case guarantees provide theoretical reliability, they often fail to capture instance-dependent performance, which can lead to more efficient and practical solutions. Another key challenge lies in generalizing to new, unseen environments, a crucial requirement for deploying these methods in dynamic and unpredictable settings. To address these limitations, my research aims to develop more efficient, robust, instance-adaptive, and generalizable sequential decision-making algorithms for both reinforcement learning and bandits. Towards this end, I focus on developing more efficient, robust, instance-adaptive, and generalizable for both general reinforcement learning (RL) and bandits.

强化学习决策优化鲁棒性泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。