arXiv:2605.08378cs.LGcs.AI2026-05

让强化学习更高效可信,解决分布式部署与人类对齐难题

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems

论文配图:Reinforcement Learning for Scalable and Trustworthy Intelligent Systems
图 1 · 摘自论文原文
  • 用联邦优化提升多智能体系统通信效率和异构计算适应性
  • 通过偏好对齐减少大模型在对话中泄露隐私等不当信息
  • 适合关注AI可扩展性与安全性的研究者与工程师

强化学习已成为提升智能系统能力的重要范式,但实际部署面临两大挑战:其一,在通信带宽受限、计算资源异构的分布式环境中,强化学习需实现高效扩展;其二,随着强化学习用于大语言模型后训练和自主代理,所优化策略必须符合人类偏好并满足隐私等安全要求。本论文从联邦优化、偏好对齐与情境安全三个方向提出四项互补贡献。第一部分研究联邦环境下的可扩展强化学习,第二部分聚焦大语言模型的可信强化学习。这些工作使强化学习在两个维度上取得进展:一方面通过通信高效与异步联邦优化提升可扩展性;另一方面通过增强与人类偏好的对齐及降低语言系统中的情境不当信息泄露,提升可信度。整体而言,本论文主张下一代智能系统需兼具高效优化与可信行为,而强化学习为此提供了统一框架。

原文摘要 · Abstract (English)

Reinforcement learning has become a powerful paradigm for improving the capability of intelligent systems, but its practical deployment faces two central challenges. First, reinforcement learning must scale efficiently in distributed environments where communication bandwidth is limited and computation is heterogeneous across agents. Second, as reinforcement learning is increasingly used in post-training large language models and autonomous agents, the optimized policies must also be aligned with human preferences and satisfy safety requirements such as privacy-aware information disclosure. This dissertation addresses both challenges through four complementary contributions spanning federated optimization, preference alignment, and contextual safety. The first part of the dissertation studies scalable reinforcement learning in federated settings. The second part of the dissertation studies trustworthy reinforcement learning for large language models. Together, these contributions advance reinforcement learning along two complementary dimensions. On the one hand, they make reinforcement learning more scalable through communication-efficient and asynchronous federated optimization. On the other hand, they make reinforcement learning more trustworthy by improving alignment with human preferences and by reducing contextually inappropriate information disclosure in language-based intelligent systems. As a whole, this dissertation argues that the next generation of intelligent systems will require both efficient optimization and trustworthy behavior, and that reinforcement learning provides a unifying framework for addressing both goals.

强化学习联邦学习可信AI大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。