arXiv:2511.03939cs.LGcs.AI2025-11综述被引 2

综述多模态、文化公平与低延迟对齐方法,为构建更公平高效的AI提供指南。

RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods

  • 梳理强化学习人类反馈基础算法,涵盖PPO、DPO与GRPO
  • 系统分析多模态对齐、文化公平性及低延迟优化最新进展
  • 适合关注AI伦理、高效训练与跨模态应用的研究者

强化学习从人类反馈(RLHF)是大语言模型对齐的标准方法,但近期进展已超越传统的文本对齐。本文综述对齐研究的新前沿,聚焦多模态对齐、文化公平性与低延迟优化三大关键缺口。通过回顾PPO、DPO和GRPO等基础算法,并深入分析最新技术,本文提供对各类方法的比较性综述,指出开放挑战,为构建更鲁棒、高效且公平的AI系统提供重要路线图。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) is the standard for aligning Large Language Models (LLMs), yet recent progress has moved beyond canonical text-based methods. This survey synthesizes the new frontier of alignment research by addressing critical gaps in multi-modal alignment, cultural fairness, and low-latency optimization. To systematically explore these domains, we first review foundational algo- rithms, including PPO, DPO, and GRPO, before presenting a detailed analysis of the latest innovations. By providing a comparative synthesis of these techniques and outlining open challenges, this work serves as an essential roadmap for researchers building more robust, efficient, and equitable AI systems.

RLHF多模态文化公平低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。