用信息检索思路优化大模型对齐,提升生成可信度。
LLM Alignment as Retriever Optimization: An Information Retrieval Perspective
- 将大模型生成类比为信息检索中的召回优化
- 在AlpacaEval2和MixEval-Hard上分别提升38.9%和13.7%
- 适合关注模型安全与对齐的科研人员
大型语言模型(LLMs)在推理、编程和沟通方面展现出革命性能力,推动了跨行业的创新。其真正潜力取决于有效对齐,以确保行为正确、可信且符合伦理,应对虚假信息、幻觉、偏见和滥用等挑战。现有基于强化学习的对齐方法复杂难控,直接优化方法则更为简洁。本文提出一种新颖的直接优化对齐方法,借鉴成熟的信息检索(IR)原理,构建了将大模型生成与奖励模型映射到检索器-重排序器范式的系统框架。在此基础上,提出LLM对齐作为检索偏好优化(LarPO),显著提升了整体对齐质量。大量实验验证了该方法的有效性:在AlpacaEval2和MixEval-Hard上平均分别提升38.9%和13.7%。本工作通过融合信息检索基础,为推进大模型对齐开辟新路径,提供了未来研究的重要方向。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have revolutionized artificial intelligence with capabilities in reasoning, coding, and communication, driving innovation across industries. Their true potential depends on effective alignment to ensure correct, trustworthy and ethical behavior, addressing challenges like misinformation, hallucinations, bias and misuse. While existing Reinforcement Learning (RL)-based alignment methods are notoriously complex, direct optimization approaches offer a simpler alternative. In this work, we introduce a novel direct optimization approach for LLM alignment by drawing on established Information Retrieval (IR) principles. We present a systematic framework that bridges LLM alignment and IR methodologies, mapping LLM generation and reward models to IR's retriever-reranker paradigm. Building on this foundation, we propose LLM Alignment as Retriever Preference Optimization (LarPO), a new alignment method that enhances overall alignment quality. Extensive experiments validate LarPO's effectiveness with 38.9 % and 13.7 % averaged improvement on AlpacaEval2 and MixEval-Hard respectively. Our work opens new avenues for advancing LLM alignment by integrating IR foundations, offering a promising direction for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。