梳理提升大模型对话可靠性的三大类对齐技术。
Factors That Support Grounded Responses in LLM Conversations: A Rapid Review
- 按模型生命周期分三类对齐方法:推理时、后训练、强化学习。
- 推理时方法无需重训,能有效对齐用户意图并减少幻觉。
- 适合关注大模型对话质量与可信度的研究者与开发者。
大语言模型在对话中可能生成与用户意图不符、缺乏上下文依据或出现幻觉的内容,影响基于LLM应用的可靠性。本研究采用基于PRISMA框架和PICO策略的快速综述方法,系统识别并分析了提升对话对齐性、确保上下文接地性、减少幻觉与话题漂移的技术。所识别的方法按其作用于模型生命周期的阶段分为三类:推理时、后训练及基于强化学习的方法。其中,推理时方法表现尤为高效,可在不重新训练的前提下实现输出对齐,同时支持用户意图理解、保持上下文连贯性并抑制幻觉。这些技术为提升大模型对话在关键对齐目标上的质量与可靠性提供了结构化机制。
原文摘要 · Abstract (English)
Large language models (LLMs) may generate outputs that are misaligned with user intent, lack contextual grounding, or exhibit hallucinations during conversation, which compromises the reliability of LLM-based applications. This review aimed to identify and analyze techniques that align LLM responses with conversational goals, ensure grounding, and reduce hallucination and topic drift. We conducted a Rapid Review guided by the PRISMA framework and the PICO strategy to structure the search, filtering, and selection processes. The alignment strategies identified were categorized according to the LLM lifecycle phase in which they operate: inference-time, post-training, and reinforcement learning-based methods. Among these, inference-time approaches emerged as particularly efficient, aligning outputs without retraining while supporting user intent, contextual grounding, and hallucination mitigation. The reviewed techniques provided structured mechanisms for improving the quality and reliability of LLM responses across key alignment objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。