让大模型自己检查并改进答案,提升生物医学搜索准确性。
Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
- 大模型通过自我反馈机制反复优化查询和答案。
- 不同模型在问答任务中表现差异明显,部分任务提升显著。
- 适合关注AI自主纠错与专业领域搜索的研究者。
代理式检索增强生成(RAG)与‘深度研究’系统旨在实现大型语言模型(LLMs)的自主搜索流程,通过迭代方式不断优化输出。然而,在生物医学等专业领域,自动化系统可能降低用户参与度,并偏离专家的信息需求。专业搜索任务需要高水准的用户知识与透明性。2025年BioASQ CLEF挑战赛采用专家设计的问题,为研究此类问题提供了平台。我们评估了Gemini-Flash 2.0、o3-mini、o4-mini和DeepSeek-R1等推理与非推理型模型的表现。核心方法是引入自反馈机制:模型自行生成、评估并修正其输出,用于查询扩展及多种答案类型(是/否、事实型、列表、理想答案)。实验发现,自反馈策略在不同模型与任务间效果不一。本研究揭示了大模型自纠错的能力边界,也为未来比较模型生成反馈与人类专家输入的有效性提供参考。
原文摘要 · Abstract (English)
Agentic Retrieval Augmented Generation (RAG) and 'deep research' systems aim to enable autonomous search processes where Large Language Models (LLMs) iteratively refine outputs. However, applying these systems to domain-specific professional search, such as biomedical research, presents challenges, as automated systems may reduce user involvement and misalign with expert information needs. Professional search tasks often demand high levels of user expertise and transparency. The BioASQ CLEF 2025 challenge, using expert-formulated questions, can serve as a platform to study these issues. We explored the performance of current reasoning and nonreasoning LLMs like Gemini-Flash 2.0, o3-mini, o4-mini and DeepSeek-R1. A key aspect of our methodology was a self-feedback mechanism where LLMs generated, evaluated, and then refined their outputs for query expansion and for multiple answer types (yes/no, factoid, list, ideal). We investigated whether this iterative self-correction improves performance and if reasoning models are more capable of generating useful feedback. Preliminary results indicate varied performance for the self-feedback strategy across models and tasks. This work offers insights into LLM self-correction and informs future work on comparing the effectiveness of LLM-generated feedback with direct human expert input in these search systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。