提出改进科学问答中引用真实性的方法,提升生成答案与引用文献的严格对应。
DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA

- 通过预生成过滤和后生成蕴含约束,增强答案引用真实性
- 前沿模型虽流畅但常脱离引用内容生成答案
- 适合关注可信生成与科学问答质量评估的研究者
本文描述了DS@GT ARC在CLEF 2026 LongEval Task 4(检索增强生成RAG)中的提交方案。研究发现传统自然语言评价指标与引用完整性之间存在偏差。我们评估了基于Corrective RAG(CRAG)和CiteFix的修正管道,在基准与前沿RAG QA模型上的表现。尽管前沿模型在答案相关性和流畅性上表现优异,但其大型语言模型作为裁判的诊断显示,这些模型可在不使用上下文的情况下正确识别相关文档。通过生成前的段落过滤和生成后的主张与引用材料严格蕴含关系验证,我们的修正管道在引用忠实度和答案根基性方面略有提升。研究建议,可信RAG QA的评估需引入强调严格答案根基性的新指标。
原文摘要 · Abstract (English)
This paper describes DS@GT ARC's submission to the CLEF 2026 LongEval Task 4 on Retrieval-Augmented Generation (RAG). In this submission, we examine a divergence between traditional natural language evaluation metrics and citation integrity as applied to RAG QA systems. We evaluate a corrective pipeline using Corrective RAG (CRAG) and CiteFix against baseline and frontier model benchmark RAG QA scores. While frontier models maximized answer relevance and fluency scores, our RAGAs LLM-as-judge diagnostics indicate that frontier models would correctly identify relevant documents without using their context in answer generation. Conversely, by filtering chunks pre-generation and enforcing strict entailment of generated claims to the cited material post-generation, our corrective pipeline marginally improved citation faithfulness and answer grounding. We propose that evaluation of trustworthy RAG QA requires metrics that reward strict answer grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。