arXiv:2606.07130cs.CL2026-06

让AI生成带精准引用的文本,确保每句话都有据可查。

Explicit Evidence Grounding via Structured Inline Citation Generation

论文配图:Explicit Evidence Grounding via Structured Inline Citation Generation
图 1 · 摘自论文原文
  • 通过三种方法自动生成结构化引用,将论点与原文档和证据段落关联。
  • 在三个问答数据集上测试,发现大模型能找对文档但难定位精确证据段。
  • 适合需要高可信度生成的场景,如学术写作、医疗咨询等。

随着AI系统广泛应用,生成内容的准确性和可追溯性愈发重要。本文提出FullCite框架,与以往工作不同,该框架为每个陈述生成结构化内联引用,将论点与源文档及支持证据段落明确关联。提出了三种引文生成策略:基于提示的生成、基于引文语法的约束解码,以及事后跨度对齐。在ASQA、BioASQ和ExpertQA三个问答基准上,从文档正确性、证据段识别和论点-引用忠实度三个维度评估引文质量。结果显示,尽管大语言模型能有效识别相关文档,但在文档内精确定位支持性证据段方面表现不佳。这一差距表明,实现真正可信的有引用问答,需更重视精确证据段识别的研究。

原文摘要 · Abstract (English)

As AI systems become more widely adopted, the demand for factual and faithful generation grows. Properly attributing information through citations becomes, therefore, crucial. This work introduces FullCite, a framework that, in contrast to most previous works, generates structured inline citations linking each claim to both its source document and supporting evidence. FullCite proposes three strategies to inline citation generation: prompt-based generation, constrained decoding over a citation grammar, and posthoc span alignment. Using three question answering benchmarks, namely, ASQA, BioASQ, and ExpertQA, we assess citation quality and faithfulness along three dimensions: document-level correctness, evidence span identification, and claim-citation faithfulness. Our evaluation shows that while LLMs are generally effective at identifying relevant documents, they struggle to identify the precise supporting spans within them. This gap suggests that achieving faithful attributed QA will require research to place greater emphasis on precise evidence span identification.

引用生成可信生成问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。