arXiv:2503.15289cs.CL2025-03ACL被引 8

让大模型追溯文本来源并识别生成关系,提升内容可信度。

TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship Classification

  • 通过句子级溯源与关系分类,精准追踪目标句的来源。
  • 在多文档长文本场景下,11个模型平均准确率超60%。
  • 适合关注AI内容可解释性与可信度的研究者使用。

大语言模型在文本生成中表现出极高的流畅性和连贯性,但其广泛应用引发了内容可靠性与责任归属的担忧。在高风险领域,理解内容的生成来源与方式至关重要。为此,我们提出文本溯源挑战TROVE,旨在将目标文本中的每个句子回溯至可能较长或跨多文档输入中的具体源句。除了识别来源,TROVE还标注细粒度的关系类型(引用、压缩、推断等),深入揭示每个目标句的形成过程。为构建基准数据集,我们利用三个公开数据集,覆盖英语和中文共11种不同场景(如问答、摘要),涵盖0–5k、5–10k、10k+长度的源文本,强调多文档与长文档设置的重要性。为确保数据质量,采用三阶段标注流程:句子检索、GPT-4o溯源与人工溯源。我们在直接提示与检索增强两种范式下评估11个大模型,结果显示:检索对性能提升至关重要;更大模型在复杂关系分类上表现更优;闭源模型普遍领先,但开源模型在检索增强下展现出显著潜力。数据集已公开:https://github.com/ZNLP/ZNLP-Dataset。

原文摘要 · Abstract (English)

LLMs have achieved remarkable fluency and coherence in text generation, yet their widespread adoption has raised concerns about content reliability and accountability. In high-stakes domains, it is crucial to understand where and how the content is created. To address this, we introduce the Text pROVEnance (TROVE) challenge, designed to trace each sentence of a target text back to specific source sentences within potentially lengthy or multi-document inputs. Beyond identifying sources, TROVE annotates the fine-grained relationships (quotation, compression, inference, and others), providing a deep understanding of how each target sentence is formed. To benchmark TROVE, we construct our dataset by leveraging three public datasets covering 11 diverse scenarios (e.g., QA and summarization) in English and Chinese, spanning source texts of varying lengths (0-5k, 5-10k, 10k+), emphasizing the multi-document and long-document settings essential for provenance. To ensure high-quality data, we employ a three-stage annotation process: sentence retrieval, GPT-4o provenance, and human provenance. We evaluate 11 LLMs under direct prompting and retrieval-augmented paradigms, revealing that retrieval is essential for robust performance, larger models perform better in complex relationship classification, and closed-source models often lead, yet open-source models show significant promise, particularly with retrieval augmentation. We make our dataset available here: https://github.com/ZNLP/ZNLP-Dataset.

文本溯源大模型可信度关系分类多文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。