arXiv:2608.05228cs.AIcs.CL2026-08

提出可自适应拆分事实的框架,兼顾精确性与上下文完整性。

TriQua: Reconciling Granularity and Context in Factuality Evaluation

  • 根据事实复杂度动态拆分为三元组或带附加条件的超关系事实
  • 在证据验证任务中表现优于现有方法,与人工评分高度一致
  • 适合需要精准纠错和可解释性的大模型事实核查场景

大语言模型的事实性评估常采用‘分解-验证’范式,但存在根本矛盾:原子事实(单句表达单一信息)常丢失必要上下文,而宽泛陈述又缺乏精细评估所需的粒度。为此,我们提出TriQua框架,根据事实复杂度灵活建模:简单声明以标准三元组形式提取,复杂声明则通过附加辅助上下文限定词表示为超关系事实。该自适应结构在保持准确检索与验证能力的同时,不牺牲原子性。此外,TriQua的验证过程可直接标注具体三元组与限定词中的错误,提供细粒度可解释性。我们还提出TriQuaScore来量化这些结构化事实单元的真实性。实证评估显示,TriQuaScore与人工标注的事实性评分高度一致,且其分解质量稳健,优于现有基于分解的框架,在基于证据的事实验证任务中表现更优。

原文摘要 · Abstract (English)

The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit essential context, while broader statements lack the granularity needed for precise assessment. To address this, we introduce TriQua, a framework that flexibly models facts based on their complexity. Simple claims are extracted as standard triples, while complex claims are represented as hyperrelational facts by attaching auxiliary contextual qualifiers. This adaptive structure preserves the necessary context for accurate retrieval and verification without sacrificing atomicity. Furthermore, TriQua's verification process directly annotates concrete errors within specific triples and qualifiers, providing fine-grained explainability for error detection. Alongside the framework, we propose TriQuaScore to quantify the factuality of these structured fact units. Empirical evaluations show that TriQuaScore strongly aligns with human annotated factuality scores, TriQua achieves robust decomposition quality, and outperforms existing decomposition-based frameworks in evidence-based fact verification.

事实核查可解释性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。