不用参考文本也能精准检测生成文本的语义错误。
Cross-Examination Framework: A Task-Agnostic Diagnostic for Information Fidelity in Text-to-Text Generation
- 把原文和生成文当作独立知识库,互相提问验证。
- 三项评分可识别遗漏、矛盾等关键错误,准确率超传统指标。
- 适合评估翻译、摘要、医疗记录生成等任务,尤其擅长发现实体与关系错误。
传统指标如BLEU和BERTScore难以捕捉文本生成中的语义保真度。本文将交叉审查框架(CEF)应用于无参考文本的多维评估,将源文本与候选文本视为独立知识库,从中生成可验证问题并进行交叉审查,得出覆盖度、一致性与符合性三个可解释评分。在机器翻译、摘要和临床笔记生成任务中验证,该框架能有效识别内容遗漏与事实矛盾等标准指标忽略的错误。关键贡献包括系统性鲁棒性分析以选择稳定判断模型;参考文本有无模式间强相关性证明了其无需黄金参考即可可靠运行。人类专家验证表明,CEF检测出的不一致问题与语义层面重大错误高度吻合,尤其在实体与关系扭曲的识别上表现突出。
原文摘要 · Abstract (English)
Traditional metrics like BLEU and BERTScore fail to capture semantic fidelity in generative text-to-text tasks. We adapt the Cross-Examination Framework (CEF) for a reference-free, multi-dimensional evaluation by treating the source and candidate as independent knowledge bases. CEF generates verifiable questions from each text and performs a cross-examination to derive three interpretable scores: Coverage, Conformity, and Consistency. Validated across translation, summarization and clinical note-generation, our framework identifies critical errors, such as content omissions and factual contradictions, missed by standard metrics. A key contribution is a systematic robustness analysis to select a stable judge model. Crucially, the strong correlation between our reference-free and with-reference modes validates CEF's reliability without gold references. Furthermore, human expert validation demonstrates that CEF mismatching questions align with meaning-altering semantic errors higher than with non-semantic errors, particularly excelling at identifying entity-based and relational distortions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。