arXiv:2501.14956cs.CLcs.AI2025-01ACL被引 32

用AI自动生成可解释的个性化长文评估报告。

ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation

  • 让大模型分析生成文本与参考文本的细节匹配度。
  • 比现有方法更贴近人工判断,提升7.2%一致性。
  • 适合需要透明评估结果的研究者和开发者。

个性化长文本生成的评估极具挑战性,因为只有原始提示撰写者能可靠评价输出,但跨研究重复邀请同一用户不现实。本文提出ExPerT——一种可解释的基于参考文本的评估框架。ExPerT利用大语言模型从生成文本与参考文本中提取原子级特征及其证据,匹配特征并基于内容与写作风格进行对齐评估,这两项是个性化文本的核心属性。此外,ExPerT为每一步评估生成细粒度、可解释的说明,显著提升透明度。实验表明,与当前最优评估方法相比,ExPerT在与人工判断的一致性上实现7.2%的相对提升。人类评估者对解释可用性的评分高达4.7/5,证明其在增强决策可解释性方面的有效性。

原文摘要 · Abstract (English)

Evaluating personalized text generated by large language models (LLMs) is challenging, as only the LLM user, i.e., prompt author, can reliably assess the output, but re-engaging the same individuals across studies is infeasible. This paper addresses the challenge of evaluating personalized text generation by introducing ExPerT, an explainable reference-based evaluation framework. ExPerT leverages an LLM to extract atomic aspects and their evidence from the generated and reference texts, match the aspects, and evaluate their alignment based on content and writing style -- two key attributes in personalized text generation. Additionally, ExPerT generates detailed, fine-grained explanations for every step of the evaluation process, enhancing transparency and interpretability. Our experiments demonstrate that ExPerT achieves a 7.2% relative improvement in alignment with human judgments compared to the state-of-the-art text generation evaluation methods. Furthermore, human evaluators rated the usability of ExPerT's explanations at 4.7 out of 5, highlighting its effectiveness in making evaluation decisions more interpretable.

文本评估可解释性个性化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。