arXiv:2602.15778cs.CL2026-02中稿 · *SEM 2026被引 1

用个性化提示提升文本评估精度,兼顾速度与人工判断一致性。

*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation

  • 基于困惑度设计无需生成的判断机制,降低计算开销。
  • 个性化提示使评估结果与人工评分相关性显著提高。
  • 适合需要高效、精准文本质量评估的研究者和开发者。

自动文本质量评估常依赖大模型作为评判者(LLM-judge),但此类方法计算成本高且需后处理。为此,我们基于ParaPLUIE——一种基于困惑度、无需生成文本即可估计“是/否”答案置信度的指标——提出*-PLUIE,即针对特定任务的提示变体,并评估其与人工判断的一致性。实验表明,个性化*-PLUIE在保持低计算成本的同时,相比原有方法与人工评分的相关性更强。

原文摘要 · Abstract (English)

Evaluating the quality of automatically generated text often relies on LLM-as-a-judge (LLM-judge) methods. While effective, these approaches are computationally expensive and require post-processing. To address these limitations, we build upon ParaPLUIE, a perplexity-based LLM-judge metric that estimates confidence over ``Yes/No'' answers without generating text. We introduce *-PLUIE, task specific prompting variants of ParaPLUIE and evaluate their alignment with human judgement. Our experiments show that personalised *-PLUIE achieves stronger correlations with human ratings while maintaining low computational cost.

文本评估LLM-judge个性化提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。