arXiv:2509.25903cs.CLcs.AI2025-09

提出高效评估文本个性化质量的方法PerQ,降低大模型评测成本。

PerQ: Efficient Evaluation of Multilingual Text Personalization Quality

  • 用多个大模型融合评估文本个性化质量,避免单个模型偏差。
  • 相比传统方法,减少计算资源消耗,提升评测效率。
  • 适合对比大小模型生成能力的研究者使用。

由于缺乏评估文本特定方面(如个性化质量)的指标,研究者通常依赖大型语言模型进行元评估。然而,单个语言模型存在内部偏差,因此推荐使用多个模型联合评估,这直接增加了评测成本。本文提出一种计算高效的文本个性化质量评估方法——PerQ。通过案例研究对比大、小语言模型的生成能力,验证了该度量在实际研究中的可用性,有效减少了资源浪费。

原文摘要 · Abstract (English)

Since no metrics are available to evaluate specific aspects of a text, such as its personalization quality, the researchers often rely solely on large language models to meta-evaluate such texts. Due to internal biases of individual language models, it is recommended to use multiple of them for combined evaluation, which directly increases costs of such meta-evaluation. In this paper, a computationally efficient method for evaluation of personalization quality of a given text (generated by a language model) is introduced, called PerQ. A case study of comparison of generation capabilities of large and small language models shows the usability of the proposed metric in research, effectively reducing the waste of resources.

文本评估个性化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。