arXiv:2503.17755cs.CLcs.LG2025-03ACL被引 10

用分类探针挖掘大模型隐含知识,提升文本评价准确性。

Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying Probes

  • 通过对比提示差异训练线性分类探针,直接读取模型隐知识。
  • 在6个数据集上优于生成式判断,且计算开销相近。
  • 方法可解释性强,适合需要透明评判的场景。

大型语言模型常被用作自动文本评价工具,但其效果易受无意偏差影响。本文提出利用线性分类探针,通过对比提示对之间的差异进行训练,直接访问模型的隐含知识,以提取更准确的偏好判断。在四个不同模型家族、多种规模的模型及六个涵盖文本质量评估与常识推理的多样化数据集上开展广泛实验,结果表明,无论是监督还是无监督的探针方法,均持续优于传统的生成式判断,且计算成本相近。该方法在领域迁移下仍具泛化能力,甚至超越使用相同训练数据量的微调评价器。结果表明,线性探针为大模型作为评判者任务提供了一种准确、鲁棒且计算高效的方案,并能揭示模型编码判断相关知识的方式。数据与代码将未来公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are often used as automated judges to evaluate text, but their effectiveness can be hindered by various unintentional biases. We propose using linear classifying probes, trained by leveraging differences between contrasting pairs of prompts, to directly access LLMs' latent knowledge and extract more accurate preferences. Through extensive experiments using models of varying size from four different families and six diverse datasets assessing text quality evaluation and common sense reasoning, we demonstrate that both supervised and unsupervised probing approaches consistently outperform traditional generation-based judgement while maintaining similar computational costs. These probes generalise under domain shifts and can even outperform finetuned evaluators with the same training data size. Our results suggest linear probing offers an accurate, robust and computationally efficient approach for LLM-as-judge tasks while providing interpretable insights into how models encode judgement-relevant knowledge. Our data and code will be openly released in the future.

大模型评判探针分析偏好提取可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。