用隐空间几何特性实现低成本高精度对话模型满意度评估
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
- 基于隐空间极化指数自动生成评分标准
- 在工业数据集上优于生成式基线,且推理成本降几个数量级
- 适合需要高频迭代的对话系统开发团队使用
准确评估用户满意度对对话AI的迭代至关重要。然而,对于开放性助手,传统A/B测试缺乏可靠指标:显式反馈稀疏,隐式指标模糊。为此,我们提出BoRP(Bootstrapped Regression Probing),一种可扩展的高保真满意度评估框架。不同于生成式方法,BoRP利用大语言模型隐空间的几何特性,采用基于极化指数的自举机制自动化生成评分规则,并使用偏最小二乘法(PLS)将隐藏状态映射为连续得分。在工业数据集上的实验表明,BoRP(Qwen3-8B/14B)显著优于生成式基线(甚至超过Qwen3-Max),且推理成本降低数个数量级,支持全规模监控与基于CUPED的高敏感度A/B测试。
原文摘要 · Abstract (English)
Accurate evaluation of user satisfaction is critical for iterative development of conversational AI. However, for open-ended assistants, traditional A/B testing lacks reliable metrics: explicit feedback is sparse, while implicit metrics are ambiguous. To bridge this gap, we introduce BoRP (Bootstrapped Regression Probing), a scalable framework for high-fidelity satisfaction evaluation. Unlike generative approaches, BoRP leverages the geometric properties of LLM latent space. It employs a polarization-index-based bootstrapping mechanism to automate rubric generation and utilizes Partial Least Squares (PLS) to map hidden states to continuous scores. Experiments on industrial datasets show that BoRP (Qwen3-8B/14B) significantly outperforms generative baselines (even Qwen3-Max) in alignment with human judgments. Furthermore, BoRP reduces inference costs by orders of magnitude, enabling full-scale monitoring and highly sensitive A/B testing via CUPED.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。