用用户体验评估大模型,比单纯看回答对错更真实。
QoNext: Towards Next-generation QoE for Foundation Models
- 引入网络多媒体的体验质量理念,量化用户交互感受。
- 通过可控实验收集人类评分,建立可预测体验的数据库。
- 能指导产品优化,让大模型服务更符合真实用户需求。
现有大模型评估方法,包括近期以人类为中心的方案,未能捕捉交互中真正重要的因素:用户实际体验。当前方法仅关注输出正确性,忽视了响应质量与交互过程共同作用带来的用户满意度,限制了对用户体验机制的理解。为此,我们提出 QoNext,首个将网络与多媒体领域的用户体验(QoE)原则应用于大模型评估的框架。QoNext 识别影响用户体验的体验因素,并在受控实验中收集人类评分,涵盖不同配置。基于这些研究,构建了一个面向用户体验的数据库,并训练出可从可测量系统参数预测感知体验的模型。结果表明,QoNext 不仅支持主动、细粒度的评估,还能为实际产品化服务中优化大模型提供可操作指导。
原文摘要 · Abstract (English)
Existing evaluations of foundation models, including recent human-centric approaches, fail to capture what truly matters: user's experience during interaction. Current methods treat evaluation as a matter of output correctness alone, overlooking that user satisfaction emerges from the interplay between response quality and interaction, which limits their ability to account for the mechanisms underlying user experience. To address this gap, we introduce QoNext, the first framework that adapts Quality of Experience (QoE) principles from networking and multimedia to the assessment of foundation models. QoNext identifies experiential factors that shape user experience and incorporates them into controlled experiments, where human ratings are collected under varied configurations. From these studies we construct a QoE-oriented database and train predictive models that estimate perceived user experience from measurable system parameters. Our results demonstrate that QoNext not only enables proactive and fine-grained evaluation but also provides actionable guidance for productized services of optimizing foundation models in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。