用统计方法同时评估大模型文本的流畅、多样、连贯等多维度质量。
Statistical Multicriteria Evaluation of LLM-Generated Text
- 基于广义随机占优理论,不加权地统一评估多种文本质量指标。
- 能发现人类与生成文本在多个维度上的显著差异。
- 适合关注文本生成质量综合评估的研究者或工程师。
评估大语言模型生成文本的质量仍是自然语言处理中的核心挑战。现有评估方法常依赖单一指标或简单加权,无法捕捉连贯性、多样性、流畅性等指标间的复杂权衡。本文引入基于广义随机占优(GSD)的统计推断框架,解决三大问题:单指标评估不足、自动指标(基数)与人工判断(序数)不兼容、缺乏统计推断保障。GSD-front 方法能同时评估多个质量维度,尊重其不同量纲,基于解码策略的偏序关系,避免人为加权。应用于常见解码策略与人类生成文本的对比,该方法识别出显著性能差异,并考虑了采样设计中非独立同分布的潜在偏差。
原文摘要 · Abstract (English)
Assessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplistic aggregations that fail to capture the nuanced trade-offs between coherence, diversity, fluency, and other relevant indicators of text quality. In this work, we adapt a recently proposed framework for statistical inference based on Generalized Stochastic Dominance (GSD) that addresses three critical limitations in existing benchmarking methodologies: the inadequacy of single-metric evaluation, the incompatibility between cardinal automatic metrics and ordinal human judgments, and the lack of inferential statistical guarantees. The GSD-front approach enables simultaneous evaluation across multiple quality dimensions while respecting their different measurement scales, building upon partial orders of decoding strategies, thus avoiding arbitrary weighting of the involved metrics. By applying this framework to evaluate common decoding strategies against human-generated text, we demonstrate its ability to identify statistically significant performance differences while accounting for potential deviations from the i.i.d. assumption of the sampling design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。