构建德语文本生成体验质量评估数据集,提升模型评价与人类感知一致性。
Assessing Quality of Experience in Natural Language Generation of German Text
- 基于人类反馈构建德语文本生成的体验质量评估数据集。
- 混合模型在预测人类评分上优于纯Transformer模型。
- 适合关注多维度生成质量评估的研究者和开发者。
自然语言生成(NLG)技术的快速发展使得可靠评估生成文本质量变得愈发重要,尤其在大型语言模型(LLMs)广泛应用于实际场景的背景下。然而,传统自动评估指标难以捕捉人类对生成质量的多维度感知。本文提出TextQ-German,一个面向德语NLG任务(包括自动摘要和机器翻译)的人类中心化质量体验(QoE)评估数据集。通过面向德语母语者的众包实验,收集了人类质量评分,并识别出各任务的关键感知维度。我们开发了基于Transformer、语言学特征及混合方法的自动QoE预测模型。混合模型在几乎所有实验设置中均优于纯Transformer基线,而仅使用语言学特征的方法也能接近微调语言模型的表现。数据集进一步扩展了由大语言模型生成的输出,并标注了总体QoE分数。在预留测试集上的最终验证表明模型具备对未见数据的泛化能力。本工作贡献了一个公开可获取的资源与自动QoE预测基准,为构建更符合人类感知的NLG系统奠定了基础。
原文摘要 · Abstract (English)
The rapid advancement of Natural Language Generation (NLG) has made the reliable evaluation of generated text increasingly critical, as these systems, such as large language models (LLMs), are now widely deployed in real-world applications. However, traditional automatic metrics fail to capture the multifaceted nature of perceived quality. In this paper, we introduce TextQ-German, a novel dataset suite for human-centered evaluation of German NLG from a Quality of Experience (QoE) perspective, covering automatic text summarization and machine translation. Through crowdsourcing studies with German speakers, we collect human quality ratings and identify relevant perceptual quality dimensions for each task. We develop automatic QoE prediction models, including transformer-based, linguistic feature-based, and hybrid approaches. Hybrid models outperform pure transformer baselines in almost all experimental settings, while linguistic features alone can approach the performance of fine-tuned language models. The dataset is extended with LLM-generated outputs annotated with overall QoE scores. Final validation on held-out sets indicates generalization to unseen data. Our work contributes a publicly accessible resource for NLG evaluation and baselines for automatic QoE prediction, providing a foundation for developing NLG systems that better align with human quality perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。