首个专用于德语简化的评估指标,提升简化质量判断准确性。
DETECT: Determining Ease and Textual Clarity of German Text Simplifications
- 基于大模型生成合成评分,无需人工标注数据
- 在语义保留和流畅性上与人工评价相关性显著更高
- 适合关注德语可读性、自动评估的研究者
当前德语自动文本简化(ATS)评估依赖SARI、BLEU、BERTScore等通用指标,难以全面衡量简化程度、语义保留和流畅性。尽管英语已有LENS等专用指标,但德语因缺乏人工标注语料而进展滞后。为此,我们提出DETECT,首个面向德语的综合评估指标,涵盖简化度、语义保留和流畅性三维度,完全基于大语言模型(LLM)生成的数据训练。方法上,我们适配了LENS框架,引入(i)通过LLM生成合成评分的流水线,实现无标注数据集构建;(ii)基于LLM的优化步骤,使评分标准契合简化需求。据我们所知,还构建了目前最大的德语文本简化人工评价数据集以直接验证指标。实验表明,DETECT在与人工判断的相关性上显著优于现有指标,尤其在语义保留和流畅性方面提升明显。研究还揭示了大模型在自动评估中的潜力与局限,为语言可及性任务提供可迁移指导。
原文摘要 · Abstract (English)
Current evaluation of German automatic text simplification (ATS) relies on general-purpose metrics such as SARI, BLEU, and BERTScore, which insufficiently capture simplification quality in terms of simplicity, meaning preservation, and fluency. While specialized metrics like LENS have been developed for English, corresponding efforts for German have lagged behind due to the absence of human-annotated corpora. To close this gap, we introduce DETECT, the first German-specific metric that holistically evaluates ATS quality across all three dimensions of simplicity, meaning preservation, and fluency, and is trained entirely on synthetic large language model (LLM) responses. Our approach adapts the LENS framework to German and extends it with (i) a pipeline for generating synthetic quality scores via LLMs, enabling dataset creation without human annotation, and (ii) an LLM-based refinement step for aligning grading criteria with simplification requirements. To the best of our knowledge, we also construct the largest German human evaluation dataset for text simplification to validate our metric directly. Experimental results show that DETECT achieves substantially higher correlations with human judgments than widely used ATS metrics, with particularly strong gains in meaning preservation and fluency. Beyond ATS, our findings highlight both the potential and the limitations of LLMs for automatic evaluation and provide transferable guidelines for general language accessibility tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。