为提升德语可持续报告可读性,构建句子级可读性评分体系
Towards Empowering Consumers through Sentence-level Readability Scoring in German ESG Reports
- 基于众包标注扩展德语文本可读性数据集
- 微调小规模Transformer模型预测人类可读性误差最低
- 大模型提示虽有潜力但需权衡速度与精度
随着经济与社会对可持续性的日益关注,信息量激增,消费者亟需可靠的信息获取渠道。为此,企业开始自愿或依法发布环境、社会与治理(ESG)报告。为服务公众,这些报告不仅应面向金融专家,也需对非专业人士清晰易懂。但其表述是否足够清晰?本文在现有德语ESG报告句子级数据集基础上,引入众包可读性标注。结果显示,母语者普遍认为报告句子易读,但可读性具主观性。我们评估了多种可读性评分方法在预测误差和与人工评分相关性上的表现。分析表明,尽管大语言模型提示具备区分清晰与难读句子的潜力,但小型微调的Transformer模型在预测人类可读性时误差最小。多个模型预测结果平均可小幅提升性能,但会降低推理速度。
原文摘要 · Abstract (English)
With the ever-growing urgency of sustainability in the economy and society, and the massive stream of information that comes with it, consumers need reliable access to that information. To address this need, companies began publishing so called Environmental, Social, and Governance (ESG) reports, both voluntarily and forced by law. To serve the public, these reports must be addressed not only to financial experts but also to non-expert audiences. But are they written clearly enough? In this work, we extend an existing sentence-level dataset of German ESG reports with crowdsourced readability annotations. We find that, in general, native speakers perceive sentences in ESG reports as easy to read, but also that readability is subjective. We apply various readability scoring methods and evaluate them regarding their prediction error and correlation with human rankings. Our analysis shows that, while LLM prompting has potential for distinguishing clear from hard-to-read sentences, a small finetuned transformer predicts human readability with the lowest error. Averaging predictions of multiple models can slightly improve the performance at the cost of slower inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。