首个符合欧洲标准的易读文本生成数据集,助力认知障碍者获取信息。
Inclusive Easy-to-Read Generation for Individuals with Cognitive Impairments
- 构建首个符合欧洲标准的易读文本生成数据集ETR-fr。
- 小样本微调下预训练模型在跨领域文本上表现接近大模型。
- 结合自动评估与人工问卷,确保输出可读性与可用性。
保障认知障碍者的信息可及性对自主权、自我决定和完整公民身份至关重要。然而,手动制作易读文本(ETR)耗时、成本高且难以扩展,限制了医疗、教育和公共生活中的关键信息获取。人工智能驱动的ETR生成提供了可扩展的解决方案,但面临数据稀缺、领域适应和大语言模型轻量化学习等挑战。本文提出ETR-fr,首个完全符合欧洲ETR指南的易读文本生成数据集。我们在预训练模型(PLMs)和大语言模型(LLMs)上实现参数高效微调,建立生成基线。为确保输出质量与可访问性,引入基于自动指标与人工评估相结合的评测框架,后者采用36题评估问卷,与指南一致。总体结果表明,PLMs在性能上可媲美LLMs,且对外部领域文本具有良好的适应能力。
原文摘要 · Abstract (English)
Ensuring accessibility for individuals with cognitive impairments is essential for autonomy, self-determination, and full citizenship. However, manual Easy-to-Read (ETR) text adaptations are slow, costly, and difficult to scale, limiting access to crucial information in healthcare, education, and civic life. AI-driven ETR generation offers a scalable solution but faces key challenges, including dataset scarcity, domain adaptation, and balancing lightweight learning of Large Language Models (LLMs). In this paper, we introduce ETR-fr, the first dataset for ETR text generation fully compliant with European ETR guidelines. We implement parameter-efficient fine-tuning on PLMs and LLMs to establish generative baselines. To ensure high-quality and accessible outputs, we introduce an evaluation framework based on automatic metrics supplemented by human assessments. The latter is conducted using a 36-question evaluation form that is aligned with the guidelines. Overall results show that PLMs perform comparably to LLMs and adapt effectively to out-of-domain texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。