arXiv:2608.01158cs.CL2026-08中稿 · KONVENS 2026

构建多层级医学文本语料库,助力医疗信息精准传达。

PlainMedScale: A Corpus of Multi-Level Simplified Medical Texts in German and English

  • 按沟通功能分四层构建德英双语医学语料库
  • 发现现有可读性指标无法跨层级通用
  • 揭示大模型仍保留原文难度,适配健康传播研究

我们提出PlainMedScale,一个涵盖德语和英语的多层级医学语料库,包含四个可读性层级:参考、解释、决策支持和获取。该语料库源自MSD(专业与大众)、Gesund.Bund、Apotheken Umschau Einfache Sprache及NHS。四个层级对应不同沟通功能,突破了以往语料库中专家-大众二元对立的局限。通过两个试点研究,我们发现许多在两层间建立的可读性指标无法推广至完整梯度;同时,即使使用SOTA开源大模型进行简明语言转换,仍部分保留输入文本的难度。代码(https://github.com/GS-Uni-Heidelberg/PlainMedScale)与数据(https://doi.org/10.5281/zenodo.21728290)已公开。

原文摘要 · Abstract (English)

We introduce PlainMedScale, a topic-aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD (professional and consumer), Gesund.Bund, Apotheken Umschau Einfache Sprache, and the NHS. The four tiers correspond to distinct communicative functions --- reference, explanation, decision support, and access --- and move beyond the binary expert--lay contrast of prior corpora. In two pilot studies enabled by the alignments, we show that many readability metrics established on two registers fail to generalize across the full gradient, and that a SOTA open-weight LLM prompted for Plain Language still partially preserves the difficulty of its input. Code (https://github.com/GS-Uni-Heidelberg/PlainMedScale) and data (https://doi.org/10.5281/zenodo.21728290) are made available.

医疗文本可读性多语言语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。