德国临床文本数据难获取,研究者用翻译、合成和领域代理数据替代。
Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data
- 用翻译英文数据、生成虚构内容及医疗指南等作为替代数据源。
- 共发现71个独特语料库,其中46个真实、5个翻译、6个合成。
- 适合关注德国医疗NLP但无法获取原始数据的研究者参考。
本文综述德国临床文档语料库,因德国严格的数据隐私法规,多数真实临床数据被锁在医院内部系统中,外部研究者难以访问。为应对这一数据困境,研究者探索了多种替代方案:包括英文学术语料的机器翻译、虚构临床内容的合成生成,以及各类领域代理数据。后者包括医学期刊、治疗指南、药品说明书等近端代理,还有百科类医学文章或社交媒体上的医疗内容等远端代理。通过符合PRISM标准的文献检索,从4个数据库中筛选出362条相关文献,最终确定78篇相关论文,涵盖92种语料库版本,其中71个为独立语料库。统计显示,真实临床语料库46个,翻译语料库5个,合成语料库6个;近端代理18个,远端代理17个。尽管公开替代数据数量众多,但其体裁风格、术语使用与真实临床文档仍存在显著差异,因此其有效性仍需审慎评估。
原文摘要 · Abstract (English)
We survey clinical document corpora, with focus on German textual data. Due to rigid data privacy legislation in Germany these resources, with only few exceptions, are stored in safe clinical data spaces and locked against clinic-external researchers. This situation stands in stark contrast with established workflows in the field of natural language processing where easy accessibility and reuse of data collections are common practice. Hence, alternative corpus designs have been examined to escape from this data poverty. Besides machine translation of English clinical datasets and the generation of synthetic corpora with fictitious clinical contents, several other types of domain proxies have come up as substitutes for clinical documents. Common instances of close proxies are medical journal publications, therapy guidelines, drug labels, etc., more distant proxies include online encyclopedic medical articles or medical contents from social media channels. After PRISM-conformant identification of 362 hits from 4 bibliographic systems, 78 relevant documents were finally selected for this review. They contained overall 92 different published versions of corpora from which 71 were truly unique in terms of their underlying document sets. Out of these, the majority were clinical corpora -- 46 real ones, 5 translated ones, and 6 synthetic ones. As to domain proxies, we identified 18 close and 17 distant ones. There is a clear divide between the large number of non-accessible authentic clinical German-language corpora and their publicly accessible substitutes: translated or synthetic, close or more distant proxies. So on first sight, the data bottleneck seems broken. Yet differences in genre-specific writing style, wording and medical domain expertise in this typological space are also obvious. This raises the question how valid alternative corpus designs really are.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。