arXiv:2409.19737cs.CL2024-09综述被引 13

系统梳理240篇论文,揭示NLP在痴呆研究中的应用与空白。

A Systematic Review of NLP for Dementia -- Tasks, Datasets and Opportunities

  • 综述240篇文献,聚焦痴呆检测、语言生物标志物等四类任务
  • 半数研究仅用临床数据做痴呆检测,其他方向仍待开拓
  • 呼吁加强跨领域合作,关注伦理与模型可信度

语言与认知衰退密切相关,推动了自然语言处理(NLP)与医学界在痴呆研究中的长期合作。本文系统回顾了超过240篇将NLP应用于痴呆相关研究的论文,涵盖医学、技术及NLP领域的文献。识别出关键研究方向:痴呆检测、语言生物标志物提取、照护者支持和患者辅助。其中,一半论文仅专注于使用临床数据进行痴呆检测。然而,人工退化语言模型、合成数据、数字孪生等方向仍待探索。文章指出信任、科学严谨性、可应用性及跨社区协作方面的缺口,并揭示了多样化的数据来源:录音、书面文本、结构化数据、自发语言、合成数据、临床记录、社交媒体数据等。本综述旨在激发更具创新性、影响力与严谨性的痴呆NLP研究。

原文摘要 · Abstract (English)

The close link between cognitive decline and language has fostered long-standing collaboration between the NLP and medical communities in dementia research. To examine this, we reviewed over 240 papers applying NLP to dementia-related efforts, drawing from medical, technological, and NLP-focused literature. We identify key research areas, including dementia detection, linguistic biomarker extraction, caregiver support, and patient assistance, showing that half of all papers focus solely on dementia detection using clinical data. Yet, many directions remain unexplored -- artificially degraded language models, synthetic data, digital twins, and more. We highlight gaps and opportunities around trust, scientific rigor, applicability and cross-community collaboration. We raise ethical dilemmas in the field, and highlight the diverse datasets encountered throughout our review -- recorded, written, structured, spontaneous, synthetic, clinical, social media-based, and more. This review aims to inspire more creative, impactful, and rigorous research on NLP for dementia.

痴呆研究NLP应用数据多样性跨学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。