arXiv:2412.17669cs.CL2024-12被引 1

用大模型补全布罗卡失语症患者的不连贯句子,提升语言康复辅助工具效果。

Generating Completions for Broca's Aphasic Sentences Using Large Language Models

  • 基于规则生成模拟失语症语料,微调大模型补全语法缺失句。
  • 模型在更长输入上表现更好,合成与真实数据均验证有效。
  • 适合语言康复、临床辅助系统研究者参考。

布罗卡失语症以非流利、费力、语法缺失的言语表达为特征,但理解能力相对保留。传统治疗耗时耗力,且难以模拟真实对话场景。本文探索使用序列到序列的大语言模型(LLMs)完成布罗卡失语症患者的句子。首先,通过基于规则的系统生成模拟失语症语料,模仿其语言特征;随后,在无真实失语样本的情况下,对四个预训练大模型进行微调,任务是补全语法缺失的句子。在合成数据和真实失语症数据上评估微调模型性能。结果表明,大模型具备重建语法缺失句子的能力,且在输入句长增加时表现更优。研究展示了大模型在提升失语症患者沟通辅助工具方面的潜力,或可推广至其他临床人群。

原文摘要 · Abstract (English)

Broca's aphasia is a type of aphasia characterized by non-fluent, effortful and agrammatic speech production with relatively good comprehension. Since traditional aphasia treatment methods are often time-consuming, labour-intensive, and do not reflect real-world conversations, applying natural language processing based approaches such as Large Language Models (LLMs) could potentially contribute to improving existing treatment approaches. To address this issue, we explore the use of sequence-to-sequence LLMs for completing Broca's aphasic sentences. We first generate synthetic Broca's aphasic data using a rule-based system designed to mirror the linguistic characteristics of Broca's aphasic speech. Using this synthetic data (without authentic aphasic samples), we then fine-tune four pre-trained LLMs on the task of completing agrammatic sentences. We evaluate our fine-tuned models on both synthetic and authentic Broca's aphasic data. We demonstrate LLMs' capability for reconstructing agrammatic sentences, with the models showing improved performance with longer input utterances. Our result highlights the LLMs' potential in advancing communication aids for individuals with Broca's aphasia and possibly other clinical populations.

大模型失语症语言康复句子补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。