arXiv:2510.24817cs.CLcs.AI2025-10被引 2

用合成数据模拟失语症患者语言特征,解决标注数据稀缺问题。

Towards a Method for Synthetic Generation of Persons with Aphasia Transcripts

  • 通过删词、插入填充词、错语替换生成四等级失语症语音文本。
  • Mistral 7b Instruct 生成的文本在词数、词长等指标上最接近真实失语症表现。
  • 适合失语症研究与自然语言处理中缺乏真实数据的场景使用。

在失语症研究中,言语语言病理学家需耗费大量时间手动标注语音样本的正确信息单元(CIUs),衡量其信息量。然而,自动化识别失语语言的能力受限于数据稀缺——如AphasiaBank仅约600份转录文本,远低于大语言模型(LLMs)训练所需的数十亿词元。为此,本研究构建并验证了两种生成AphasiaBank「猫救援」图片描述任务合成转录文本的方法:一种基于过程式编程,另一种利用Mistral 7b Instruct与Llama 3.1 8b Instruct LLMs。方法通过删词、插入填充词及错语替换,在轻度、中度、重度、极重度四个严重程度层级生成文本。结果表明,相较于真人转录,Mistral 7b Instruct生成的文本在非词密度(NDW)、词数和词长方面展现出更真实的语言退化趋势。未来工作应构建更大规模数据集,微调模型以更好表征失语特征,并由言语语言病理学家评估合成文本的真实性和实用性。

原文摘要 · Abstract (English)

In aphasia research, Speech-Language Pathologists (SLPs) devote extensive time to manually coding speech samples using Correct Information Units (CIUs), a measure of how informative an individual sample of speech is. Developing automated systems to recognize aphasic language is limited by data scarcity. For example, only about 600 transcripts are available in AphasiaBank yet billions of tokens are used to train large language models (LLMs). In the broader field of machine learning (ML), researchers increasingly turn to synthetic data when such are sparse. Therefore, this study constructs and validates two methods to generate synthetic transcripts of the AphasiaBank Cat Rescue picture description task. One method leverages a procedural programming approach while the second uses Mistral 7b Instruct and Llama 3.1 8b Instruct LLMs. The methods generate transcripts across four severity levels (Mild, Moderate, Severe, Very Severe) through word dropping, filler insertion, and paraphasia substitution. Overall, we found, compared to human-elicited transcripts, Mistral 7b Instruct best captures key aspects of linguistic degradation observed in aphasia, showing realistic directional changes in NDW, word count, and word length amongst the synthetic generation methods. Based on the results, future work should plan to create a larger dataset, fine-tune models for better aphasic representation, and have SLPs assess the realism and usefulness of the synthetic transcripts.

失语症合成数据LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。