arXiv:2606.15696cs.AIcs.CL2026-06

大模型可辅助识别失语症言语中的有效信息单元,提升评估效率。

Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?

  • 用少量示例提示让大模型进行逐词信息单元分类。
  • 三种模型少样本准确率F1达0.776至0.817,表现接近人工标注。
  • 适合用于人机协同的失语症语言评估系统,非完全自动化。

正确信息单元(CIUs)是失语症话语评估的核心,能衡量交流的信息量而非仅语言形式。但其标注耗时且需专业评分员。本研究检验指令微调的大语言模型(LLMs)是否能从失语症话语转录文本中可靠完成逐词CIU分类。使用猫救援刺激任务获取16段图片描述文本,按Nicholas和Brookshire(1993)标准标注CIU状态,涵盖控制组、轻度、中度和重度失语症四类严重程度。对比四种公开可用的指令微调LLM在零样本与两种少样本提示下的表现,采用五次分层随机种子测试。以人类共识标签为基准,评估准确率、精确率、召回率、F1和Cohen's kappa。零样本表现不足;少样本提示显著提升性能,三款模型表现良好:Llama-3.1-8B、Qwen2.5-7B、Mistral-7B平均F1在0.776至0.817之间,固定全局与按块局部示例选择无显著差异。Phi-3-mini表现不稳定,不可靠。优秀模型召回率高但精确率偏低,显示对非信息单元存在系统性误判。性能随话语严重程度变化,重度失语症下最差。少样本提示可在不进行梯度训练的前提下支持自动化CIU识别,但与人工标注一致性仍不足,无法实现完全自主。研究支持将大模型用于人机协同的话语评估系统。

原文摘要 · Abstract (English)

Correct Information Units (CIUs) are central to discourse assessment in aphasia because they quantify communicative informativeness rather than linguistic form alone. However, CIU scoring is time intensive and requires trained raters. This study examined whether instruction-tuned large language models (LLMs) can reliably perform token-level CIU classification from aphasic discourse transcripts. Sixteen picture-description transcripts elicited with the Cat Rescue stimulus were annotated for CIU status according to Nicholas and Brookshire (1993). The sample spanned four severity strata: control, mild, moderate, and severe aphasia. Four publicly available instruction-tuned LLMs were benchmarked under zero-shot and two few-shot prompting conditions across five stratified random seeds. Performance was evaluated against consensus human labels using accuracy, precision, recall, F1, and Cohen's kappa. Zero-shot prompting was insufficient across models. In contrast, few-shot prompting yielded substantial gains and produced competitive performance for three viable models. Mean few-shot F1 scores ranged from 0.776 to 0.817 across Llama-3.1-8B, Qwen2.5-7B, and Mistral-7B, with no significant differences between fixed global and per-chunk local example selection. Phi-3-mini was unstable and did not yield reliable performance. Viable models showed high recall but lower precision, suggesting systematic over-classification of tokens as CIUs. Performance also varied by discourse severity, with the weakest results in more severe aphasia. Few-shot LLM prompting can support automated CIU identification without gradient-based task training, but agreement with human annotation remains insufficient for fully autonomous use. These findings support LLM-based CIU scoring as a promising human-in-the-loop component of discourse assessment systems.

失语症大模型信息单元人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。