为大模型设计可量化的失语症测试基准,评估语言缺陷。
The Text Aphasia Battery (TAB): A Clinically-Grounded Benchmark for Aphasia-Like Deficits in Language Models
- 基于临床失语测试改编文本评测框架,四类子测试
- 自动化评估与人工评分一致性达0.255(加权卡帕系数)
- 适合研究大模型语言缺陷的科研人员使用
大型语言模型(LLMs)已成为研究人类语言的计算基础的理想‘模型生物’,为探索失语等语言障碍提供了前所未有的机遇。然而,传统临床评估不适用于LLMs,因其预设人类特有的语用压力,并探测非人工架构固有的认知过程。本文提出文本失语症检测量表(Text Aphasia Battery, TAB),该量表改编自快速失语症量表(QAB),专用于评估大模型的失语样缺陷。TAB包含四个子测试:连贯文本生成、词汇理解、句子理解与复述。本文详细阐述了TAB的设计、子测试内容及评分标准。为支持大规模应用,我们验证了一种基于Gemini 2.5 Flash的自动化评估协议,其可靠性与专家人工评分相当(模型-共识一致性的加权卡帕系数为0.255,人-人一致性为0.286)。我们公开发布TAB,作为临床根基明确、可扩展的分析框架,用于研究人工系统中的语言缺陷。
原文摘要 · Abstract (English)
Large language models (LLMs) have emerged as a candidate "model organism" for human language, offering an unprecedented opportunity to study the computational basis of linguistic disorders like aphasia. However, traditional clinical assessments are ill-suited for LLMs, as they presuppose human-like pragmatic pressures and probe cognitive processes not inherent to artificial architectures. We introduce the Text Aphasia Battery (TAB), a text-only benchmark adapted from the Quick Aphasia Battery (QAB) to assess aphasic-like deficits in LLMs. The TAB comprises four subtests: Connected Text, Word Comprehension, Sentence Comprehension, and Repetition. This paper details the TAB's design, subtests, and scoring criteria. To facilitate large-scale use, we validate an automated evaluation protocol using Gemini 2.5 Flash, which achieves reliability comparable to expert human raters (prevalence-weighted Cohen's kappa = 0.255 for model--consensus agreement vs. 0.286 for human--human agreement). We release TAB as a clinically-grounded, scalable framework for analyzing language deficits in artificial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。