用智能对话框架主动发现自闭症语言障碍特征,提升评估效率。
A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism

- 设计多代理系统,医生代理先思考再选问题,主动触发潜在语言异常。
- 在35名患者484段对话中实现82.1%的障碍特征覆盖,比人工复现高16.6%。
- 适合临床辅助筛查系统研发者及自闭症语言评估研究者参考。
自闭症谱系障碍中的社会语言障碍(SLD)特征,如重复话语、代词错用和刻板引用媒体内容,在自然对话中常隐匿,仅在特定对话条件下显现。在结构化临床评估中,提问策略的选择是决定诊断信息量的关键但被忽视的因素。大型语言模型能否主动选择策略以系统性暴露这些隐藏特征,尚无深入探索。本文提出TPA(Think, Plan, Ask)——一种应用于孤独症诊断观察量表第4模块(ADOS-2)语言评估的主动式多代理对话框架。其中,医生代理在生成问题前显式推理尚未观测到的特征,并据此选择临床依据充分的策略。患者代理基于真实ADOS-2临床数据构建,实现无需真人参与的可复现评估,经三项独立实验验证具备良好真实性。在35名患者的484个对话片段上,TPA优于六种对比基线,在所有主指标上表现更优,达到82.1%的SLD特征覆盖率,较训练有素临床医生人工复现的65.5%高出16.6%,且每轮诊断效率显著提升(AUCC:0.628 vs. 0.458,绝对提升+0.170)。结果表明,主动提问策略选择能大幅提升自动化SLD特征评估效率,对可扩展的AI辅助临床筛查具有直接意义。
原文摘要 · Abstract (English)
Characteristic linguistic behaviors associated with Social Language Disorder (SLD) in autism spectrum disorder, including echoic repetition, pronoun displacement, and stereotyped media quoting, are largely absent from spontaneous conversation and only emerge under specific conversational conditions. In structured clinical assessments, this latency means that questioning strategy selection is a critical yet underappreciated determinant of how much diagnostic information a conversation yields. Whether large language models (LLMs) can be guided to proactively select questioning strategies that systematically surface these latent traits remains largely unexplored. Here we present TPA (Think, Plan, Ask), a proactive multi-agent dialogue framework applied to the language assessment component of the Autism Diagnostic Observation Schedule Module 4 (ADOS-2), in which a doctor agent explicitly reasons about which traits remain unobserved before selecting a clinically grounded strategy and generating a targeted question. A patient agent grounded in real ADOS-2 clinical data enables reproducible evaluation without real patient participation, validated across three independent experiments confirming adequate fidelity to real patient language. Evaluated on 484 episodes from 35 patients, TPA outperforms six competitive dialogue planning baselines across all primary metrics, achieving 82.1% SLD trait coverage, 16.6% higher than automated replay of real clinical dialogues conducted by trained clinicians (65.5%), with substantially greater per-turn diagnostic efficiency (AUCC: 0.628 vs. 0.458, absolute gain +0.170). These results demonstrate that proactive questioning strategy selection substantially improves the efficiency of automated SLD trait assessment, with direct implications for scalable AI-assisted clinical screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。