arXiv:2605.11091cs.LGcs.AI2026-05

首个四轴多年龄自闭症筛查模型评测基准,揭示不同年龄段诊断难点。

ASD-Bench: A Four-Axis Comprehensive Benchmark of AI Models for Autism Spectrum Disorder

论文配图:ASD-Bench: A Four-Axis Comprehensive Benchmark of AI Models for Autism Spectrum Disorder
图 1 · 摘自论文原文
  • 构建跨儿童、青少年、成人的四轴评测框架,覆盖性能、校准、可解释性与抗干扰性。
  • 青少年组准确率上限仅0.837,显著低于儿童组的0.915,提示早期干预窗口差异。
  • 发现社交动机、模式识别等特征重要性随年龄动态变化,适合临床部署参考。

自动化自闭症谱系障碍(ASD)筛查工具受限于单一架构评估、轴向局限及成人群体主导,难以揭示关键的年龄特异性诊断模式。本文提出ASD-Bench,一个系统性的表格型基准,评估机器学习、深度学习与基础模型在三个年龄组(1-11岁儿童、12-16岁青少年、17-64岁成人)上的表现,涵盖预测性能、校准度、可解释性与对抗鲁棒性四方面。基于4,068条AQ-10问卷数据,评测包括XGBoost、AdaBoost、随机森林、逻辑回归、MLP、TabNet、TabTransformer、FT-Transformer及TabPFN v2等模型。引入启发式综合惩罚(HAP)指标,对假阴性加重惩罚并考虑交叉验证方差以保障部署稳定性。成人组表现优异(10/17模型达到完美F1与AUC),青少年组更具挑战(F1上限0.837,儿童组为0.915)。特征重要性随年龄演变:儿童以A9(社交动机)为主导,青少年转向A5(模式识别),成人则呈现扁平化分布,符合发育期社会掩饰现象。准确率与校准度分离明显:如AdaBoost在成人组获F1=1.000但ECE=0.302,表明单一指标评价不足。研究提供分年龄部署建议。所有结果均基于问卷标签,属概念验证,非临床诊断认证。

原文摘要 · Abstract (English)

Automated ASD screening tools remain limited by single-architecture evaluations, axis-restricted assessment, and near-exclusive focus on adult cohorts, obscuring age-specific diagnostic patterns critical for early intervention. We introduce ASD-Bench, a systematic tabular benchmark evaluating ML, deep learning, and foundation model configurations across three age cohorts (children 1-11 yr, adolescents 12-16 yr, adults 17-64 yr) on four axes: predictive performance, calibration, interpretability, and adversarial robustness. Applied to a curated v3 dataset of 4,068 AQ-10 records, our benchmark spans classical models (XGBoost, AdaBoost, Random Forest, Logistic Regression), neural networks (MLP), deep tabular transformers (TabNet, TabTransformer, FT-Transformer), and TabPFN v2. We introduce the Heuristic Aggregate Penalty (HAP): a cost-sensitive metric penalising false negatives more heavily and incorporating cross-validation variance for deployment stability. Adult classification yields high performance (10/17 models achieve perfect F1 and AUC), while adolescents present a harder task (F1 ceiling 0.837 vs. 0.915 for children). Feature hierarchies shift across cohorts: A9 (social motivation) dominates for children, A5 (pattern recognition) leads for adolescents, and adults exhibit a flatter importance profile consistent with developmental social masking. Accuracy and calibration are dissociated: AdaBoost achieves F1=1.000 on adults with ECE=0.302, confirming single-metric evaluation is insufficient for clinical AI. Cohort-specific deployment recommendations are provided. All findings should be interpreted as proof-of-concept evidence on questionnaire-derived labels rather than clinically validated diagnostic performance.

自闭症筛查多年龄评测模型可解释性临床AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。