arXiv:2606.24960cs.LGcs.AI2026-06

用不确定性感知融合多专家模型,提升中风康复评估的准确性与可信度。

Enhancing Clinician Decision-Making via Uncertainty-Aware Multi-Expert Fusion for Stroke Rehabilitation

论文配图:Enhancing Clinician Decision-Making via Uncertainty-Aware Multi-Expert Fusion for Stroke Rehabilitation
图 1 · 摘自论文原文
  • 通过动态贝叶斯网络融合692个模态模型,分层级输出带置信度的评估结果
  • 在105名患者788次动作中实现任务准确率94.2%(kappa=0.934)
  • 结果可解释且符合临床规则,获四名医生认可并愿采纳

中风康复需评估运动组织方式而非仅成功与否,但当前评估受限于工具如行动研究臂测验(ARAT)将丰富行为压缩为单一序数评分,丢失关键质量细节。自动化方案常追求噪声标签下的高准确率,输出黑箱分数,难以落地临床。为此,我们提出xAARA:一种增强而非替代临床判断的系统。基于多视角视频,xAARA在任务、动作阶段和运动质量层面返回带校准不确定性和解释的ARAT评估。将临床评分视为不适定反问题,xAARA通过熵门控构建692个校准的多模态模型,并依据临床有效规则验证结果,低置信度情况自动延迟。在105名中风患者(788次练习)中,任务准确率达94.2%(Cohen's kappa=0.934),动作阶段准确率为81.3%(kappa=0.727),预测不确定性较单临床师评分降低96.1%。对主观案例,其评估始终匹配至少一位评分者,且从未出现超范围评分。四位独立临床医生验证了评估有效性,并表示愿意采用该系统。我们认为,严谨的不确定性量化与临床对齐的可解释性,是推动自动化评估从技术演示走向临床部署的关键桥梁。

原文摘要 · Abstract (English)

Tailoring stroke rehabilitation requires assessing how movements are organized, not merely if they succeed. Currently, this assessment is a rate-limiting bottleneck. Instruments like the Action Research Arm Test (ARAT) compress rich behavioral observations into single ordinal endpoints, discarding the movement-quality details that distinguish recovery from compensation. Automated alternatives typically chase accuracy on noisy, single-observer labels to output opaque scores - a technology-centric approach that rarely reaches clinical practice. To address this, we present xAARA: an engine designed to augment rather than replace clinical judgment. From multi-view video, xAARA returns ARAT assessments with calibrated uncertainty and explanations across task, movement-phase, and movement-quality levels. Treating clinical scoring as an ill-posed inference problem, xAARA composes 692 calibrated multimodal models via a Dynamic Bayesian Network with entropy-based gating. It qualifies results against clinical validity rules and defers low-confidence cases. In 105 stroke survivors (788 exercises), xAARA achieved 94.2% task accuracy (Cohen's kappa=0.934) and 81.3% movement-phase accuracy (kappa=0.727), reducing predictive uncertainty by 96.1% compared to single-clinician scoring. For subjective cases, it matched at least one rater 100% of the time and never returned out-of-range scores. Four independent clinicians validated the assessments and indicated willingness to adopt the system. We argue that principled uncertainty quantification and clinician-aligned explainability are the critical bridges moving automated assessment from technical demonstration to a deployable clinical tool.

中风康复不确定性量化多专家融合可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。