arXiv:2608.20587cs.CVcs.AI2026-08

不微调模型,通过聚合个体步态数据提升帕金森步态评估准确率。

Aggregate, Don't Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity

  • 基于受试者级后验概率聚合,避免对模型进行微调。
  • 在隐藏测试集上达到0.6945的宏平均F1,领先第二名11个百分点。
  • 适合医疗计算、跨机构数据建模的研究者参考。

本文描述了在2026年MoCha帕金森步态挑战赛中的优胜方案,该方案从临床机构未见数据中预测MDS-UPDRS步态严重程度。系统使用冻结的公开运动编码器,仅通过一个$4\times512$的线性层,实现了0.6945的宏平均F1,在隐藏测试集上排名第一,优于第二名的0.5807和组织方基线的0.4289。几乎所有性能提升来自三个常被视为辅助流程的步骤:复现基准评测的精确头部设计、按受试者分组内对每步后验概率求平均、以及无需标签的特征均值与决策点的归纳校准。四种形式的编码器微调全部失败,十种替代编码器表现更差。所有消融实验结果均需在隐藏测试集上验证,因我们自建的留两队列外交叉验证显示其与最终得分呈负相关。我们完整公开负面记录,并将受试者级聚合确认为本基准的性能上限。

原文摘要 · Abstract (English)

We describe the winning entry to the MoCha 2026 Benchmark and Challenge on Parkinsonian Gait, which predicts MDS-UPDRS gait severity from canonicalized SMPL motion recorded at clinical sites unseen during training. The system reaches 0.6945 macro-F1 on the hidden test and ranked first of 58 entries, ahead of the runner-up at 0.5807 and the organizers' baseline at 0.4289, on a frozen public motion encoder with a single $4\times512$ linear layer. Nearly all of the margin comes from three stages usually treated as bookkeeping: reproducing the reference benchmark's exact head recipe, averaging per-walk posteriors within the subject grouping the organizers ship, and a label-free transductive calibration of the feature mean and the decision operating point. Fine-tuning the encoder lost in four distinct forms, and ten alternative encoders were worse. Every ablation number is a paid read on the hidden test, because our own leave-two-cohort-out cross-validation proved anti-correlated with the deciding score over eleven configurations. We give the negative record in full, and identify our largest gain, subject-level aggregation, as the binding ceiling on this benchmark.

帕金森步态评估跨机构建模后验聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。