arXiv:2604.15203cs.CL2026-04ACL

构建医疗设备不良事件多标签分类动态基准,支持不确定性量化评估。

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events

论文配图:MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events
图 1 · 摘自论文原文
  • 基于真实不良事件报告构建持续更新的动态数据集,防止训练污染。
  • 发现小模型在准确率与不确定性量化间平衡最佳,大模型虽提升罕见标签性能但不确定性弱。
  • 适合关注医疗AI可解释性与可靠性评估的研究者与临床应用开发者。

高风险领域如医疗中的机器学习不仅需要强预测性能,还需可靠的不确定性量化(UQ)以支持人工监督。多标签文本分类(MLTC)是该领域核心任务,但受标签不平衡、依赖关系及组合复杂性挑战。现有基准日益饱和且可能受训练数据污染影响,难以区分真实推理能力与记忆现象。我们提出MADE,一个源自医疗设备不良事件报告并持续更新的动态MLTC基准,避免数据污染。MADE具有层次化长尾标签分布,支持严格时间划分下的可复现评估。我们在超过20种编码器与解码器模型上建立基线,涵盖微调与少样本设置(指令微调/推理变体,本地/云端访问)。系统评估熵/一致性基与自述式不确定性量化方法。结果表明:小规模判别性微调解码器在头尾标签准确率上表现最强且保持良好UQ;生成式微调提供最可靠UQ;大型推理模型虽提升罕见标签性能,但不确定性表现意外薄弱;自述置信度并非不确定性可靠代理。相关工作已公开于 https://hhi.fraunhofer.de/aml-demonstrator/made-benchmark。

原文摘要 · Abstract (English)

Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification (UQ) to support human oversight. Multi-label text classification (MLTC) is a central task in this domain, yet remains challenging due to label imbalances, dependencies, and combinatorial complexity. Existing MLTC benchmarks are increasingly saturated and may be affected by training data contamination, making it difficult to distinguish genuine reasoning capabilities from memorization. We introduce MADE, a living MLTC benchmark derived from {m}edical device {ad}verse {e}vent reports and continuously updated with newly published reports to prevent contamination. MADE features a long-tailed distribution of hierarchical labels and enables reproducible evaluation with strict temporal splits. We establish baselines across more than 20 encoder- and decoder-only models under fine-tuning and few-shot settings (instruction-tuned/reasoning variants, local/API-accessible). We systematically assess entropy-/consistency-based and self-verbalized UQ methods. Results show clear trade-offs: smaller discriminatively fine-tuned decoders achieve the strongest head-to-tail accuracy while maintaining competitive UQ; generative fine-tuning delivers the most reliable UQ; large reasoning models improve performance on rare labels yet exhibit surprisingly weak UQ; and self-verbalized confidence is not a reliable proxy for uncertainty. Our work is publicly available at https://hhi.fraunhofer.de/aml-demonstrator/made-benchmark.

多标签分类医疗AI不确定性量化动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。