arXiv:2605.31512cs.CL2026-05

多语言骨科临床文本分类,提升低资源医疗场景下的决策可靠性。

Reliable Multilingual Orthopedic Decision Support from Clinical Narratives: Language-Aware Adaptation and Verification-Guided Deferral

论文配图:Reliable Multilingual Orthopedic Decision Support from Clinical Narratives: Language-Aware Adaptation and Verification-Guided Deferral
图 1 · 摘自论文原文
  • 用语言感知适配器增强多语言模型,实现跨语言骨科文本理解。
  • 在自然临床数据分布下,平均宏F1达0.8792,优于其他模型。
  • 引入可解释的拒答机制,覆盖72.3%时准确率提升至84.4%。

多语言骨科决策支持在低资源医疗环境中面临挑战,临床文本常含专业术语、混合文字、证据不全、标签不平衡及语言依赖性记录模式。本文提出一种面向可靠性的框架,用于分类英文、印地语和旁遮普语的自由文本骨科笔记。比较了任务对齐的多语言变压器编码器、微调的DistilBERT基线、零样本指令微调的大语言模型(LLMs)以及领域自适应编码器IndicBERT-HPA。IndicBERT-HPA通过语言感知的骨科适配头增强IndicBERT,支持临床相关的多语言表示学习。评估不仅包括总体准确率,还涵盖每类表现、ROC-AUC、AUPRC、预期校准误差、跨语言稳定性及在平衡与自然流行分布下的鲁棒性。零样本LLMs在封闭集分类中仍显著逊于任务适配编码器,且存在语言依赖性不稳定性。在自然临床流行分布下,IndicBERT-HPA表现最佳,平均宏F1达0.8792,宏AUROC为0.894,AUPRC为0.902。进一步引入确定性选择性验证层,结合置信度门控、证据一致性检查和语言风险筛查。在随机选取的5,000条记录子集上,该机制在72.3%覆盖率下实现84.4%的选择性准确率和0.76的选择性宏F1,优于全接受预测的71.5%准确率和0.65宏F1。结果支持具有显式拒答的可靠性导向多语言临床决策支持。

原文摘要 · Abstract (English)

Multilingual orthopedic decision support remains challenging in low-resource healthcare settings, where clinical narratives contain specialized terminology, mixed scripts, incomplete evidence, label imbalance and language-dependent documentation patterns. This article presents a reliability-oriented framework for classifying free-text orthopedic notes in English, Hindi and Punjabi. We compare task-aligned multilingual transformer encoders, a task-fine-tuned DistilBERT baseline, zero-shot instruction-tuned large language models (LLMs) and a domain-adaptive encoder, IndicBERT-HPA. IndicBERT-HPA augments IndicBERT with language-aware orthopedic adapter heads to support clinically relevant multilingual representation learning. Evaluation extends beyond aggregate accuracy to per-class performance, ROC-AUC, AUPRC, expected calibration error, cross-language stability and robustness under controlled balanced and natural-prevalence distributions. The evaluated zero-shot LLMs remain substantially less effective than task-adapted encoders for closed-set classification, with language-dependent instability. Under natural clinical prevalence, IndicBERT-HPA achieves the strongest overall performance, reaching an averaged Macro-F1 of 0.8792, Macro-AUROC of 0.894 and AUPRC of 0.902. We further implement a deterministic selective-verification layer combining confidence gating, evidence-consistency checking and language-risk screening. On a randomly selected held-out 5,000-record subset, it achieves 84.4% selective accuracy and 0.76 selective Macro-F1 at 72.3% coverage, compared with 71.5% accuracy and 0.65 Macro-F1 for accept-all prediction. These results support reliability-oriented multilingual clinical decision support with explicit deferral.

多语言骨科决策支持可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。