针对标签不平衡的结构化预测,提出新方法提升模型鲁棒性与准确性。
STaR-DRO: Stateful Tsallis Reweighting for Group-Robust Structured Prediction

- 用模块化提示工程解决格式漂移、标签歧义等常见错误。
- 在10亿到700亿参数模型上,标签F1提升14.46,实体跨度F1提升17.40。
- 专为持续困难群体加权,避免误扰简单组,适合医疗文本分析场景。
大语言模型进行结构化预测时,需保证输出在标签准确、本体约束、结构合法及证据支撑方面表现良好,尤其在标签不平衡和群体难度异质的情况下。本文提出统一框架以实现本体约束生成。首先,设计模块化提示工程架构,融合XML式结构、专家消歧规则、思维链推理、元数据感知决策逻辑、模式契约与自验证门控,应对上下文中的格式漂移、标签模糊、证据幻觉及元数据干扰等问题。其次,提出STaR-DRO方法,结合塔利斯镜像上升、稀疏entmax式原像映射、EMA平滑分组损失追踪、重标度上升信号与有界超额多倍器,相比传统DRO(依赖密集香农熵指数梯度更新)可避免高方差随机重加权、不对持久困难组加权、且减少单纯形竞争成本。实验在临床高风险任务EPPC Miner上验证,要求从患者-医生安全消息中提取层级标签与证据跨度。在1B-70B Llama模型上,提示工程使零样本提取平均标签F1提升+14.46,跨度F1提升+17.40。基于监督微调后,STaR-DRO进一步提高准确率与鲁棒性,平均标签F1再增+1.08与+2.20,分组验证交叉熵分别降低21.3%与14.8%(相较SFT与标准DRO)。结果推动面向以患者为中心的临床沟通挖掘技术发展。
原文摘要 · Abstract (English)
Structured prediction with large language models requires outputs that are label-accurate, ontology-constrained, structurally valid, and evidence-grounded under label imbalance and heterogeneous group difficulty. We present a unified framework for ontology-constrained generation. First, we introduce a modular prompt-engineering architecture combining XML-style structure, expert disambiguation rules, chain-of-thought reasoning, metadata-aware decision logic, schema contracts, and a self-validation gate. It targets recurrent in-context failures, including format drift, label ambiguity, evidence hallucination, and metadata-conditioned confusion. Second, we propose STaR-DRO, combining Tsallis mirror ascent, sparse entmax-style primal mapback, EMA-smoothed group-loss tracking, rescaled ascent signals, and bounded excess-only multipliers. Unlike conventional DRO, which relies on dense Shannon-entropy exponentiated-gradient updates, can introduce high-variance stochastic reweighting, assigns positive adversarial mass to groups that are not persistently hard, and incurs costs through simplex competition, STaR-DRO upweights only persistently hard groups without suppressing easier ones. We evaluate the framework on EPPC Miner, a clinically grounded high-stakes structured-prediction task requiring hierarchical label prediction and evidence-span extraction from patient-provider secure messages. Across 1B-70B Llama models, prompt engineering improves zero-shot extraction, yielding an average label F1 gain of +14.46 and a Span F1 gain of +17.40. Building on supervised fine-tuning, STaR-DRO further improves accuracy and robustness, increasing average label F1 by +1.08 and +2.20 while reducing mean groupwise validation cross-entropy by 21.3% and 14.8% relative to SFT and standard DRO, respectively. These results advance reliable automated communication mining for patient-centered clinical care analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。