arXiv:2609.07766cs.CLcs.AI2026-09

用任务知识筛选有效方法,提升自杀风险评估的可靠性

Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment

论文配图:Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment
图 1 · 摘自论文原文
  • 基于临床知识设计针对性模型结构,避免盲目堆砌技巧
  • 在1635篇标注数据上验证31种方法,仅5种带来可靠提升
  • 适合医疗文本分析、小样本高风险场景的研究与应用

从社交媒体文本中评估自杀风险属于小样本、高风险场景,不仅需预测严重程度,还需提供依据和临床相关因素。然而,常用NLP方法如模型扩展、合成数据、损失重加权、集成学习和阈值调整,常未在严重类别不平衡、多输出及有限作者级数据条件下验证其有效性。本研究基于1,635条临床标注帖子,对7类共31种预设技术进行约300次控制实验,首次系统审计该方法手册在此场景下的表现。结果发现仅5项技术有显著收益。系统最终实现4级风险预测(0.8203)、证据片段识别(0.7953)及24个临床因素的宏F1(0.7045),综合得分0.7781,排名53支队伍第三。关键创新包括:将因素预测重构为文本与术语库定义之间的蕴含关系,采用多样化集成与类别平衡训练;风险预测引导7模型证据标签器;证据用于约束符号规则;困难类别独立处理。此外通过部署一致校准修正验证与测试时评分不匹配问题,显著提升因素系统性能。提出‘任务条件化技术选择’原则:仅当任务知识或实证支持时才保留特定方法。

原文摘要 · Abstract (English)

Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity prediction but also supporting evidence and clinically relevant risk and protective factors. Yet common NLP techniques, including model scaling, synthetic data, loss reweighting, ensembling, and threshold tuning, are often applied without testing whether their gains hold up under severe class imbalance, coupled outputs, and limited author-level data. We study 1,635 clinician-annotated posts and audit 31 pre-specified techniques from 7 methodological families through roughly 300 controlled experiments on author-disjoint partitions. We found no prior audit of this playbook in this regime. The findings guide a task-grounded system for three outputs: 4-level suicide risk, evidence spans, and 24 clinical risk and protective factors. Only 5 of 31 comparisons produced reliable gains. We reformulate factor prediction as entailment between each post and its codebook definitions, using an architecturally diverse ensemble with class-balanced training and score rescaling. Risk predictions condition a 7-model evidence tagger ensemble; evidence restricts symbolic risk rules; and a difficult risk class is routed separately. The factor predictor remains independent because risk evidence provides no additional factor signal. We also correct a mismatch between validation scores used for threshold fitting and test-time ensemble scores through deployment-consistent calibration, yielding the largest improvement to the factor system. The final system achieves 0.8203 for risk, 0.7953 for evidence, and 0.7045 macro-F1 for factors, with a 0.7781 composite, ranking third among 53 teams. We call the underlying principle task-conditioned technique selection: retain techniques only when task-specific knowledge, structure, or empirical evidence justifies them.

自杀风险评估小样本学习可解释性AI临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。