arXiv:2605.21845cs.CLcs.AI2026-05中稿 · IEEE ICHI 2026

用复杂度评分优化提示策略,提升自杀原因分析的准确性

Comparing LLM and Fine-Tuned Model Performance on NVDRS Circumstance Extraction with Varying Prompt Complexity

  • 根据编码手册结构设计复杂度评分,动态选择提示方式
  • 大模型在低频复杂情形下显著优于微调模型(如RoBERTa)
  • 方法可推广至GPT-5.2、Gemini 2.5 Pro等主流大模型

自杀是美国主要死因之一,理解其前因需要从死亡调查文本中提取结构化信息。许多情境需超越关键词匹配的语义推理。我们提出“复杂度评分”算法,分析编码手册结构,预测在哪些情况下包含完整编码指南的详细提示优于仅含名称的提示。据此构建混合方法,按情形动态选择提示策略。在国家暴力死亡报告系统(NVDRS)的25个高推理复杂度情形上,评估大语言模型(LLMs)与微调的RoBERTa性能。结果显示,在训练数据不足的低频情形中,大模型表现显著更优。进一步验证表明,该框架在前沿大模型(GPT-5.2、Gemini 2.5 Pro、Llama-3 70B)上具有一致性能模式。研究支持一种混合架构:大模型处理罕见且复杂的场景,微调模型处理常见场景。

原文摘要 · Abstract (English)

Suicide is a leading cause of death in the United States, and understanding the circumstances that precede it requires extracting structured information from death investigation narratives. Many of these circumstances require semantic inference beyond simple keyword matching. We develop a ``Complexity Score'' algorithm that analyzes coding manual structure to predict when detailed prompts with full coding guidelines improve over name-only prompts. We then construct a hybrid approach that selects prompt strategy per circumstance. We evaluate large language models (LLMs) against fine-tuned RoBERTa on 25 inferentially complex circumstances from the National Violent Death Reporting System (NVDRS). We found that LLMs substantially outperform on low-prevalence circumstances where training data is insufficient. We further demonstrate that our framework generalizes across frontier LLMs, with GPT-5.2, Gemini 2.5 Pro and Llama-3 70B showing consistent performance patterns. These findings support a hybrid architecture where LLMs handle rare, inferentially complex circumstances while fine-tuned models handle common ones.

大模型信息抽取自杀研究提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。