arXiv:2509.04462cs.CLcs.AI2025-09被引 6

GPT-5在生物医学NLP任务中表现优于GPT-4o,尤其在推理和多模态问答上。

Benchmarking GPT-5 for biomedical natural language processing

  • 在五类核心生物医学任务中统一评测GPT-5与GPT-4o的零次、一次、五次提示性能。
  • GPT-5在诊断推理和多模态问答任务中提升显著,且单位正确预测成本低30%~50%。
  • 适合高精度医疗问答场景,推荐分层提示策略以平衡效率与准确性。

生物医学文献与临床叙述对自然语言理解提出多重挑战,涵盖实体抽取、文档合成及多步诊断推理。本研究扩展统一基准,评估GPT-5与GPT-4o在零次、一次、五次提示下于五大核心生物医学NLP任务的表现:命名实体识别、关系抽取、多标签文档分类、摘要生成与简化,并覆盖九个扩展的生物医学QA数据集,涵盖事实知识、临床推理与多模态视觉理解。采用标准化提示、固定解码参数与一致推理流程,评估模型性能、延迟与按令牌计费的成本。GPT-5持续优于GPT-4o,尤其在诊断推理密集型数据集如MedXpertQA和DiagnosisArena中表现突出,多模态问答亦稳定提升。核心任务中,GPT-5在化学实体识别与ChemProt得分更高,但疾病实体识别与摘要仍低于领域微调基线。尽管输出更长,其延迟相当,单位正确预测成本降低30%至50%。细粒度分析显示诊断、治疗与推理子类型均有改进,但边界敏感抽取与证据密集摘要仍具挑战。总体而言,GPT-5已接近部署可用的生物医学问答性能,兼具准确率、可解释性与经济效率。结果支持分层提示策略:大规模或成本敏感应用使用直接提示,复杂或高风险场景采用思维链支架,凸显精确性与事实一致性关键时仍需混合方案。

原文摘要 · Abstract (English)

Biomedical literature and clinical narratives pose multifaceted challenges for natural language understanding, from precise entity extraction and document synthesis to multi-step diagnostic reasoning. This study extends a unified benchmark to evaluate GPT-5 and GPT-4o under zero-, one-, and five-shot prompting across five core biomedical NLP tasks: named entity recognition, relation extraction, multi-label document classification, summarization, and simplification, and nine expanded biomedical QA datasets covering factual knowledge, clinical reasoning, and multimodal visual understanding. Using standardized prompts, fixed decoding parameters, and consistent inference pipelines, we assessed model performance, latency, and token-normalized cost under official pricing. GPT-5 consistently outperformed GPT-4o, with the largest gains on reasoning-intensive datasets such as MedXpertQA and DiagnosisArena and stable improvements in multimodal QA. In core tasks, GPT-5 achieved better chemical NER and ChemProt scores but remained below domain-tuned baselines for disease NER and summarization. Despite producing longer outputs, GPT-5 showed comparable latency and 30 to 50 percent lower effective cost per correct prediction. Fine-grained analyses revealed improvements in diagnosis, treatment, and reasoning subtypes, whereas boundary-sensitive extraction and evidence-dense summarization remain challenging. Overall, GPT-5 approaches deployment-ready performance for biomedical QA while offering a favorable balance of accuracy, interpretability, and economic efficiency. The results support a tiered prompting strategy: direct prompting for large-scale or cost-sensitive applications, and chain-of-thought scaffolds for analytically complex or high-stakes scenarios, highlighting the continued need for hybrid solutions where precision and factual fidelity are critical.

GPT-5生物医学NLP推理能力成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。