通过分阶段智能推理,提升临床文本摘要的准确性与可信度。
AgenticSum: An Agentic Inference-Time Framework for Faithful Clinical Text Summarization
- 将摘要任务拆解为选上下文、生成、验证和修正四阶段协同处理
- 在两个公开数据集上显著降低幻觉内容,各项指标优于基线模型
- 适合需要高可靠性医疗文本处理的研究者与临床系统开发者
大型语言模型(LLMs)在自动化临床文本摘要方面潜力巨大,但因临床文档长度长、噪声多、异质性强,保持事实一致性仍具挑战。我们提出AgenticSum,一种推理时的智能体框架,通过分离上下文选择、生成、验证与针对性修正环节,减少幻觉内容。该框架将摘要过程分解为协调的四个阶段:压缩相关上下文、生成初稿、利用内部注意力信号识别支持薄弱的片段,并在监督控制下对标注内容进行选择性修正。我们在两个公开数据集上评估AgenticSum,采用基于参考的指标、LLM作为裁判的评估以及人工评价。在多种衡量标准下,AgenticSum相较于原始LLM及其他强基线模型均表现出持续改进。结果表明,带有针对性修正的结构化智能体设计,是利用LLM提升临床笔记摘要质量的有效推理时解决方案。
原文摘要 · Abstract (English)
Large language models (LLMs) offer substantial promise for automating clinical text summarization, yet maintaining factual consistency remains challenging due to the length, noise, and heterogeneity of clinical documentation. We present AgenticSum, an inference-time, agentic framework that separates context selection, generation, verification, and targeted correction to reduce hallucinated content. The framework decomposes summarization into coordinated stages that compress task-relevant context, generate an initial draft, identify weakly supported spans using internal attention grounding signals, and selectively revise flagged content under supervisory control. We evaluate AgenticSum on two public datasets, using reference-based metrics, LLM-as-a-judge assessment, and human evaluation. Across various measures, AgenticSum demonstrates consistent improvements compared to vanilla LLMs and other strong baselines. Our results indicate that structured, agentic design with targeted correction offers an effective inference time solution to improve clinical note summarization using LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。