用分层结构提升医学多文档摘要的清晰度与可信度
Leveraging Hierarchical Organization for Medical Multi-document Summarization
- 将文档分层组织输入模型,增强跨文档信息关联
- 专家评估显示生成摘要优于人工摘要,且事实性、覆盖性更优
- 适合医疗AI、临床决策支持系统开发者参考
医学多文档摘要(MDS)需有效处理跨文档关系。本文探究在输入中引入分层结构是否能提升模型组织和上下文化信息的能力,相比传统扁平方法。我们在三个大语言模型上测试了两种分层组织方式,并通过自动指标、模型评估及领域专家评价(偏好、可理解性、清晰度、复杂度、相关性、覆盖率、事实性、连贯性)进行全面评估。结果表明,人类专家更偏好模型生成的摘要;分层方法普遍保持事实性、覆盖率与连贯性,同时提升摘要偏好度。此外,我们检验GPT-4模拟判断与人类判断的一致性,发现客观维度上一致性更高。研究证明,分层结构可提升模型生成医学摘要的清晰度,同时保证内容覆盖,为提升生成摘要的人类接受度提供实用路径。
原文摘要 · Abstract (English)
Medical multi-document summarization (MDS) is a complex task that requires effectively managing cross-document relationships. This paper investigates whether incorporating hierarchical structures in the inputs of MDS can improve a model's ability to organize and contextualize information across documents compared to traditional flat summarization methods. We investigate two ways of incorporating hierarchical organization across three large language models (LLMs), and conduct comprehensive evaluations of the resulting summaries using automated metrics, model-based metrics, and domain expert evaluation of preference, understandability, clarity, complexity, relevance, coverage, factuality, and coherence. Our results show that human experts prefer model-generated summaries over human-written summaries. Hierarchical approaches generally preserve factuality, coverage, and coherence of information, while also increasing human preference for summaries. Additionally, we examine whether simulated judgments from GPT-4 align with human judgments, finding higher agreement along more objective evaluation facets. Our findings demonstrate that hierarchical structures can improve the clarity of medical summaries generated by models while maintaining content coverage, providing a practical way to improve human preference for generated summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。