用保留上下文的分词过滤和知识图谱提升临床文本摘要质量
ConTextual: Improving Clinical Text Summarization in LLMs with Context-preserving Token Filtering and Knowledge Graphs
- 通过保留关键临床术语的分词过滤,精准提取重要信息
- 在两个公开数据集上超越现有基线,提升摘要语言连贯性和临床准确性
- 适合医疗AI研究者与临床决策支持系统开发者参考
非结构化临床数据是丰富且独特的信息来源,对优化患者诊疗决策至关重要。准确提取其中关键上下文信息,是释放其价值的核心。以往研究多对所有输入分词一视同仁,或依赖启发式过滤,常忽略细微临床线索,难以聚焦决策关键内容。本文提出ConTextual框架,融合上下文保持的分词过滤方法与领域专用知识图谱(KG),在保留关键临床术语的同时,通过结构化知识增强信息。实验在两个公开基准数据集上验证,该方法显著优于现有基线,同时提升摘要的语言连贯性与临床保真度。结果表明,分词级过滤与结构化检索具有互补性,为提升临床文本生成的精度提供了可扩展方案。
原文摘要 · Abstract (English)
Unstructured clinical data can serve as a unique and rich source of information that can meaningfully inform clinical practice. Extracting the most pertinent context from such data is critical for exploiting its true potential toward optimal and timely decision-making in patient care. While prior research has explored various methods for clinical text summarization, most prior studies either process all input tokens uniformly or rely on heuristic-based filters, which can overlook nuanced clinical cues and fail to prioritize information critical for decision-making. In this study, we propose Contextual, a novel framework that integrates a Context-Preserving Token Filtering method with a Domain-Specific Knowledge Graph (KG) for contextual augmentation. By preserving context-specific important tokens and enriching them with structured knowledge, ConTextual improves both linguistic coherence and clinical fidelity. Our extensive empirical evaluations on two public benchmark datasets demonstrate that ConTextual consistently outperforms other baselines. Our proposed approach highlights the complementary role of token-level filtering and structured retrieval in enhancing both linguistic and clinical integrity, as well as offering a scalable solution for improving precision in clinical text generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。