arXiv:2512.23637cs.CL2025-12

构建首个专家标注的医疗问答摘要数据集,助力精准理解患者提问。

A Dataset and Benchmark for Consumer Healthcare Question Summarization

  • 从社交平台收集1507条真实医疗问题,由领域专家标注摘要。
  • 基于该数据集评估多个主流摘要模型,验证其在医疗场景下的性能。
  • 为社交媒体上的患者诉求分析提供可靠基准,适合医疗NLP研究者使用。

公众寻求健康信息的行为导致网络上充斥着大量与健康相关的用户提问。通常,用户会使用过于详细和冗余的信息来描述自身病情或医疗需求,给自然语言理解带来挑战。一种应对策略是将问题进行摘要,提炼原始问题的核心信息。近年来,大规模数据集显著推动了多文档摘要和对话摘要等任务的发展。然而,缺乏针对消费者医疗问题摘要任务的领域专家标注数据集,制约了高效摘要系统的发展。为此,我们引入了一个新数据集CHQ-Sum,包含1507个经领域专家标注的消费者健康问题及其对应摘要。该数据集源自社区问答论坛,为理解社交媒体上的患者健康相关内容提供了宝贵资源。我们在多个最先进的摘要模型上对该数据集进行了基准测试,展示了其有效性。

原文摘要 · Abstract (English)

The quest for seeking health information has swamped the web with consumers health-related questions. Generally, consumers use overly descriptive and peripheral information to express their medical condition or other healthcare needs, contributing to the challenges of natural language understanding. One way to address this challenge is to summarize the questions and distill the key information of the original question. Recently, large-scale datasets have significantly propelled the development of several summarization tasks, such as multi-document summarization and dialogue summarization. However, a lack of a domain-expert annotated dataset for the consumer healthcare questions summarization task inhibits the development of an efficient summarization system. To address this issue, we introduce a new dataset, CHQ-Sum,m that contains 1507 domain-expert annotated consumer health questions and corresponding summaries. The dataset is derived from the community question answering forum and therefore provides a valuable resource for understanding consumer health-related posts on social media. We benchmark the dataset on multiple state-of-the-art summarization models to show the effectiveness of the dataset

医疗NLP问答摘要数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。