用用户真实对话发现知识库盲区,精准补内容提升心理援助检索效果
Mind the Gap: Aligning Knowledge Bases with User Needs to Enhance Mental Health Retrieval
- 基于论坛数据识别真实用户需求中的知识缺口,指导有目标地扩充知识库
- 仅需42%-318%的新增内容即可接近全量知识库95%的检索性能
- 适合医疗信息平台、生成式AI在高风险场景下的可信内容建设
获取可靠心理健康信息对早期求助至关重要,但扩展知识库成本高且常与用户实际需求脱节,导致检索系统在面对非正式或情境化表达时表现不佳。本文提出一种基于AI的、以缺口为导向的知识库增强框架,通过叠加自然语言用户数据(如论坛帖子)来识别被忽视的主题,从而按覆盖度和实用性优先级进行扩充。案例研究对比了有目标的定向扩充与随机非定向扩充,评估四种检索增强生成(RAG)管道的检索相关性与实用性。定向扩充仅需42%(查询转换)、74%(重排序和分层)、318%(基线)的内容增长即可达到约95%全量参考语料库的性能;而非定向扩充则需232%、318%、403%、763%的大幅增加才能实现类似效果。结果表明,有针对性的知识库扩展可显著降低内容生产负担,同时保持高质量的检索与信息提供,为构建可信健康信息库及支持高风险领域生成式AI应用提供了可扩展方案。
原文摘要 · Abstract (English)
Access to reliable mental health information is vital for early help-seeking, yet expanding knowledge bases is resource-intensive and often misaligned with user needs. This results in poor performance of retrieval systems when presented concerns are not covered or expressed in informal or contextualized language. We present an AI-based gap-informed framework for corpus augmentation that authentically identifies underrepresented topics (gaps) by overlaying naturalistic user data such as forum posts in order to prioritize expansions based on coverage and usefulness. In a case study, we compare Directed (gap-informed augmentations) with Non-Directed augmentation (random additions), evaluating the relevance and usefulness of retrieved information across four retrieval-augmented generation (RAG) pipelines. Directed augmentation achieved near-optimal performance with modest expansions--requiring only a 42% increase for Query Transformation, 74% for Reranking and Hierarchical, and 318% for Baseline--to reach ~95% of the performance of an exhaustive reference corpus. In contrast, Non-Directed augmentation required substantially larger and thus practically infeasible expansions to achieve comparable performance (232%, 318%, 403%, and 763%, respectively). These results show that strategically targeted corpus growth can reduce content creation demands while sustaining high retrieval and provision quality, offering a scalable approach for building trusted health information repositories and supporting generative AI applications in high-stakes domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。