通过聚类与新损失函数提升医疗搜索意图识别准确率。
Enhancing Healthcare Search Intent Recognition with Query Representation Learning and Session Context
- 用聚类整合相似查询,改进查询表示学习。
- 在两个真实数据集上,意图分类准确率显著提升。
- 适合医疗信息检索、搜索系统优化的研究者参考。
医疗搜索意图分类对改善在线健康信息的精准推送至关重要。由于医疗查询复杂且高质量标注数据稀缺,模型构建面临挑战。以往研究利用用户点击日志和成对损失函数建模共点击行为,但许多医疗查询存在多重意图,导致点击行为模糊或分歧。仅基于全局统计学习单一主流意图,可能在特定会话中造成偏差与性能下降。为此,本文通过聚类聚合相似查询,并引入新损失函数捕捉医疗查询的多面性,实现更高效准确的表示学习。同时提出一致性率(CR)评分,量化查询歧义及全局与会话级意图间的不一致程度,并验证了将学习到的查询表示融入上下文会话分类的有效方法。在真实医疗搜索日志(HS数据集)和公开的TripClick数据集上的实验表明,该方法不仅提升了查询表示的聚类效果,还显著增强了后续意图分类精度。
原文摘要 · Abstract (English)
Classifying the intent behind healthcare search queries is crucial for improving the delivery of online healthcare information. The intricate nature of medical search queries, coupled with the limited availability of high-quality labeled data, presents substantial challenges for developing efficient classification models. Previous studies have exploited user interaction data, such as user clicks from search logs and employed pairwise loss functions to model co-click behavior for query representation learning. However, many health queries could have multiple intents, resulting in ambiguous or divergent click behavior. Furthermore, learning the single most popular intent of queries as inferred from global statistics based on the aggregate behavior of different users could potentially lead to disparity and performance drop when classifying the query intent within specific search sessions. To address these limitations, our work improves the query representation learning by aggregating similar queries via clustering, and introducing a novel loss function designed to capture the multifaceted nature of health search queries, resulting in a more scalable and accurate learning procedure. Furthermore, we quantify the ambiguity of health queries and the misalignment between global search intents and those discerned from individual sessions, by introducing the concordance rate (CR) score, and demonstrate a simple and effective method for incorporating our learned query representation into contextual, session-based search intent classification. Our extensive experimental results and analysis on two real-world search log datasets, i.e., a Health Search (HS) dataset and the publicly available TripClick dataset, demonstrate that our approach not only improves the intrinsic clustering metrics for query representation learning but also enhances accuracy for subsequent search intent classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。