arXiv:2412.14533cs.CL2024-12中稿 · SIGIR Demo Paper T…

用多特征搜索让海量文献探索更高效,支持聚类、问答和时间线分析。

ClusterChat: Multi-Feature Search for Corpus Exploration

  • 融合文本嵌入聚类与关键词/语义搜索,实现文档组织与检索一体化。
  • 在400万篇PubMed摘要上验证,提升趋势洞察力且响应迅速。
  • 适合科研人员、分析师快速掌握文献全景,尤其适合医学、金融等领域。

在生物医学、金融和法律领域,持续生成的海量文档使大规模文本语料探索面临挑战。传统基于关键词的搜索方法常孤立地返回文档,难以发现整体趋势与关联。我们提出ClusterChat(演示视频与源代码见:https://github.com/achouhan93/ClusterChat),一个开源语料探索系统,整合了基于文本嵌入的文档聚类、词法与语义搜索、时间线驱动探索以及语料与文档级问答等多种特征搜索能力。通过在包含四百万篇摘要的PubMed数据集上的两个案例研究,验证了ClusterChat在保持大规模文档集合可扩展性与响应速度的同时,显著提升了上下文感知的洞察力。

原文摘要 · Abstract (English)

Exploring large-scale text corpora presents a significant challenge in biomedical, finance, and legal domains, where vast amounts of documents are continuously published. Traditional search methods, such as keyword-based search, often retrieve documents in isolation, limiting the user's ability to easily inspect corpus-wide trends and relationships. We present ClusterChat (The demo video and source code are available at: https://github.com/achouhan93/ClusterChat), an open-source system for corpus exploration that integrates cluster-based organization of documents using textual embeddings with lexical and semantic search, timeline-driven exploration, and corpus and document-level question answering (QA) as multi-feature search capabilities. We validate the system with two case studies on a four million abstract PubMed dataset, demonstrating that ClusterChat enhances corpus exploration by delivering context-aware insights while maintaining scalability and responsiveness on large-scale document collections.

语料探索聚类搜索多特征检索PubMed

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。