Eliot让研究者实时追踪科学文献演化,透明可视地分析热点变化。
Eliot: Interactively $\underline{E}$xploring Fast-Changing Scientific $\underline{Li}$terature Trends with $\underline{O}$nline Da$\underline{t}$a and Learning

- 查询时动态抓取arXiv论文,用嵌入+聚类生成主题并展示发表年份分布。
- 在8个领域测试中,MiniLM+10维UMAP+层次聚类效果最佳,主题一致性高。
- 适合需要可审计、快速掌握前沿动向的研究者,尤其适用于技术更新快的领域。
科学出版物的快速增长使跟踪快速变化的研究领域愈发困难。搜索引擎和基于大模型的助手虽能检索或摘要论文,但常隐藏数据筛选、组织方式及时间关联。我们提出公开部署的交互式系统Eliot,实现可追溯的科学文献演化探索。基于对大语言模型和自动规划调度领域的两项研究,Eliot将文献演化分析推广至无需手工分类体系和领域特定脚本的通用场景。用户输入查询词与过滤条件后,系统在查询时从arXiv获取论文,以标题和摘要表示每篇论文,聚类成主题,分配代表性关键词,并可视化各主题的发表年份分布。我们评估Eliot作为应用系统和交互式研究辅助工具的表现:在八个arXiv领域进行离线配置研究,比较不同文档表示、降维方法与聚类算法,使用内在聚类与主题连贯性指标,结果支持以MiniLM嵌入搭配10维UMAP与层次聚类为实用默认方案;情景式调查与专家焦点小组评估可解释性与使用场景,参与者在85%的情景回应中认为聚类标签有意义,反馈表明Eliot最适用于需要可审计的快速技术领域概览。
原文摘要 · Abstract (English)
The rapid growth of scientific publishing has made it increasingly difficult to track how fast-moving areas evolve. Search engines and LLM-based assistants retrieve or summarize papers, but often hide how the corpus was selected, organized, or connected to temporal patterns. We present $\texttt{Eliot}$, a publicly deployed interactive system for traceable exploration of evolving scientific literature. Motivated by two studies on Large Language Models (LLMs) and Automated Planning and Scheduling (APS), $\texttt{Eliot}$ generalizes literature-evolution analysis beyond hand-built taxonomies and domain-specific scripts. Given explicit query terms and filters, it retrieves arXiv papers at query time, represents each paper by title and abstract, clusters the corpus into themes, assigns representative keywords, and visualizes each cluster's publication-year distribution. We evaluate $\texttt{Eliot}$ as both an applied system and an interactive research aid. An offline configuration study across eight arXiv domains compares document representations, dimensionality reduction methods, and clustering algorithms using intrinsic clustering and topic-coherence metrics; the results support MiniLM embeddings with 10-dimensional UMAP and Agglomerative Clustering as a practical default. A scenario-based survey and expert focus group assess interpretability and use contexts: participants rated cluster labels as meaningful in 85% of scenario responses, and feedback indicated that $\texttt{Eliot}$ is most valuable for auditable overviews of rapidly changing technical areas. These results suggest that query-time clustering and temporal inspection can complement search and generation tools by helping researchers inspect and refine the evidence behind literature trends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。