arXiv:2604.25057cs.LGcs.DL2026-04

CiteRadar帮学者一键生成含全球分布的引文地图与作者画像。

CiteRadar: A Citation Intelligence Platform for Researcher Profiling and Geographic Visualization

论文配图:CiteRadar: A Citation Intelligence Platform for Researcher Profiling and Geographic Visualization
图 1 · 摘自论文原文
  • 输入谷歌学术ID,自动整合五类数据源生成完整引文档案。
  • 修复引文解析错误,使同名作者误合并问题导致的h指数偏差减少9倍。
  • 自动生成可离线查看的交互式世界地图,支持城市级研究人员定位。

理解学术引用的地理覆盖范围和社群结构对职业发展、项目申请和合作发现日益重要,但现有工具或需昂贵订阅,或仅提供聚合数据而缺乏细粒度作者元信息。我们提出CiteRadar,一个开源系统,仅需输入单个Google Scholar用户标识符,即可通过命令行自动输出包含:作者完整论文列表、所有被引文献及增强后的作者元数据、按引用频次与h指数排序的作者表、纯文本统计摘要,以及自包含的交互式HTML世界地图。该系统整合了Google Scholar、OpenAlex、CrossRef、Semantic Scholar和OpenStreetMap Nominatim五个异构数据源,采用五阶段处理流程。关键技术贡献包括:(1)抗非断行空格干扰的学者元字符串解析器,解决谷歌学术网页中普遍存在的未文档化字符问题;(2)基于停用词过滤机构名称相似度的两阶段作者消歧系统,有效防止同名作者误合并,实证消除高达9倍的h指数误判;(3)将OpenAlex网页链接转换为API链接的修正方案,使带城市层级位置数据的作者记录比例从0%提升至约60%;(4)采用对数缩放的Folium交互式世界地图,以每城市研究员弹窗形式呈现,生成独立运行的HTML文件。

原文摘要 · Abstract (English)

Understanding the geographic reach and community structure of one's scholarly citations is increasingly valuable for career development, grant applications, and collaboration discovery -- yet accessible tools for answering these questions remain scarce. Existing bibliometric platforms either require costly institutional subscriptions or expose only aggregate citation counts without granular per-author metadata. We present CiteRadar, an open-source system that accepts a single Google Scholar user identifier and automatically produces a structured output folder containing: the author's complete publication list, all retrieved citing papers with enriched author metadata, two ranked author tables (by citation frequency and by h-index), a plain-text statistical summary, and a self-contained interactive HTML world map -- all from a single command-line invocation. CiteRadar integrates five heterogeneous data sources -- Google Scholar, OpenAlex, CrossRef, Semantic Scholar, and OpenStreetMap Nominatim -- through a carefully engineered five-stage pipeline. Key technical contributions include: (1) a Scholar meta-string parser resilient to Unicode non-breaking-space separators, a pervasive but undocumented quirk in Scholar's HTML that silently corrupts venue and year fields when unhandled; (2) a two-stage author disambiguation system using stop-word-filtered institution name similarity to guard against the well-known same-name entity-merging failure mode in bibliometric databases, demonstrated to eliminate h-index attribution errors of up to 9x the correct value; (3) an OpenAlex web-URL to API-URL conversion fix that raises the fraction of author records with city-level location data from 0% to ~60%; and (4) a logarithmically-scaled interactive Folium world map with per-city researcher popups, rendered as a fully self-contained HTML file.

引文分析学术画像地理可视化开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。