arXiv:2609.03971cs.IR2026-09

构建德国网页双维度分类平台,兼顾主题与地理信息

The WebKurator.de Platform: Combined Regional and Topical Web Curation

论文配图:The WebKurator.de Platform: Combined Regional and Topical Web Curation
图 1 · 摘自论文原文
  • 采用主题与地理分离的二维分类模型
  • 基于554万网站数据,85.17%地理定位在德国
  • 支持用户建议与审核,适合数字保存机构使用

国家图书馆等记忆机构面临系统化网络内容保存的挑战。现有目录式方法(如Curlie)以主题为中心,地理信息与主题、语言混杂。为此,我们推出WebKurator.de平台,面向德国网络内容,实现主题与地理的双重精准标注。该平台采用二维分类模型,分离主题分类与地理标注,并融合大语言模型主题识别、印鉴页地址提取与地理编码技术,支持用户提交建议并经审核后加入。平台基于德国印鉴数据集启动,包含554万网站,其中314万有印鉴页,成功提取并地理编码了这些地址。上述数据中258万(占85.17%)位于德国且已标注主题,构成初始数据基础,未来可通过用户持续补充。

原文摘要 · Abstract (English)

The systematic curation of the Web remains a central challenge for national libraries and memory institutions that aim to preserve culturally and regionally relevant content. Existing directory-based approaches such as Curlie implement a predominantly topic-centric, one-dimensional hierarchy, where geographic aspects are intertwined with topical and linguistic categories. To address this limitation, we present WebKurator.de, a collaborative platform for combined regional and topical web curation, initially focused on the German web. WebKurator introduces a two-dimensional curation model that explicitly separates topical categorization and geographic annotation. The system integrates LLM-based topic classification and imprint-based address extraction with geocoding, and supports user suggestions together with moderated review. The platform is bootstrapped from the German Imprints Dataset, a large-scale collection of 5.54 million websites. Among them, 3.14 million contain imprint pages, for which we successfully extracted and geocoded postal addresses. Of these, 2.58 million (85.17%) are located in Germany and also have an assigned topic label. These websites form the initial foundation of WebKurator.de and can be continuously extended through user suggestions.

网络保存双维度分类地理标注德国数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。