Svarna整合5大类希腊语语料,提供超5亿词免费在线分析工具。
Svarna: An Open Corpus Workbench for Modern Greek

- 整合5个数据库,覆盖多种语体,总规模超507百万词
- 支持关键词共现、频次分析、正则搜索等10余种查询功能
- 无需登录或安装,可自建语料库,适合语言学研究者
本文介绍Svarna——一个面向现代希腊语的免费开源网络语料库工作台。Svarna整合了涵盖不同语体、机构文本、文学作品、方言、社交媒体和历史文献的5个数据库,总计超过507百万词和约2900万句。该平台解决了希腊语技术资源分散、访问受限或已下线的问题。用户无需登录、安装或培训即可使用。系统提供关键词在上下文(KWIC)标记、按语体归一化的频率分析、互信息计算的共现提取、包含93个话语标记的词典及其分布特征、文本级分析工具(如n-gram、变体、共现网络)、对数比值法进行语体对比、正则表达式搜索,以及可选的大语言模型层用于语用标注和自由研究模式。系统基于FastAPI后端与SQLite FTS5全文索引,部署于Azure的Docker容器中,源码、构建脚本和部署配置均在GitHub公开,用户可自行添加语料并部署实例。本文详述系统设计、语料结构及典型应用案例。Svarna是探索现有数据的第一步,有望为未来更深入的研究奠定基础。
原文摘要 · Abstract (English)
This paper introduces Svarna, a free, open-source, web-based corpus workbench for modern Greek. Svarna integrates five databases covering various registers, institutional, literary, dialectal, social media, and historical, to provide a total of more than 507 million words and around 29 million sentences. This platform addresses the chronic gaps in Greek language technology. Although various corpus resources exist, they are scattered across different platforms, and in many cases, institutional access is restricted or they are no longer available online. Svarna integrates these resources into a single interface that can be used without logging in, installation, or specialized training. This system provides a concordancer with KWIC marking capabilities, frequency analysis including register-by-register normalization, collocation extraction using mutual information, a dictionary of 93 Greek discourse markers providing distribution profiles, text-level analysis tools including n-grams, variants, and collocation networks, register comparison using log-ratio, regular expression search, and an optional LLM layer for pragmatic annotation and free research mode. This platform is built upon SQLite FTS5 full-text indexes provided via a FastAPI backend, deployed as Docker containers on Azure, and released under the MIT license. Source code, build scripts, and deployment configurations are publicly available on GitHub. Users can add their own corpora and deploy their own instances. This document describes the system design, corpus structure, and use cases demonstrating the various queries supported by the platform. Svarna serves as the first step in exploring available data and is expected to lay the foundation for more comprehensive research in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。