arXiv:2603.00620cs.CL2026-03被引 1

解决多语言NLP中海量语言元数据管理难题

QQ: A Language Metadata Toolkit for Multilingual NLP

  • 构建语言变体、脚本、地区等关联图谱
  • 支持标识符归一化与跨资源语言发现
  • 适合需要规范语言标注的研究者使用

多语言自然语言处理研究涉及数百甚至上千种语言,语言元数据的管理、发现和报告已成为普遍挑战。我们提出QQ,一个元数据工具包与浏览器探索器。QQ将多种语言元数据来源整合为包含语言变体、脚本、地区、标识符、名称及关系的图谱,并通过Python API、命令行接口和基于浏览器的探索器提供访问。用户可进行标识符归一化、获取元数据、遍历关系,以及发现哪些外部资源包含特定语言。我们在三个工作流中验证了其有效性:HuggingFace Hub的审计、不同标识系统资源的关联,以及生成可复现的语言报告表格。QQ通过版本控制、开放格式和可重用接口,支持符合FAIR原则的元数据实践。

原文摘要 · Abstract (English)

Multilingual NLP research increasingly involves hundreds or thousands of languages across different datasets. Managing, discovering, and reporting language metadata becomes a common hurdle at these scales. We present QQ, a metadata toolkit and browser explorer. QQ compiles language metadata sources into a graph of language varieties, scripts, regions, identifiers, names, and relations, and exposes it through a Python API, a command-line interface, and a browser-based explorer. Users can normalize identifiers, retrieve metadata, traverse relations, and discover which external resources contain a language. We demonstrate QQ on three workflows: an audit of the HuggingFace Hub, linking resources that use different identifier systems, and generating reproducible language-reporting tables. QQ supports FAIR-oriented metadata practices through versioning, open formats, and reusable interfaces.

多语言NLP元数据管理语言识别工具包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。