arXiv:2501.05606cs.CLcs.IR2025-01被引 7

统一语言资源元数据,让多源数据更易查询与使用

Harmonizing Metadata of Language Resources for Enhanced Querying and Accessibility

  • 用链接数据整合多源语言资源元数据
  • 真实用户查询验证显示多数需求可满足
  • 适合语言资源管理与数据开放者参考

本文针对多源语言资源(LRs)元数据的不一致性问题,利用链接数据和RDF技术,将多个来源的数据集成到基于DCAT和META-SHARE OWL本体的统一模型中。通过新开发的Linghub门户,支持文本搜索、分面浏览及高级SPARQL查询。基于语料库邮件列表(CML)的真实用户查询评估表明,尽管仍存在一些局限性,但多数用户需求可被有效响应。研究揭示了显著的元数据问题,倡导采用开放词汇表与标准以提升元数据一致性。初步研究表明,基于API的资源访问方式有助于机器可用性和数据子集提取,为更高效、标准化的语言资源利用铺平道路。

原文摘要 · Abstract (English)

This paper addresses the harmonization of metadata from diverse repositories of language resources (LRs). Leveraging linked data and RDF techniques, we integrate data from multiple sources into a unified model based on DCAT and META-SHARE OWL ontology. Our methodology supports text-based search, faceted browsing, and advanced SPARQL queries through Linghub, a newly developed portal. Real user queries from the Corpora Mailing List (CML) were evaluated to assess Linghub capability to satisfy actual user needs. Results indicate that while some limitations persist, many user requests can be successfully addressed. The study highlights significant metadata issues and advocates for adherence to open vocabularies and standards to enhance metadata harmonization. This initial research underscores the importance of API-based access to LRs, promoting machine usability and data subset extraction for specific purposes, paving the way for more efficient and standardized LR utilization.

元数据语言资源知识图谱开放数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。