arXiv:2506.01938cs.CL2025-06

构建污水管理多语言命名实体识别基准,支持智能决策。

Novel Benchmark for NER in the Wastewater and Stormwater Domain

  • 构建法语-意大利语污水管理领域文本语料库
  • 评估主流NER方法性能,提供可靠基线
  • 探索跨语言自动标注,便于扩展至新语言

有效的污水处理与雨水管理对城市可持续性和环境保护至关重要。由于领域术语复杂且涉及多语言环境,从报告和法规中提取结构化知识极具挑战。本文聚焦于特定领域的命名实体识别(NER),作为支持决策的信息抽取第一步。建立多语言基准对于评估相关方法至关重要。本研究构建了法语-意大利语的污水管理领域文本语料库,评估了当前最先进的NER方法(包括基于大模型的方法),为未来策略提供可靠基线,并探索了自动化标注投影技术,以实现语料库向新语言的扩展。

原文摘要 · Abstract (English)

Effective wastewater and stormwater management is essential for urban sustainability and environmental protection. Extracting structured knowledge from reports and regulations is challenging due to domainspecific terminology and multilingual contexts. This work focuses on domain-specific Named Entity Recognition (NER) as a first step towards effective relation and information extraction to support decision making. A multilingual benchmark is crucial for evaluating these methods. This study develops a French-Italian domain-specific text corpus for wastewater management. It evaluates state-of-the-art NER methods, including LLM-based approaches, to provide a reliable baseline for future strategies and explores automated annotation projection in view of an extension of the corpus to new languages.

命名实体识别多语言水环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。