构建北欧四国客服自助语料库,助力多语言NLP研究。
Telenor Nordics Customer Service self-help corpus

- 从四大运营商官网收集并人工验证1122份文档
- 含27万词、188万字符,覆盖网络硬件到账单管理
- 适合做跨语言检索与客服智能体的研究者使用
本文介绍了一个多语言客服自助语料库,包含芬兰语、丹麦语、挪威语和瑞典语的1,122份经人工验证的文档,总计274,599词、1,884,833字符。数据源自北欧四家电信运营商的公开自助页面,经大模型与人工联合筛选,剔除个人身份信息并确保相关性。当前北欧语言领域特定数据集稀缺,尤其在客服这一日益重要的领域,对检索增强生成、跨语言迁移学习及新型代理服务架构具有价值。分析显示,不同运营商间文档长度与结构差异显著,体现不同的编辑策略,且主题覆盖广泛,涵盖网络硬件、移动服务、电视与流媒体、账单及账户管理。该数据集以CC-BY-NC-SA-4.0许可公开,可于https://zenodo.org/records/20732652获取,旨在支持北欧NLP与信息检索领域的可复现研究。
原文摘要 · Abstract (English)
This paper presents a multilingual customer service self-help corpus comprising 1,122 manually validated documents in Finnish, Danish, Norwegian, and Swedish, totaling 274,599 words and 1,884,833 characters. The documents have been sourced from the public self-help pages of four Nordic telecommunications operators and subsequently filtered for person-identifiable information and relevance through a combined LLM and human annotation pipeline. Domain-specific datasets for Nordic languages remain scarce, particularly in customer service: a domain of growing importance for retrieval-augmented generation, cross-lingual transfer learning, and emerging agent-based service architectures. An analysis of the corpus reveals substantial variation in document length and structure across operators, reflecting distinct editorial strategies, as well as broad topical coverage spanning network hardware, mobile services, TV and streaming, billing, and account management. The dataset is publicly available under a CC-BY-NC-SA-4.0 license at https://zenodo.org/records/20732652, intended to support reproducible research in Nordic NLP and information retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。