arXiv:2507.02958cs.CL2025-07被引 3

开源9万+真实客服通话数据集,支持AI客服研发。

Real-World En Call Center Transcripts Dataset with PII Redaction

  • 构建大规模英文客服对话数据集,覆盖多国口音
  • 9.17万通通话,超1万小时音频,已脱敏可商用
  • 填补公开客服数据空白,适合非商业研究使用

我们推出CallCenterEN,一个大规模(91,706通对话,共10,448小时音频)的英语真实世界客服通话文本数据集,旨在支持客户服务与销售AI系统的研究与开发。该数据集是目前最大规模的开源同类数据集,包含来自印度、菲律宾和美国口音的来电与外呼对话。所有内容均提供高质量、已去除个人身份信息(PII)的人类可读转录文本,确保符合全球数据保护法规。由于生物特征隐私顾虑,音频未在公开发布中提供。鉴于真实世界客服数据集的稀缺性,CallCenterEN填补了语音识别语料库中的关键空白,采用CC BY-NC 4.0许可,仅限非商业研究使用。

原文摘要 · Abstract (English)

We introduce CallCenterEN, a large-scale (91,706 conversations, corresponding to 10448 audio hours), real-world English call center transcript dataset designed to support research and development in customer support and sales AI systems. This is the largest release to-date of open source call center transcript data of this kind. The dataset includes inbound and outbound calls between agents and customers, with accents from India, the Philippines and the United States. The dataset includes high-quality, PII-redacted human-readable transcriptions. All personally identifiable information (PII) has been rigorously removed to ensure compliance with global data protection laws. The audio is not included in the public release due to biometric privacy concerns. Given the scarcity of publicly available real-world call center datasets, CallCenterEN fills a critical gap in the landscape of available ASR corpora, and is released under a CC BY-NC 4.0 license for non-commercial research use.

客服数据语音识别数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。