arXiv:2606.30312cs.CL2026-06

构建多语言对话数据集,用于检测个人身份信息。

DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

论文配图:DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information
图 1 · 摘自论文原文
  • 用大模型生成8类场景的合成对话,覆盖11种语言
  • 每条对话经语音合成与转录,实现文本与语音对齐
  • 提供标注数据和基线模型,适合隐私保护研究者使用

医疗、社会科学等领域的对话数据是研究与自动分析的重要资源,但数据共享需去除个人身份与敏感信息以保护隐私。为支持自动化去标识化系统的发展与评估,我们提出DialogPII,一个包含合成对话与语音转录的多语言数据集。该数据集涵盖8种交互场景(如急诊电话、医疗问诊、心理咨询、保险沟通等)、19类实体类型及11种语言(英语、阿拉伯语、芬兰语、法语、德语、印地语、意大利语、波兰语、葡萄牙语、西班牙语、土耳其语)。对话通过大语言模型半自动生成,经人工校验确保真实性和多样性,并本地化至具体国家与城市背景。所有对话经文本转语音合成后,使用Whisper进行转录,通过自动投影与人工修正完成标注,实现跨语言的文本与语音资源对齐。我们还发布了多语言命名实体识别基线模型,并通过标注者一致性分析、翻译质量评估、标注投影验证及基于Transformer的序列标注模型实验,提供技术验证。

原文摘要 · Abstract (English)

Conversational data collected in domains such as healthcare or social sciences is a valuable resource for research and automated analysis. However, responsible data sharing requires the detection and removal of personally identifiable and sensitive information to protect individual privacy. To support the development and evaluation of automatic de-identification systems, we present DialogPII, a multilingual dataset of synthetic dialogs and speech-derived transcripts for personal information detection. DialogPII covers eight interaction scenarios (emergency calls, medical anamnesis interviews, therapy sessions, insurance communication, customer support, clinical interviews regarding an AI-supported dashboard, police reports, and group therapy discussions), 19 entity types, and 11 languages (English, Arabic, Finnish, French, German, Hindi, Italian, Polish, Portuguese, Spanish, and Turkish). Dialogs were generated semi-automatically using large language models, manually curated for plausibility and diversity, and localized to country- and city-specific contexts. All dialogs were additionally converted to speech via text-to-speech synthesis, transcribed with Whisper, and annotated through automatic projection and manual correction, yielding aligned written and speech-derived resources across all languages. We further release baseline multilingual named entity recognition models and provide technical validation through inter-annotator agreement analysis, translation quality evaluation, annotation projection assessment, and benchmark experiments with transformer-based sequence labeling models.

隐私保护多语言对话数据去标识化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。