首个面向电信诈骗的音文慢思考数据集,助力智能反诈研究。
TeleAntiFraud-28k: An Audio-Text Slow-Thinking Dataset for Telecom Fraud Detection
- 通过ASR转录+TTS重生成,构建隐私保护的真实对话文本
- 用大模型自指导扩充场景,覆盖28,511对语音文本对,含欺诈推理标注
- 支持诈骗类型识别、场景分类等任务,适合反诈模型训练与评测
电信诈骗检测因缺乏高质量多模态训练数据而面临挑战,尤其是融合音频信号与推理导向文本分析的数据稀缺。为此,我们提出TeleAntiFraud-28k,首个开源的音文慢思考数据集,专为自动化电信诈骗分析设计。数据集通过三种策略构建:(1) 使用自动语音识别(ASR)转录通话记录并匿名化原始音频,结合文本到语音(TTS)模型再生,实现隐私保护下的真实文本生成;(2) 基于大语言模型(LLM)的自指导采样,增强语义多样性,扩大真实场景覆盖;(3) 多智能体对抗合成,模拟新兴诈骗手法,基于预设通信场景与诈骗类型。最终数据集包含28,511个经过严格处理的语音-文本对,附带详细的欺诈推理标注,涵盖三类任务:场景分类、欺诈检测、诈骗类型分类。同时构建TeleAntiFraud-Bench标准化评估基准,采用比例抽样样本以系统化测试模型性能。我们还提供在混合真实/合成数据上训练的生产优化监督微调(SFT)模型,并开源数据处理框架,推动社区共建。本工作为多模态反欺诈研究建立基础框架,解决数据隐私与场景多样性难题。项目将发布于 https://github.com/JimmyMa99/TeleAntiFraud。
原文摘要 · Abstract (English)
The detection of telecom fraud faces significant challenges due to the lack of high-quality multimodal training data that integrates audio signals with reasoning-oriented textual analysis. To address this gap, we present TeleAntiFraud-28k, the first open-source audio-text slow-thinking dataset specifically designed for automated telecom fraud analysis. Our dataset is constructed through three strategies: (1) Privacy-preserved text-truth sample generation using automatically speech recognition (ASR)-transcribed call recordings (with anonymized original audio), ensuring real-world consistency through text-to-speech (TTS) model regeneration; (2) Semantic enhancement via large language model (LLM)-based self-instruction sampling on authentic ASR outputs to expand scenario coverage; (3) Multi-agent adversarial synthesis that simulates emerging fraud tactics through predefined communication scenarios and fraud typologies. The generated dataset contains 28,511 rigorously processed speech-text pairs, complete with detailed annotations for fraud reasoning. The dataset is divided into three tasks: scenario classification, fraud detection, fraud type classification. Furthermore, we construct TeleAntiFraud-Bench, a standardized evaluation benchmark comprising proportionally sampled instances from the dataset, to facilitate systematic testing of model performance on telecom fraud detection tasks. We also contribute a production-optimized supervised fine-tuning (SFT) model trained on hybrid real/synthetic data, while open-sourcing the data processing framework to enable community-driven dataset expansion. This work establishes a foundational framework for multimodal anti-fraud research while addressing critical challenges in data privacy and scenario diversity. The project will be released at https://github.com/JimmyMa99/TeleAntiFraud.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。