arXiv:2604.09016cs.LGcs.AI2026-04

从电报平台提取数据并匿名化敏感信息,助力合规的网络犯罪研究。

Identification and Anonymization of Named Entities in Unstructured Information Sources for Use in Social Engineering Detection

  • 融合语音增强与Transformer模型,提升多模态信息中的实体识别准确率。
  • 所提NER方案在检测敏感信息时达到最高F1分数,音频转录使用Parakeet表现最优。
  • 兼顾数据结构完整性和隐私保护,适合法律合规下的网络安全研究。

本研究针对符合《通用数据保护条例》(GDPR)及西班牙《1995年刑事法典组织法》等法规要求的前提下,构建网络犯罪分析数据集的挑战。提出一套从Telegram平台收集文本、音频和图像信息的系统,集成信号增强技术的语音转文字模型,并评估了包括Microsoft Presidio和基于Transformer架构的自研AI模型在内的多种命名实体识别(NER)方案。实验表明,Parakeet在音频转录任务中表现最佳,所提NER方案在敏感信息检测中取得最高F1值。同时,提出了可量化评估数据结构一致性与个人隐私保护程度的匿名化指标,支持在现行法律框架内开展网络安全研究。

原文摘要 · Abstract (English)

This study addresses the challenge of creating datasets for cybercrime analysis while complying with the requirements of regulations such as the General Data Protection Regulation (GDPR) and Organic Law 10/1995 of the Penal Code. To this end, a system is proposed for collecting information from the Telegram platform, including text, audio, and images; the implementation of speech-to-text transcription models incorporating signal enhancement techniques; and the evaluation of different Named Entity Recognition (NER) solutions, including Microsoft Presidio and AI models designed using a transformer-based architecture. Experimental results indicate that Parakeet achieves the best performance in audio transcription, while the proposed NER solutions achieve the highest f1-score values in detecting sensitive information. In addition, anonymization metrics are presented that allow evaluation of the preservation of structural coherence in the data, while simultaneously guaranteeing the protection of personal information and supporting cybersecurity research within the current legal framework.

隐私保护实体识别数据匿名化社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。