首个土耳其诈骗电话多模态数据集,验证语音转文本更有效
Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams
- 用7个大模型对比音频与转录文本的检测效果
- 转录文本比直接处理音频准确率更高,人工修正无明显提升
- 为低资源语言反诈研究提供数据与方法参考
诈骗电话侵害全球弱势群体,但现有检测研究几乎仅聚焦英语等高资源语言。在土耳其等低资源语境下,标注数据稀缺,技术防御能力有限。本研究探索大语言模型(LLMs)在土耳其语诈骗检测中的应用,首次发布包含100对对齐音频-转录文本的公开多模态数据集。评估了三个模型族共七种LLM:Gemini 2.5(Flash、Flash-Lite、Pro)、GPT-4o、Qwen(Max、Plus、Turbo),在原始音频、自动语音识别转录、母语者校正转录三种输入条件下表现。结果表明,基于转录文本的输入始终优于直接音频处理,而人工校正与未校正转录性能相当。该工作凸显低资源语言与真实威胁场景下具文化与语言包容性的AI安全研究的紧迫性,推动更鲁棒的多模态反欺诈系统发展。
原文摘要 · Abstract (English)
Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages. In low-resource settings such as Turkish, detection is especially difficult, as annotated data is scarce and technological defenses remain limited. This research investigates how large language models (LLMs) can support scam detection in Turkish by introducing the first public multi-modal dataset of 100 aligned audio-transcript pairs of scam and benign conversations. We evaluate seven LLMs spanning three model families: Gemini 2.5 (Flash, Flash-Lite, Pro), GPT-4o, and Qwen (Max, Plus, Turbo), under three input conditions: raw audio, automatic speech-to-text transcripts, and transcripts refined by a native speaker. Our results suggest that transcript-based inputs consistently outperform direct audio processing, while human-corrected and uncorrected transcripts perform comparably. By centering a low-resource language and real world threat, this work highlights the urgent need for culturally and linguistically inclusive AI safety research and more robust multi-modal systems for fraud prevention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。