arXiv:2506.06180cs.CL2025-06被引 3

用小模型精调实现高精度语音诈骗检测,效果媲美大模型。

Detecting Voice Phishing with Precision: Fine-Tuning Small Language Models

  • 用专家设计的评估标准+思维链提示,精调Llama3小模型。
  • 在挑战性数据集上,性能超越其他小模型,接近GPT-4水平。
  • 首次构建语音诈骗语料库,适用于安全与反诈研究者。

我们通过精调Llama3这一代表性开源小语言模型,构建语音诈骗(VP)检测器。在提示中引入精心设计的VP评估标准,并应用思维链(CoT)技术。为评估模型鲁棒性并凸显性能差异,我们构建了对抗性测试数据集,使模型处于高难度场景。此外,为弥补现有语音诈骗语料不足,我们基于已有或新型诈骗手法生成语料。实验对比了是否加入评估标准、是否使用CoT,或二者结合的情况。结果表明,使用包含评估标准的提示进行微调的Llama3-8B模型,在小模型中表现最佳,其性能可媲美基于GPT-4的检测器。这说明对小模型而言,将人类专家知识嵌入提示比单纯使用思维链更有效。

原文摘要 · Abstract (English)

We develop a voice phishing (VP) detector by fine-tuning Llama3, a representative open-source, small language model (LM). In the prompt, we provide carefully-designed VP evaluation criteria and apply the Chain-of-Thought (CoT) technique. To evaluate the robustness of LMs and highlight differences in their performance, we construct an adversarial test dataset that places the models under challenging conditions. Moreover, to address the lack of VP transcripts, we create transcripts by referencing existing or new types of VP techniques. We compare cases where evaluation criteria are included, the CoT technique is applied, or both are used together. In the experiment, our results show that the Llama3-8B model, fine-tuned with a dataset that includes a prompt with VP evaluation criteria, yields the best performance among small LMs and is comparable to that of a GPT-4-based VP detector. These findings indicate that incorporating human expert knowledge into the prompt is more effective than using the CoT technique for small LMs in VP detection.

语音诈骗小模型提示工程安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。