arXiv:2606.16659cs.CL2026-06

构建遮蔽链接的短信骗术检测基准,评估大模型真实判断能力。

FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection

论文配图:FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
图 1 · 摘自论文原文
  • 用遮蔽网址的短信-网页链路评测模型,避免依赖域名信誉
  • 9个智能体在699条案例中误判率高,良性案例召回率低
  • 适合研究对抗性攻击、可解释性及可信AI的开发者使用

短信诈骗正呈现跨渠道特征:短信引导用户访问网页,最终风险取决于短信内容与页面信息及用户操作的一致性。现有评估要么仅关注短信本身,要么暴露网址和域名线索,使模型依赖声誉捷径。为此,我们提出FraudSMSWalker,一个针对遮蔽网址的短信到网页欺诈判断控制型基准。该基准包含699条双语链路,涵盖10种服务场景,其中332条为欺诈案例,367条为良性案例。模型可见输入仅为短信上下文与净化后的网页证据,原始网址、主机、域名、IP、跳转路径及声誉元数据均被隐藏。基准还包含10类难以区分的良性案例,其页面含登录、支付、验证或账户管理元素,在特定服务背景下合理但常出现在诈骗流程中。我们在遮蔽浏览器代理协议下评估9个网络智能体,并进行网址可见性消融实验。结果显示,当前智能体虽能识别可疑线索,却难以保持良性召回率,常基于弱证据做出阳性预测。这些发现表明,FraudSMSWalker是衡量智能体在无直接声誉捷径时能否做出准确且有证据支撑的欺诈判断的关键基准。相关代码与数据集可通过匿名链接获取。

原文摘要 · Abstract (English)

SMS fraud is increasingly cross-channel: a message directs the user to a webpage, and the final risk depends on how the SMS claim aligns with the page content and requested user action. However, existing evaluations either focus on message-only smishing classification or expose URL and domain cues that allow models to rely on reputation shortcuts. To address this gap, we introduce \textbf{FraudSMSWalker}, a controlled benchmark for URL-masked SMS-to-webpage fraud judgment. FraudSMSWalker contains 699 bilingual chains, including 332 fraudulent and 367 benign cases, across ten service scenarios. The model-visible input consists of the SMS context and sanitized webpage evidence, while raw URLs, hosts, domains, IPs, redirects, and reputation metadata are withheld. The benchmark further includes hard benign cases whose pages contain login, payment, verification, or account-management elements that are plausible under the service context but also appear in scam flows. We evaluate nine web agents under masked browser-agent protocols and conduct URL-visibility ablations. The results show that current agents can detect suspicious cues, but struggle to preserve benign recall and often produce positive predictions that are weakly supported by the observed evidence. These findings position FraudSMSWalker as a benchmark for measuring whether web agents can make fraud judgments that remain both accurate and evidence-grounded when direct reputation shortcuts are suppressed. The associated code and dataset are accessible at the \href{https://anonymous.4open.science/w/FraudMessageWalker-Bench}{anonymous link}.

安全检测智能体评测反欺诈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。