arXiv:2607.13078cs.CRcs.AI2026-07

揭示大模型在风控流程中缺乏真实部署证据,提出评估框架与最低证据清单。

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

  • 构建FORTE框架,分类大模型在风控中的角色
  • 发现欺诈类研究普遍缺少延迟、成本等实测数据
  • 提出部署必备的五项证据清单,指导可信落地

大模型被广泛用于欺诈检测、内容审核等安全工作流,但现有研究多聚焦模型性能,忽视其在实际流程中的运行表现。本文通过分析49个运营相关来源(18个欺诈、14个审核、17个通用鲁棒性),发现证据严重失衡:欺诈研究仅报告离线任务准确率或案例研究,无任何一篇提供单次决策延迟、成本或校准数据;而内容审核研究更常公开延迟、成本、公平性等实测结果。本文提出‘FORTE’角色-证据框架,将大模型定位为分类器、检索接口、解释生成器等六类组件,并制定最小部署证据清单,涵盖延迟预算、每决策成本、阈值设定、解释完整性与对抗压力等五项要求,明确未来研究需填补的证据缺口。

原文摘要 · Abstract (English)

LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational pipelines. This creates a practical evidence question: what would justify placing an LLM inside a live workflow with latency, cost, escalation, human-review, and adversarial-risk constraints? We address this question through a fraud-first survey of deployment evidence. We code 49 operationally relevant sources on LLM use in fraud detection, investigation support, content moderation, and cross-cutting robustness (18 fraud, 14 moderation, 17 cross-cutting), supplemented by 15 contextual references that establish the survey boundaries. These sources include systems, benchmarks, frameworks, and deployment-relevant surveys, not 49 production deployments. The main finding is an evidence imbalance. Fraud supplies the largest task-specific portion of the coded corpus. The moderation papers, however, include more explicit public evidence on latency, cost, governance, and fairness. Among the 18 fraud and investigation sources, none report clean per-decision latency, per-decision dollar cost, or calibration evidence; most report offline task performance, retrieval gains, or case-study accuracy instead. The survey contributes a role-and-evidence organizing frame, FORTE, for locating LLMs as classifiers, retrieval interfaces, explanation generators, reviewer assistants, agents, feature extractors, or escalation components. It also contributes a minimum deployment-evidence checklist covering latency budget, cost per decision, decision threshold, explanation integrity, and adversarial pressure. The resulting agenda identifies studies needed to support deployment claims for LLM-based fraud and trust-and-safety work.

大模型部署风控系统证据评估安全工作流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。