用合成数据构建对话欺诈检测框架,融合语音与文本多模态分析。
Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection

- 构建仿真理赔场景的多模态数据生成流程,含对话与双人音频。
- 通过语音识别与说话人分离,识别重复叙述和声音特征异常。
- 适合保险风控、AI反欺诈团队使用,提供可复现的检测基线。
保险欺诈造成重大财务损失和运营低效,推高保费并损害诚信投保人信任。在理赔初审(FNOL)阶段早期发现仍具挑战性。现有方法多依赖私有文本数据,限制了融合语言、行为与说话人特征的多模态方法发展。本文提出一种合成多模态框架,模拟真实FNOL情境,生成代理-客户对话文本与双人音频,经自动语音识别(ASR)与说话人分离(diarisation)处理。下游模块结合命名实体识别(NER)、基于正则的特征提取、大模型-检索增强生成(LLM-RAG)检索及说话人嵌入,构建规则驱动的风险评分系统,用于标记叙事重复、结构矛盾以及跨案件声音复用,同时平衡敏感性与误报率。数据集验证与组件级评估显示其稳定性与迁移潜力,为超越纯文本的欺诈检测提供可复现基线。
原文摘要 · Abstract (English)
Insurance fraud imposes substantial financial losses and operational inefficiencies, raising premiums and impacting trust among legitimate policyholders. Early detection at FNOL remains a persistent challenge. Existing approaches rely largely on private, text-only datasets, limiting progress on multimodal methods that integrate linguistic, behavioural, and speaker-based indicators. We introduce a synthetic multimodal framework that replicates FNOL conditions. It generates agent-customer dialogue transcripts and two-speaker audios, performs ASR and diarisation. Downstream modules combine NER, regex-based feature extraction, LLM-RAG retrieval, and speaker embeddings in a rule-based risk score to flag narrative reuse, structural inconsistencies, and cross-case voice repetition while balancing sensitivity and false positives. Dataset validation and component-level evaluations show stability and transfer potential, offering a reproducible baseline beyond text-only fraud detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。