arXiv:2604.09644cs.CYcs.AI2026-04

用多模态不一致性检测识别企业夸大AI能力的虚假宣传。

Detecting Corporate AI-Washing via Cross-Modal Semantic Inconsistency Learning

  • 通过文本、图像、视频三模态比对,判断财报与披露内容是否自洽。
  • 在8.8万组数据上实现F1 0.882,比单文本方法提升17.4个百分点。
  • 适合监管机构快速筛查上市公司AI宣传真实性。

企业利用生成式AI广泛传播夸大或虚构的AI能力信息,已构成资本市场信息披露完整性的系统性威胁。现有检测方法依赖单一文本频率分析,易受对抗性改写和跨渠道模糊化影响。本文提出AWASH框架,将AI洗白检测重构为跨模态事实推理任务,并构建首个大规模三模态基准AW-Bench,包含4892家A股上市公司2019Q1至2025Q2的88,412组对齐财报文本、披露图片与业绩会视频数据。提出跨模态不一致性检测(CMID)网络,融合三模态编码器、结构化自然语言蕴含模块及操作性证据层,交叉验证AI声明与可核实物理证据(专利申请轨迹、特定人才招聘、算力基础设施代理)。在六种基线对比中,CMID达到F1 0.882、AUC-ROC 0.921,优于最强文本基线17.4个百分点,优于最新多模态对手11.3个百分点。经14名监管分析师预注册用户研究验证,CMID生成的证据报告使案例审查时间减少43%,真阳性检出率提升28%。结果证实结构化多模态推理在大规模企业披露监控中的技术优势与实际应用价值。

原文摘要 · Abstract (English)

Corporate AI-washing-the strategic misrepresentation of AI capabilities via exaggerated or fabricated cross-channel disclosures-has emerged as a systemic threat to capital market information integrity with the widespread adoption of generative AI. Existing detection methods rely on single-modal text frequency analysis, suffering from vulnerability to adversarial reformulation and cross-channel obfuscation. This paper presents AWASH, a multimodal framework that redefines AI-washing detection as cross-modal claim-evidence reasoning (instead of surface-level similarity measurement), built on AW-Bench-the first large-scale trimodal benchmark for this task, including 88412 aligned annual report text, disclosure image, and earnings call video triplets from 4892 A-share listed firms during 2019Q1-2025Q2. We propose the Cross-Modal Inconsistency Detection (CMID) network, integrating a tri-modal encoder, a structured natural language inference module for claim-evidence entailment reasoning, and an operational grounding layer that cross-validates AI claims against verifiable physical evidence (patent filing trajectories, AI-specific talent recruitment, compute infrastructure proxies). Evaluated against six competitive baselines, CMID achieves an F1 score of 0.882 and an AUC-ROC of 0.921, outperforming the strongest text-only baseline by 17.4 percentage points and the latest multimodal competitor by 11.3 percentage points. A pre-registered user study with 14 regulatory analysts verifies that CMID-generated evidence reports cut case review time by 43% while increasing true positive detection rates by 28%. These findings confirm the technical superiority and practical applicability of structured multimodal reasoning for large-scale corporate disclosure surveillance.

AI洗白多模态监管科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。