arXiv:2511.12658cs.CV2025-11NeurIPS被引 1

用数学分解法生成更真实的伪造文本图像,提升模型泛化能力

Toward Real-world Text Image Forgery Localization: Structured and Interpretable Data Synthesis

  • 基于傅里叶思想建模人类编辑行为,实现可解释的伪造数据生成
  • 从16,750个真实伪造案例中提取规律,构建分层行为分布模型
  • 生成数据使模型在多个真实场景测试中表现显著提升,适合安全检测研究者

现有文本图像伪造定位方法因真实数据规模有限且合成数据与真实世界存在分布差异,泛化能力差。为此,本文提出基于傅里叶级数的伪造合成框架(FSTS),通过结构化流程收集5种典型伪造类型共16,750个真实编辑实例,记录视频、PSD文件及操作日志等多格式编辑痕迹。分析个体与群体层面的行为模式,构建分层建模框架:每个伪造参数以基函数组合表示,群体分布通过行为聚合构建。该形式受傅里叶级数启发,支持可解释的逼近。采样该分布生成多样且逼真的训练数据,显著提升模型在四个评估协议下对真实数据的泛化性能。数据集已开源。

原文摘要 · Abstract (English)

Existing Text Image Forgery Localization (T-IFL) methods often suffer from poor generalization due to the limited scale of real-world datasets and the distribution gap caused by synthetic data that fails to capture the complexity of real-world tampering. To tackle this issue, we propose Fourier Series-based Tampering Synthesis (FSTS), a structured and interpretable framework for synthesizing tampered text images. FSTS first collects 16,750 real-world tampering instances from five representative tampering types, using a structured pipeline that records human-performed editing traces via multi-format logs (e.g., video, PSD, and editing logs). By analyzing these collected parameters and identifying recurring behavioral patterns at both individual and population levels, we formulate a hierarchical modeling framework. Specifically, each individual tampering parameter is represented as a compact combination of basis operation-parameter configurations, while the population-level distribution is constructed by aggregating these behaviors. Since this formulation draws inspiration from the Fourier series, it enables an interpretable approximation using basis functions and their learned weights. By sampling from this modeled distribution, FSTS synthesizes diverse and realistic training data that better reflect real-world forgery traces. Extensive experiments across four evaluation protocols demonstrate that models trained with FSTS data achieve significantly improved generalization on real-world datasets. Dataset is available at \href{https://github.com/ZeqinYu/FSTS}{Project Page}.

伪造检测数据合成可解释性图像安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。