arXiv:2512.01354cs.AIcs.CL2025-12被引 1

模拟人类认知局限生成带瑕疵的合成数据,缓解模型崩溃问题

The Necessity of Imperfection:Reversing Model Collapse via Simulating Cognitive Boundedness

  • 用认知状态解码器和扰动算子逆向生成带人类特性的文本
  • 生成文本与真人文本相似度达JS散度0.0614,显著优于常规LLM输出
  • 在股市压力测试中降低47.4%最大回撤,适合金融建模等高可靠性场景

尽管合成数据被广泛提倡,但当前主流生产范式——追求统计平滑性——系统性地去除了人类文本中基于认知的长尾不规则特征。长期训练于这类统计最优但认知贫乏的数据会加速模型崩溃。本文提出范式转变:不模仿数据表面属性,而是模拟生成人类文本的认知过程。提出提示驱动认知计算框架(PMCSF),核心为认知状态解码器(CSD),将无结构文本逆向解析为结构化认知向量;以及认知文本编码器(CTE),通过数学定义的认知扰动算子将这些状态重构为富含人类典型瑕疵的文本。框架通过两阶段评估验证:首先,在认知编解码验证中,CTE生成文本与真人文本的杰恩-申诺尔散度为0.0614(标准LLM输出为0.4431),通过双盲专业媒体评审,跨异构模型的认知特征一致性系数ICC > 0.9;其次,在功能增益评估中,同构压力测试显示,融入CTE生成数据的策略在2015年股灾期间最大回撤降低47.4%,实现8.6%防御型阿尔法,超出交易成本33倍。研究证明,建模人类认知局限——而非复制数据表层——可产生具有真实功能收益的合成数据,为解决人工智能数据崩溃危机提供可行技术路径。

原文摘要 · Abstract (English)

Although synthetic data is widely promoted as a remedy, its prevailing production paradigm -- one optimizing for statistical smoothness -- systematically removes the long-tail, cognitively grounded irregularities that characterize human text. Prolonged training on such statistically optimal but cognitively impoverished data accelerates model collapse. This paper proposes a paradigm shift: instead of imitating the surface properties of data, we simulate the cognitive processes that generate human text. We introduce the Prompt-driven Cognitive Computing Framework (PMCSF), whose core consists of a Cognitive State Decoder (CSD) that reverse-engineers unstructured text into structured cognitive vectors, and a Cognitive Text Encoder (CTE) that re-materializes these states into text enriched with human-typical imperfections via mathematically defined Cognitive Perturbation Operators. The framework is validated through a two-stage objective evaluation pipeline. First, in cognitive codec verification, CTE text yields a Jensen-Shannon divergence of 0.0614 from human text (vs. 0.4431 for standard LLM output), passes double-blind professional media review, and achieves an intraclass correlation coefficient ICC > 0.9 for cognitive profile alignment across heterogeneous models. Second, in functional gain evaluation, isomorphic stress tests in the A-share market show that strategies incorporating CTE-generated data reduce maximum drawdown by 47.4% during the 2015 crash and deliver 8.6% Defensive Alpha, exceeding transaction costs by a factor of 33. Our findings demonstrate that modelling human cognitive limitations -- not copying surface data -- enables synthetic data with genuine functional gain, offering a viable technical pathway toward resolving the AI data-collapse crisis.

合成数据认知建模模型崩溃金融应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。