用大模型自动生成可验证复杂度的决策场景,解决人工设计慢且偏倚的问题。
Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models

- 用大模型生成结构化决策场景,并通过多重理论框架验证其复杂度。
- 4238个场景验证显示复杂度分级效果显著,跨模型一致性极高。
- 适合认知科学、AI评估等需要标准化测试场景的研究者使用。
认知决策研究依赖于多样且复杂度可控的场景,但人工制作耗时、不一致且易偏倚。本文开发了一套自动化流程,利用大语言模型生成结构化决策场景,并基于任务复杂度理论构建复合验证框架。在多个领域和复杂度层级上评估了4,238个场景,测量验证达到严格心理计量标准:五组独立模型间一致性近乎完美(组内相关系数0.997,κ=0.971);已知组效度显示各层级间差异显著(η²=0.587,所有配对p<0.001);因子分析揭示主导复杂度构念(载荷0.87–0.96),互动性为次级维度(载荷0.34)。判别效度受文本长度与复杂度强相关限制(部分相关0.86),虽影响构念纯度,但不损害层级划分功能。模型分析显示吞吐量与模板通过率负相关(r=-0.967,p=0.007,n=5),体现速度-质量权衡,主要由一高吞吐模型驱动。Llama 4 Maverick生成最快(134/分钟),但复杂场景产出不足;DeepSeek Chat V3.2在领域覆盖与模板符合率间取得平衡。系统具备优异心理计量特性,可可靠区分简单、中等、复杂三类层级,为下游人工智能认知评估提供测量基础设施。
原文摘要 · Abstract (English)
Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and biased. We developed an automated pipeline that uses LLms to generate structured decision scenarios and validates their complexity through a composite framework rooted in established task-complexity theory. We evaluated 4,238 scenarios across multiple domains and complexity tiers. Measurement validation met rigorous psychometric standards. Agreement among five independent model families was nearly perfect, with an intraclass correlation coefficient of 0.997 and a kappa of 0.971. Known-groups validity demonstrated large separation between tiers, with an eta-squared of 0.587 and all pairwise comparisons significant at p less than .001. Factor analysis revealed a dominant complexity construct, with loadings between 0.87 and 0.96 across three frameworks, while interactivity formed a weaker secondary dimension at 0.34. Discriminant validity was limited by a strong relationship between complexity and text length that persisted after controlling for tier, yielding a partial correlation of 0.86. This constrains construct purity but does not undermine the instrument's tier-grading function. Model analyses showed a negative association between throughput and schema pass rate (r = -0.967, p = .007, n = 5), suggesting a speed-quality trade-off, though largely driven by one high-throughput model. Llama 4 Maverick generated scenarios fastest at 134 per minute versus 25 for DeepSeek Chat V3.2, but underproduced complex-tier scenarios, whereas DeepSeek Chat V3.2 balanced domain coverage with high schema compliance. The system demonstrated strong psychometric properties, enabling reliable classification into Simple, Moderate, and Complex tiers and providing the measurement infrastructure needed for downstream cognitive assessment of AI systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。