arXiv:2508.17675cs.LG2025-08

用大模型生成认知测试的合成数据,解决传统方法耗时费力的问题。

Towards Synthesizing Normative Data for Cognitive Assessments Using Generative Multimodal Large Language Models

  • 用提示工程优化GPT-4o生成图像任务的文本回答。
  • 高级提示使生成结果更好区分诊断群体与人口差异。
  • 适合心理评估、临床研究者快速构建新测试数据。

认知评估需要规范性数据作为个体表现的基准。然而,基于新图像刺激开发新测试因缺乏现成的规范性数据而困难重重。传统数据收集方法成本高、耗时长且更新频率低,实用性受限。近年来,生成式多模态大语言模型(MLLMs)为从现有认知测试图像中生成合成规范性数据提供了新路径。我们研究了使用MLLMs(特别是GPT-4o和GPT-4o-mini)合成已知图像认知测试(如“饼干盗窃”图片描述任务)的文本响应的可行性。采用两种提示策略:基础指令的朴素提示与加入上下文引导的高级提示。通过嵌入分析评估生成响应在区分诊断组和人口学差异方面的能力。性能指标包括BLEU、ROUGE、BERTScore及LLM-as-a-judge评估。结果显示,高级提示生成的响应在区分诊断组和捕捉人口多样性方面优于朴素提示;表现更优的模型生成了更具真实感和多样性的内容。BERTScore在语境相似性评估中最为可靠,而BLEU对创造性输出评估效果较差。LLM-as-a-judge方法提供了有前景的初步验证结果。本研究证明,在精细化提示引导下,生成式多模态大模型可可行生成稳健的合成规范性数据,为无需传统限制地开发新型图像认知测试奠定基础。

原文摘要 · Abstract (English)

Cognitive assessments require normative data as essential benchmarks for evaluating individual performance. Hence, developing new cognitive tests based on novel image stimuli is challenging due to the lack of readily available normative data. Traditional data collection methods are costly, time-consuming, and infrequently updated, limiting their practical utility. Recent advancements in generative multimodal large language models (MLLMs) offer a new approach to generate synthetic normative data from existing cognitive test images. We investigated the feasibility of using MLLMs, specifically GPT-4o and GPT-4o-mini, to synthesize normative textual responses for established image-based cognitive assessments, such as the "Cookie Theft" picture description task. Two distinct prompting strategies-naive prompts with basic instructions and advanced prompts enriched with contextual guidance-were evaluated. Responses were analyzed using embeddings to assess their capacity to distinguish diagnostic groups and demographic variations. Performance metrics included BLEU, ROUGE, BERTScore, and an LLM-as-a-judge evaluation. Advanced prompting strategies produced synthetic responses that more effectively distinguished between diagnostic groups and captured demographic diversity compared to naive prompts. Superior models generated responses exhibiting higher realism and diversity. BERTScore emerged as the most reliable metric for contextual similarity assessment, while BLEU was less effective for evaluating creative outputs. The LLM-as-a-judge approach provided promising preliminary validation results. Our study demonstrates that generative multimodal LLMs, guided by refined prompting methods, can feasibly generate robust synthetic normative data for existing cognitive tests, thereby laying the groundwork for developing novel image-based cognitive assessments without the traditional limitations.

认知评估生成模型合成数据提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。