LLMs生成的研究方法偏好单一,易导致研究思路固化。
Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods

- 用研究问题引导LLM生成方法建议,对比真实论文方法库
- 方法选择范围收缩至59-96种,远少于真实论文的1232种
- 不同LLM间方法推荐高度相似,适合需拓宽思路的研究者
大型语言模型(LLMs)被越来越多地用于指导研究方法,但其在最少提示下的默认方法倾向尚不明确。本研究对GPT-5.1、Gemini 3 Pro和DeepSeek-V3.2分别输入1000篇近期arXiv计算机科学论文中提取的研究问题,并将生成的方法建议与原始论文中的实验方法库进行对比。由于仅提供研究问题,所测差异反映的是初始建议而非最优性。我们从双方提取结构化方法特征,映射到统一分类体系,量化多个维度的偏差,包括模型提供方、数据集任务类型、评估指标类型。其中提供方选择偏差最强,杰恩-申诺分离度比其他维度高3-5倍。其他/学术单次使用模型被低估23-24个百分点,而重复使用的学术/社区模型略高4-6个百分点。此外,LLMs建议的方法范围显著更窄:有效模型实体数从1232降至59-96;跨模型排名相关性(0.55-0.68)普遍高于模型与论文间的相关性(0.33-0.56),表明偏差在模型间高度共享。流行度基线、BM25检索校准和论文级相似性测试均表明输出是针对查询的响应,但被筛选在较窄选项内。依赖LLM建议而未交叉验证的研究者,可能无意中缩小了方法搜索空间,陷入更集中的默认倾向。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear. Here, we prompt GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 with an LLM-extracted research question from each of 1,000 recent arXiv computer-science papers and compare the resulting methodology suggestions against a paper-derived experimental inventory. Since we provide only the research question, the differences we measure reflect initial suggestions and not how optimal those suggestions are. We extract structured method features from both sources, map them into a shared taxonomy, and quantify divergence across multiple taxonomy dimensions including model provider, dataset task type, and evaluation metric type. The strongest imbalance appears in provider choice, with Jensen-Shannon divergence about 3-5x larger than any other taxonomy dimension. Other/Academic single-occurrence models are underrepresented by 23-24 percentage points, while reused academic/community models are slightly overrepresented (4-6pp). LLMs also suggest a much narrower range of methods overall: the effective number of model entities contracts from 1,232 to 59-96, and inter-LLM rank correlations (0.55-0.68) generally exceed LLM-to-paper correlations (0.33-0.56), so the distortions are largely shared across models. Popularity baselines, BM25 retrieval calibration, and paper-level similarity tests confirm that the outputs are query-specific responses, but filtered through a narrower set of options. Researchers who rely on LLM suggestions without cross-checking therefore risk narrowing their methodological search space toward a more concentrated default.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。