用有限算力精选优化研究想法,提升多样性与可执行性。
Budgeted Subset Refinement for Execution-Aware LLM Research Ideation
- 只精炼部分候选想法而非全部,节省资源
- 多样性优先的精炼策略产出最多高质量无重复想法
- 适合需要高效筛选研究创意的科研团队使用
大语言模型能生成看似新颖的研究想法,但近期研究表明这些想法缺乏多样性,难以可靠评估,且难转化为有效项目。本文评估了一个预执行阶段的控制性代理基准:在噪声较大的LLM生成想法池中,如何在有限精炼资源下分配算力,构建更优、更多样、更具执行意识的想法组合?提出「预算化子集精炼」方法,仅对选定子集进行精炼。在10个随机种子和10个研究创意环境下的统一评估中,仅生成或重排序无法产生符合标准的非重复强想法,而精炼是必要条件。均匀精炼虽能生成强个体想法,但并非最优整体算力分配。随机选k法为低成本强基线,而多样性感知的MMR-k精炼表现最佳:最高非重复强想法产出率、最低重复率,以及最优单位成本效率。对72项样本的盲评验证显示,精炼效果在不同模型族间具鲁棒性,但具体排名因评审者而异。结果表明,应将LLM研究创意系统视为预算分配支持系统,而非仅想法生成器。结论限于代理评分的组合质量,不替代专家评审或执行验证。
原文摘要 · Abstract (English)
Large language models (LLMs) can generate research ideas that appear novel to expert reviewers, but recent work also shows that such ideas often lack diversity, are difficult for LLMs to evaluate reliably, and may fail to translate into strong executed projects. This paper evaluates a controlled proxy benchmark for a pre-execution scaffolding problem: given a noisy pool of LLM-generated research ideas, how should a system allocate limited refinement effort to construct a stronger, more diverse, more execution-aware portfolio for human researchers under a fixed rubric? We introduce Budgeted Subset Refinement, a family of strategies that refine only a selected subset of candidates rather than refining all candidates uniformly. In a unified shared-candidate-pool evaluation across 10 random seeds and 10 research-ideation environments, raw generation and reranking alone produce no research-strong nonduplicate ideas under the benchmark rubric, while refinement is necessary for strong proxy-rated portfolios. Uniform refinement produces strong individual ideas but is not the best portfolio-level allocation of compute. Random-k refinement is a strong low-cost baseline, while diversity-aware MMR-k refinement gives the best overall proxy tradeoff: the highest research-strong nonduplicate yield, the lowest duplicate rate among successful methods, and the best cost per research-strong nonduplicate idea. A blinded external-judge robustness check on a balanced 72-item sample supports the broad refinement effect across independent model families, while showing that per-item rankings among refined strategies vary by judge. These results suggest that LLM research ideation systems should be evaluated not only as idea generators, but as budgeted support-allocation systems. The claims are scoped to proxy-rated portfolio quality and do not substitute for expert review or execution-grounded validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。