arXiv:2608.24001cs.AI2026-08

用推理轨迹筛选多样LLM,少模型也能更准预测未来

Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction

  • 基于推理过程差异聚类模型,选代表组成高效预测群体
  • 3个模型组成的聚类代表群胜过25个模型全体投票,效率提升80%
  • 适合追求高精度与低计算成本的未来预测任务

大语言模型在未来的预测中日益重要,促使人们采用多模型协同作为群体智慧机制。然而,单纯增加模型数量并不能保证多样性,因为不同模型可能表现出冗余行为。本文提出一种行为感知的框架,通过独立开发任务中的推理轨迹刻画模型,按行为相似性聚类,并选取代表性模型进行集体预测。我们使用7个开发基准评估25个LLM的行为多样性,再用两个未来预测基准评估多样化群体的表现。结果表明,群体构成比规模更重要:基于K-means++聚类选出的3个模型代表群,在两个预测基准上均优于对全部25个模型的常规投票,同时将模型调用减少88%,推理成本降低约80%。研究进一步表明,代表性行为多样性而非单纯最大化多样性,才是构建高效LLM群体的关键。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for future prediction, motivating the use of multiple models as a wisdom-of-the-crowd mechanism. However, simply increasing crowd size does not guarantee effective diversity, as different LLMs may exhibit redundant behaviors. We propose a behavior-aware framework for constructing diverse LLM crowds. The framework characterizes models using their reasoning traces on independent development tasks, clusters models by behavioral similarity, and selects representatives for collective prediction. We evaluate 25 LLMs using seven development benchmarks for behavioral diversity modeling and two future-prediction benchmarks for evaluating diverse crowds' performance. Our results show that crowd composition can matter more than crowd size: a three-model medoid crowd based on K-means++ behavioral clustering outperforms conventional voting over all 25 models on both prediction benchmarks, while reducing model calls by 88% and inference cost by approximately 80%. The results further suggest that representative behavioral diversity, rather than simply maximizing diversity, is important for constructing effective LLM crowds

未来预测群体智慧多样性推理轨迹

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。