首次系统评估大模型类比生成多样性,发现其输出存在领域单一问题。
On the Diversity of Analogy Making in Large Language Models

- 对比10个主流大模型的类比生成能力,评估输出多样性。
- 发现模型普遍局限于少数领域,跨域连接能力弱。
- 揭示多样性与质量间的权衡机制,适合研究认知与创造力的学者。
大语言模型(LLMs)在类比生成方面展现出巨大潜力,这是推动创新与创造力的核心认知能力。尽管已有研究广泛探讨了基于LLM的类比生成应用及其内在机制,但其输出多样性仍缺乏系统探索,而多样性对拓展跨领域联系、促进科学创新至关重要。本文对十种先进的开源与闭源大模型进行了类比生成多样性的全面评估。结果表明,存在严重的领域同质性问题:模型倾向于从有限的目标领域生成类比,限制了跨查询和模型内部的多样性。此外,分析显示现有提升多样性的方法存在根本性权衡——增加多样性往往以牺牲输出质量为代价。最后,通过因果分析揭示不同模型中决定类比多样性的敏感区域存在显著差异,暗示了多样性-质量权衡的潜在机制。据我们所知,这是首个系统研究基于大模型类比生成输出多样性的研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable potential for analogy making, a core cognitive capability that drives novelty and creativity. While prior research has extensively investigated the applications and underlying mechanisms of LLM-based analogy making, its output diversity remains largely unexplored, despite being essential for broadening cross-domain connections and fostering scientific innovation. In this work, we present a comprehensive evaluation of analogy diversity across ten state-of-the-art open- and closed-source LLMs. Our findings highlight a concerning issue of domain homogeneity, a prevalent tendency for LLMs to generate analogies from a narrow set of target domains, limiting both inter-query and intra-model diversity. Furthermore, our analysis reveals a fundamental trade-off in existing LLM diversity-enhancement methods: increasing output diversity often comes at the expense of output quality. Finally, our causal analysis of LLM information flow reveals substantial differences in the model-sensitive regions governing analogy diversity across LLMs, suggesting a potential mechanism for the observed diversity-quality trade-off. To our knowledge, this is among the first studies to systematically investigate output diversity in LLM-based analogy making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。