arXiv:2505.19240cs.CLcs.AI2025-05综述被引 20

系统梳理大模型局限性研究趋势,揭示热点演变与数据分布规律。

LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models

  • 基于25万篇论文的关键词与LLM分类,构建半自动化调研框架
  • 2022至2025年,大模型局限研究增长超五倍,2025年占相关论文30%以上
  • 聚焦推理、幻觉、偏见等核心问题,开源标注数据集与方法论

大语言模型(LLM)研究快速发展,对其局限性的关注也日益增加。本文采用自下而上的数据驱动方法,对2022年至2025年初的14,648篇相关论文进行半自动化综述,涵盖25万篇ACL与arXiv论文。通过关键词筛选、基于LLM的分类及专家验证,结合两种主题聚类方法(HDBSCAN+BERTopic与LlooM),发现ACL中大模型相关论文占比增长逾五倍,arXiv中增长近八倍。自2022年起,大模型局限性研究增速更快,至2025年已占所有大模型论文的30%以上。推理能力仍是首要研究方向,其次为泛化、幻觉、偏见与安全问题。ACL主题分布相对稳定,arXiv则转向安全风险、对齐、幻觉、知识编辑与多模态等方向。本文提供量化趋势分析,并公开标注摘要数据集与验证方法,地址:https://github.com/a-kostikova/LLLMs-Survey。

原文摘要 · Abstract (English)

Large language model (LLM) research has grown rapidly, along with increasing concern about their limitations. In this survey, we conduct a data-driven, semi-automated review of research on limitations of LLMs (LLLMs) from 2022 to early 2025 using a bottom-up approach. From a corpus of 250,000 ACL and arXiv papers, we identify 14,648 relevant papers using keyword filtering, LLM-based classification, validated against expert labels, and topic clustering (via two approaches, HDBSCAN+BERTopic and LlooM). We find that the share of LLM-related papers increases over fivefold in ACL and nearly eightfold in arXiv between 2022 and 2025. Since 2022, LLLMs research grows even faster, reaching over 30% of LLM papers by 2025. Reasoning remains the most studied limitation, followed by generalization, hallucination, bias, and security. The distribution of topics in the ACL dataset stays relatively stable over time, while arXiv shifts toward security risks, alignment, hallucinations, knowledge editing, and multimodality. We offer a quantitative view of trends in LLLMs research and release a dataset of annotated abstracts and a validated methodology, available at: https://github.com/a-kostikova/LLLMs-Survey.

大模型局限研究趋势数据驱动综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。