arXiv:2603.19269cs.CL2026-03

拆解大模型六大核心机制,帮研究者判断如何用好LLM

From Tokens To Agents: A Researcher's Guide To Understanding Large Language Models

  • 从预训练数据到代理能力,系统解析大模型运作原理
  • 指出各模块的适用边界与潜在局限,避免误用
  • 适合科研人员评估LLM在自身课题中的实际价值

研究人员面临关键抉择:如何在工作中使用或不使用大语言模型。用好它们需要理解塑造其能力与限制的内在机制。本章以非技术门槛的方式解析大模型的六个核心组成:预训练数据、分词与嵌入、Transformer架构、概率生成、对齐机制以及代理能力。每个部分均从技术基础和研究意义双重视角分析,明确其具体优势与局限。文章不提供教条式建议,而是构建一个批判性推理框架,帮助研究者判断特定任务中是否适合使用大模型,并通过基于大模型代理模拟社交媒体动态的案例研究加以验证。

原文摘要 · Abstract (English)

Researchers face a critical choice: how to use -- or not use -- large language models in their work. Using them well requires understanding the mechanisms that shape what LLMs can and cannot do. This chapter makes LLMs comprehensible without requiring technical expertise, breaking down six essential components: pre-training data, tokenization and embeddings, transformer architecture, probabilistic generation, alignment, and agentic capabilities. Each component is analyzed through both technical foundations and research implications, identifying specific affordances and limitations. Rather than prescriptive guidance, the chapter develops a framework for reasoning critically about whether and how LLMs fit specific research needs, finally illustrated through an extended case study on simulating social media dynamics with LLM-based agents.

大模型原理研究工具代理系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。