arXiv:2606.24331cs.CLcs.ET2026-06

系统梳理Transformer模型架构与应用,揭示其在多领域的适配逻辑与关键权衡。

Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment

  • 构建主流Transformer变体分类体系,涵盖编码器、解码器等六类结构
  • 量化参数量与能耗关系,揭示模型规模扩张的现实成本
  • 剖析对齐技术与数据溯源如何重塑'顶尖模型'的定义,适合决策者参考

基于Transformer的语言模型已成为自然语言处理的默认基础,但新模型发布速度过快,使从业者难以区分持久性思想与增量宣传。本文从机制与应用双层面展开综述:机制层面,将主流Transformer家族归入工作型分类体系,涵盖仅编码器、仅解码器、编码器-解码器、长上下文、置换式及生成-判别器等变体;并延伸讨论2023年后改变实践的关键进展,包括指令微调、人类反馈强化学习、直接偏好优化、专家混合扩展、检索增强,以及OpenAI、Anthropic、Google、Meta、Mistral和DeepSeek等机构的当前旗舰模型。应用层面,调研模型在医疗、金融、法律、教育、客户服务、创意写作及科研中的部署情况,关联具体能力与适用场景。论文贡献在于基于调研的批判性评估:从部署决策相关四维度比较架构差异,量化参数量与能耗间的权衡关系;同时探讨对齐方法、数据来源与基准饱和如何重构‘最先进’的内涵。最后列出值得深入研究的若干科学问题。

原文摘要 · Abstract (English)

Transformer-based language models have become the default substrate for natural language processing and the pace of new releases has made it hard for practitioners to separate durable ideas from the noise of incremental announcements. This review works at two levels. At the level of mechanism, we organise the main transformer families into a working taxonomy, covering encoder-only, decoder-only, encoder-decoder, long-context, permutation-based, and generator-discriminator variants. We then extend the discussion to post-2023 developments that changed the picture in practice: instruction tuning, reinforcement learning from human feedback, direct preference optimisation, mixture-of-experts scaling, retrieval augmentation and the current flagship model families from OpenAI, Anthropic, Google, Meta, Mistral and DeepSeek. At the level of use, we survey deployments across healthcare, finance, legal, education, customer service, creative writing and scientific work. Based on this we link each to the specific capabilities that make a transformer the appropriate tool. The contribution of this paper is a critical assessment that is based on the survey. We compare architectures on four axes that matter to deployment decisions, we quantify the trade-off between parameter count and energy cost. We also discuss how alignment methods, data provenance and benchmark saturation change what it means to call a model "state of the art". The final section lists the research questions that we think deserve more attention.

Transformer模型评估多领域应用架构对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。