arXiv:2512.09552cs.CL2025-12被引 2

为语言科学中的大模型应用提供系统方法框架,提升研究可复现性。

Systematic Framework of Application Methods for Large Language Models in Language Sciences

  • 分三类方法:提示交互、微调模型、提取上下文嵌入,对应不同研究目标。
  • 通过案例实证验证框架有效性,支持从探索到验证的多阶段研究流程。
  • 适合希望规范使用大模型的语言学者,推动语言学向可验证科学转型。

大语言模型正重塑语言科学,但其广泛应用面临方法碎片化与系统性不足的问题。本文提出两个综合性方法框架,指导大模型在语言科学中的战略化、负责任应用。第一个方法选择框架系统化了三种互补路径:(1) 使用通用模型进行提示交互,用于探索性分析与假设生成;(2) 微调开源模型,支持理论驱动的确认性研究与高质量数据生成;(3) 提取上下文嵌入,用于定量分析与模型内部机制探查。每种方法均附技术实现与权衡分析,并通过实证案例验证。基于此,第二个系统框架提供多阶段研究流程的具体配置方案。通过回顾性分析、前瞻性应用及专家评估调查,验证了框架的有效性。通过将研究问题与合适的大模型方法精准匹配,该体系推动语言科学研究从零散应用迈向可复现、可批判的严谨科学范式。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are transforming language sciences. However, their widespread deployment currently suffers from methodological fragmentation and a lack of systematic soundness. This study proposes two comprehensive methodological frameworks designed to guide the strategic and responsible application of LLMs in language sciences. The first method-selection framework defines and systematizes three distinct, complementary approaches, each linked to a specific research goal: (1) prompt-based interaction with general-use models for exploratory analysis and hypothesis generation; (2) fine-tuning of open-source models for confirmatory, theory-driven investigation and high-quality data generation; and (3) extraction of contextualized embeddings for further quantitative analysis and probing of model internal mechanisms. We detail the technical implementation and inherent trade-offs of each method, supported by empirical case studies. Based on the method-selection framework, the second systematic framework proposed provides constructed configurations that guide the practical implementation of multi-stage research pipelines based on these approaches. We then conducted a series of empirical experiments to validate our proposed framework, employing retrospective analysis, prospective application, and an expert evaluation survey. By enforcing the strategic alignment of research questions with the appropriate LLM methodology, the frameworks enable a critical paradigm shift in language science research. We believe that this system is fundamental for ensuring reproducibility, facilitating the critical evaluation of LLM mechanisms, and providing the structure necessary to move traditional linguistics from ad-hoc utility to verifiable, robust science.

大模型语言科学方法框架可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。