arXiv:2511.01615cs.CLcs.AI2025-11

研究母语者西班牙语错误,提升AI对语言不完美性的理解能力。

Imperfect Language, Artificial Intelligence, and the Human Mind: An Interdisciplinary Approach to Linguistic Errors in Native Spanish Speakers

  • 构建超500条真实母语错误语料库,跨学科分析语言错误本质。
  • 测试GPT、Gemini等模型对错误的识别与纠正准确率,评估其泛化能力。
  • 为更贴近人类思维的语言模型研发提供认知科学依据,适合语言学与NLP研究者。

语言错误不仅是语法偏离,更是揭示语言认知结构的独特窗口,也暴露了当前人工智能系统在模仿人类语言时的局限性。本研究采用跨学科方法,分析母语西班牙语者产生的语言错误,旨在探讨大型语言模型(LLM)如何理解、复现或修正这些错误。研究融合理论语言学(分类错误类型)、神经语言学(结合实时脑语言处理)和自然语言处理(评估模型对错误的解释能力)。研究构建了一个包含超过500条真实母语错误的专用语料库,用于实证分析,并将其输入GPT、Gemini等模型,评估其解释准确性和对人类语言行为模式的泛化能力。该工作不仅深化对西班牙语母语认知的理解,也推动了更符合人类认知、能应对语言不完美性、可变性和模糊性的自然语言处理系统发展。

原文摘要 · Abstract (English)

Linguistic errors are not merely deviations from normative grammar; they offer a unique window into the cognitive architecture of language and expose the current limitations of artificial systems that seek to replicate them. This project proposes an interdisciplinary study of linguistic errors produced by native Spanish speakers, with the aim of analyzing how current large language models (LLM) interpret, reproduce, or correct them. The research integrates three core perspectives: theoretical linguistics, to classify and understand the nature of the errors; neurolinguistics, to contextualize them within real-time language processing in the brain; and natural language processing (NLP), to evaluate their interpretation against linguistic errors. A purpose-built corpus of authentic errors of native Spanish (+500) will serve as the foundation for empirical analysis. These errors will be tested against AI models such as GPT or Gemini to assess their interpretative accuracy and their ability to generalize patterns of human linguistic behavior. The project contributes not only to the understanding of Spanish as a native language but also to the development of NLP systems that are more cognitively informed and capable of engaging with the imperfect, variable, and often ambiguous nature of real human language.

语言错误大模型评估跨学科研究西班牙语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。