梳理意大利十年计算语言学研究脉络,揭示领域演进趋势。
Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
- 构建2014-2024年意大利计算语言学会议论文集,系统分析研究动态。
- 发现研究重心从词典资源转向大模型与多模态语言建模。
- 适合关注意大利NLP发展或语言技术演变的研究者参考。
过去十年,计算语言学(CL)与自然语言处理(NLP)迅猛发展,尤其受基于Transformer的大语言模型(LLMs)推动,研究目标与重点已从词汇语义资源转向语言建模与多模态。本研究通过分析意大利领先的计算语言学会议CLiC-it的前10届会议(2014–2024)的论文,构建了CLiC-it语料库。该语料库涵盖论文元数据(作者来源、性别、机构等)及内容,全面揭示研究主题演变。研究旨在为意大利及国际学术界提供领域发展趋势与关键进展的洞察,支持未来研究方向的决策。
原文摘要 · Abstract (English)
Over the past decade, Computational Linguistics (CL) and Natural Language Processing (NLP) have evolved rapidly, especially with the advent of Transformer-based Large Language Models (LLMs). This shift has transformed research goals and priorities, from Lexical and Semantic Resources to Language Modelling and Multimodality. In this study, we track the research trends of the Italian CL and NLP community through an analysis of the contributions to CLiC-it, arguably the leading Italian conference in the field. We compile the proceedings from the first 10 editions of the CLiC-it conference (from 2014 to 2024) into the CLiC-it Corpus, providing a comprehensive analysis of both its metadata, including author provenance, gender, affiliations, and more, as well as the content of the papers themselves, which address various topics. Our goal is to provide the Italian and international research communities with valuable insights into emerging trends and key developments over time, supporting informed decisions and future directions in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。