arXiv:2508.06435cs.CLcs.AI2025-08被引 2

用小模型跨语言分析移民话题,速度快成本低。

Learning the Topic, Not the Language: How LLMs Classify Online Immigration Discourse Across Languages

  • 用微调的LLaMA 3.2-3B模型做多语言话题分类,无需翻译。
  • 在13种语言上实现跨语言主题识别,仅需1-2种语言微调数据。
  • 推理速度比商用模型快26-168倍,成本降低超1000倍。

大型语言模型(LLMs)为在线话语的规模化分析提供了新机遇,但其在多语言社会科学研究中的应用受限于模型规模、成本和语言偏见。本文开发了一种轻量级、开源的LLM框架,基于微调的LLaMA 3.2-3B模型,对13种语言的移民相关推文进行分类。与依赖BERT类模型或翻译管道的先前工作不同,本方法结合主题分类与立场检测,证明在仅1-2种语言上微调的LLM可泛化至未见过的语言。捕捉意识形态细微差别则需多语言微调。该方法通过少量来自代表性不足语言的数据纠正预训练偏见,避免依赖专有系统。相比商用LLM,推理速度提升26-168倍,成本降低超过1000倍,支持对数十亿条推文的实时分析。这一以规模优先的框架实现了跨语言与文化语境下公众态度研究的包容性与可复现性。

原文摘要 · Abstract (English)

Large language models (LLMs) offer new opportunities for scalable analysis of online discourse. Yet their use in multilingual social science research remains constrained by model size, cost and linguistic bias. We develop a lightweight, open-source LLM framework using fine-tuned LLaMA 3.2-3B models to classify immigration-related tweets across 13 languages. Unlike prior work relying on BERT style models or translation pipelines, we combine topic classification with stance detection and demonstrate that LLMs fine-tuned in just one or two languages can generalize topic understanding to unseen languages. Capturing ideological nuance, however, benefits from multilingual fine-tuning. Our approach corrects pretraining biases with minimal data from under-represented languages and avoids reliance on proprietary systems. With 26-168x faster inference and over 1000x cost savings compared to commercial LLMs, our method supports real-time analysis of billions of tweets. This scale-first framework enables inclusive, reproducible research on public attitudes across linguistic and cultural contexts.

多语言主题分类轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。