arXiv:2602.15436cs.CL2026-02

用大模型对战后难民访谈中的参与行为分类,量化社会融合程度。

Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs

  • 构建四维参与分类框架:活动类型、社交性、频率、体力要求
  • 大模型在35万条记录上分类准确率接近专家水平
  • 为研究难民社会融入提供可量化的结构化数据资源

数字化历史档案使大规模研究日常社会生活成为可能,但文本提取的信息常难以直接回答历史学家或社会学家的定量问题。我们针对芬兰二战时期卡累利阿难民家庭访谈的大规模语料库开展研究。前期工作已从访谈中提取超过35万条休闲活动与组织参与记录,得到71,000个唯一名称——数量过大,无法直接分析。为此,我们开发了一个涵盖参与关键维度的分类框架:活动/组织类型、典型社交性、发生频率、体力要求。通过人工标注建立金标准数据集以可靠评估,测试大语言模型能否规模化应用该分类体系。采用多轮模型运行的投票机制,发现开源大模型可接近专家判断效果。最终将方法应用于全部35万条实体,生成可用于后续社会融合及相关研究的结构化资源。

原文摘要 · Abstract (English)

Digitized historical archives make it possible to study everyday social life on a large scale, but the information extracted directly from text often does not directly allow one to answer the research questions posed by historians or sociologists in a quantitative manner. We address this problem in a large collection of Finnish World War II Karelian evacuee family interviews. Prior work extracted more than 350K mentions of leisure time activities and organizational memberships from these interviews, yielding 71K unique activity and organization names -- far too many to analyze directly. We develop a categorization framework that captures key aspects of participation (the kind of activity/organization, how social it typically is, how regularly it happens, and how physically demanding it is). We annotate a gold-standard set to allow for a reliable evaluation, and then test whether large language models can apply the same schema at scale. Using a simple voting approach across multiple model runs, we find that an open-weight LLM can closely match expert judgments. Finally, we apply the method to label the 350K entities, producing a structured resource for downstream studies of social integration and related outcomes.

社会融合大模型历史文本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。