用大模型自动分析日裔美国人拘禁口述史,提升档案可读性。
Large Language Models for Oral History Understanding with Text Classification and Sentiment Analysis
- 设计多阶段流程,结合专家标注与提示工程,提升标注质量。
- 在558句样本上,ChatGPT语义分类F1达88.71%,模型表现稳定。
- 适用于历史敏感档案的自动化分析,为数字人文提供伦理范式。
口述史是记录个体经历的重要资料,尤其对遭受系统性不公与历史湮没的群体至关重要。然而,由于格式非结构化、情感复杂且标注成本高,大规模分析仍受限。本文提出一种可扩展框架,用于日本裔美国人拘禁口述史(JAIOH)的语义与情感标注。基于大语言模型(LLM),构建高质量数据集,评估ChatGPT、Llama与Qwen在零样本、少样本及RAG策略下的表现。在15位讲述者共558句文本上,语义分类中ChatGPT F1达88.71%,优于Llama(84.99%)和Qwen(83.72%);情感分析中Llama略胜(82.66%)于Qwen(82.66%)与ChatGPT(82.29%)。最优提示配置用于标注JAIOH中1,002份访谈的92,191句文本。结果表明,在良好提示引导下,LLM可高效完成大规模口述史的语义与情感标注。研究提供可复用的标注流程与实践指南,推动人工智能在数字人文与集体记忆保存中的负责任应用。
原文摘要 · Abstract (English)
Oral histories are vital records of lived experience, particularly within communities affected by systemic injustice and historical erasure. Effective and efficient analysis of their oral history archives can promote access and understanding of the oral histories. However, Large-scale analysis of these archives remains limited due to their unstructured format, emotional complexity, and high annotation costs. This paper presents a scalable framework to automate semantic and sentiment annotation for Japanese American Incarceration Oral History. Using LLMs, we construct a high-quality dataset, evaluate multiple models, and test prompt engineering strategies in historically sensitive contexts. Our multiphase approach combines expert annotation, prompt design, and LLM evaluation with ChatGPT, Llama, and Qwen. We labeled 558 sentences from 15 narrators for sentiment and semantic classification, then evaluated zero-shot, few-shot, and RAG strategies. For semantic classification, ChatGPT achieved the highest F1 score (88.71%), followed by Llama (84.99%) and Qwen (83.72%). For sentiment analysis, Llama slightly outperformed Qwen (82.66%) and ChatGPT (82.29%), with all models showing comparable results. The best prompt configurations were used to annotate 92,191 sentences from 1,002 interviews in the JAIOH collection. Our findings show that LLMs can effectively perform semantic and sentiment annotation across large oral history collections when guided by well-designed prompts. This study provides a reusable annotation pipeline and practical guidance for applying LLMs in culturally sensitive archival analysis. By bridging archival ethics with scalable NLP techniques, this work lays the groundwork for responsible use of artificial intelligence in digital humanities and preservation of collective memory. GitHub: https://github.com/kc6699c/LLM4OralHistoryAnalysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。