arXiv:2503.08192cs.CLcs.DL2025-03

用大模型自动识别古籍中的暴力描写并分类,效率远超人工

Automating Violence Detection and Categorization from Ancient Texts

  • 用大模型分析古籍暴力内容,支持多维度分类
  • 微调+数据增强后,暴力检测F1达0.93,细粒度分类达0.86
  • 适合历史、数字人文研究者快速提取暴力数据

文学中的暴力描述为人文研究提供了宝贵洞见。对历史学家而言,暴力描绘有助于分析重大战争及重要人物个人冲突的社会动态。手动收集暴力研究数据耗时费力。本研究首次评估大语言模型(LLMs)在识别古籍暴力内容并跨多维度分类方面的有效性。实验表明,大模型是规模化精准分析历史文本的有力工具,微调与数据增强显著提升性能,暴力检测最高达F1-score 0.93,细粒度分类达0.86。

原文摘要 · Abstract (English)

Violence descriptions in literature offer valuable insights for a wide range of research in the humanities. For historians, depictions of violence are of special interest for analyzing the societal dynamics surrounding large wars and individual conflicts of influential people. Harvesting data for violence research manually is laborious and time-consuming. This study is the first one to evaluate the effectiveness of large language models (LLMs) in identifying violence in ancient texts and categorizing it across multiple dimensions. Our experiments identify LLMs as a valuable tool to scale up the accurate analysis of historical texts and show the effect of fine-tuning and data augmentation, yielding an F1-score of up to 0.93 for violence detection and 0.86 for fine-grained violence categorization.

古籍分析暴力检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。