arXiv:2410.03749cs.CLcs.LG2024-10被引 3

用媒体文本分析国家和平程度,模型准确率随数据量变化

Machine Learning Classification of Peaceful Countries: A Comparative Analysis and Dataset Optimization

  • 从全球媒体文章提取语言模式,用向量相似度建模分类
  • 数据量减小导致准确率下降,揭示规模对模型影响
  • 适合关注国际关系与文本挖掘的研究者参考

本文提出一种机器学习方法,通过分析全球媒体文章中的语言模式,对国家是否和平进行分类。采用向量嵌入与余弦相似度构建监督分类模型,有效识别和平国家。同时探讨数据集规模对模型性能的影响,研究数据缩减对分类准确率的作用。结果揭示了大规模文本数据在和平研究中的挑战与机遇。

原文摘要 · Abstract (English)

This paper presents a machine learning approach to classify countries as peaceful or non-peaceful using linguistic patterns extracted from global media articles. We employ vector embeddings and cosine similarity to develop a supervised classification model that effectively identifies peaceful countries. Additionally, we explore the impact of dataset size on model performance, investigating how shrinking the dataset influences classification accuracy. Our results highlight the challenges and opportunities associated with using large-scale text data for peace studies.

文本分类和平研究机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。