arXiv:2507.07217cs.AIcs.LG2025-07

用神经符号方法从新闻中识别供应链强迫劳动,无需大量训练数据。

Neurosymbolic Feature Extraction for Identifying Forced Labor in Supply Chains

  • 通过提问树引导大模型自动提取新闻相关特征
  • 人工与机器分类结果对比验证了方法有效性
  • 适合关注供应链伦理与合规的从业者

供应链网络复杂难测,尤其当涉及伪造零件、强迫劳动或人口贩卖等非法活动时。传统机器学习依赖大量标注数据,但非法供应链数据稀疏且常被故意篡改以掩盖真实情况。本文探索神经符号方法,在不依赖大规模训练数据的前提下,自动识别与非法活动相关的新型模式,尤其针对具有时间维度的复杂数据。研究采用问题树框架调用大语言模型,从新闻报道中提取并量化信息相关性,系统评估人类与机器对供应链强迫劳动新闻分类的一致性。结果表明该方法能有效发现隐蔽关联,为供应链透明化提供新路径。

原文摘要 · Abstract (English)

Supply chain networks are complex systems that are challenging to analyze; this problem is exacerbated when there are illicit activities involved in the supply chain, such as counterfeit parts, forced labor, or human trafficking. While machine learning (ML) can find patterns in complex systems like supply chains, traditional ML techniques require large training data sets. However, illicit supply chains are characterized by very sparse data, and the data that is available is often (purposely) corrupted or unreliable in order to hide the nature of the activities. We need to be able to automatically detect new patterns that correlate with such illegal activity over complex, even temporal data, without requiring large training data sets. We explore neurosymbolic methods for identifying instances of illicit activity in supply chains and compare the effectiveness of manual and automated feature extraction from news articles accurately describing illicit activities uncovered by authorities. We propose a question tree approach for querying a large language model (LLM) to identify and quantify the relevance of articles. This enables a systematic evaluation of the differences between human and machine classification of news articles related to forced labor in supply chains.

供应链强迫劳动神经符号大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。