arXiv:2604.07956cs.AI2026-04ACL

用多模态数据自动分类企业行业,省去人工标注成本。

MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems

  • 结合文本与地理信息,构建首个多模态企业分类基准
  • 无需训练,基线模型准确率达62.1%至74.1%
  • 多轮对话+上下文增强提升22.8%准确率,适合政策研究者

行业分类体系是公共和企业数据库的核心,用于根据经济活动对机构进行归类。由于企业名录规模庞大,人工标注成本高昂,且每次行业分类更新都需大量数据重新微调模型。本文通过利用现有或易获取的多模态资源,模拟人工专家验证流程实现行业分类。提出MONETA,首个融合文本(网站、维基百科、维基数据)与地理空间信息(OpenStreetMap、卫星图像)的多模态行业分类基准。数据集包含欧洲1000家企业,依据欧盟NACE标准划分20个经济活动类别。无需训练的基线模型在开源与闭源多模态大模型上分别达到62.10%和74.10%准确率。通过多轮设计、上下文增强与分类解释,准确率最高提升22.80%。研究将发布数据集与优化指南。

原文摘要 · Abstract (English)

Industry classification schemes are integral parts of public and corporate databases as they classify businesses based on economic activity. Due to the size of the company registers, manual annotation is costly, and fine-tuning models with every update in industry classification schemes requires significant data collection. We replicate the manual expert verification by using existing or easily retrievable multimodal resources for industry classification. We present MONETA, the first multimodal industry classification benchmark with text (Website, Wikipedia, Wikidata) and geospatial sources (OpenStreetMap and satellite imagery). Our dataset enlists 1,000 businesses in Europe with 20 economic activity labels according to EU guidelines (NACE). Our training-free baseline reaches 62.10% and 74.10% with open and closed-source Multimodal Large Language Models (MLLM). We observe an increase of up to 22.80% with the combination of multi-turn design, context enrichment, and classification explanations. We will release our dataset and the enhanced guidelines.

行业分类多模态地理信息LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。