arXiv:2503.22856cs.CL2025-03被引 3

用大模型生成纯净推文数据集,揭示噪声对建筑功能分类的致命影响

Generating Synthetic Oracle Datasets to Analyze Noise Impact: A Study on Building Function Classification Using Tweets

  • 用大模型合成仅含正确标签和相关推文的纯净数据集
  • 真实推文噪声使mBERT性能退化至关键词模型水平
  • 适合研究文本噪声、跨域泛化或弱监督学习的学者

推文为地球观测任务提供宝贵的语义信息,是遥感影像的补充模态。在建筑功能分类(BFC)中,推文常通过地理启发式方法获取,并借助外部数据库标注,这一过程本质上是弱监督的,引入了标签噪声和句子级特征噪声(如无关或无信息的推文)。尽管标签噪声被广泛研究,但句子级特征噪声的影响仍缺乏深入探讨,主要受限于缺乏可控制的清洁基准数据集。本文提出一种基于大语言模型(LLM)生成合成基准数据集的方法,该数据集仅包含正确标注且语义上与建筑相关的推文。该基准数据集支持系统性分析噪声影响,而这些在真实数据中难以分离。我们通过Naive Bayes和mBERT分类器,在三种配置下评估其效用:真实数据与合成数据训练,以及跨域泛化。结果表明,真实推文中的噪声显著削弱了mBERT的上下文学习能力,使其性能降至关键词模型水平;而使用干净合成数据时,mBERT能有效学习,远超Naive Bayes。这表明在此任务中,解决特征噪声比提升模型复杂度更为关键。我们的合成数据集为未来噪声注入研究提供了新实验环境,已公开于GitHub。

原文摘要 · Abstract (English)

Tweets provides valuable semantic context for earth observation tasks and serves as a complementary modality to remote sensing imagery. In building function classification (BFC), tweets are often collected using geographic heuristics and labeled via external databases, an inherently weakly supervised process that introduces both label noise and sentence level feature noise (e.g., irrelevant or uninformative tweets). While label noise has been widely studied, the impact of sentence level feature noise remains underexplored, largely due to the lack of clean benchmark datasets for controlled analysis. In this work, we propose a method for generating a synthetic oracle dataset using LLM, designed to contain only tweets that are both correctly labeled and semantically relevant to their associated buildings. This oracle dataset enables systematic investigation of noise impacts that are otherwise difficult to isolate in real-world data. To assess its utility, we compare model performance using Naive Bayes and mBERT classifiers under three configurations: real vs. synthetic training data, and cross-domain generalization. Results show that noise in real tweets significantly degrades the contextual learning capacity of mBERT, reducing its performance to that of a simple keyword-based model. In contrast, the clean synthetic dataset allows mBERT to learn effectively, outperforming Naive Bayes Bayes by a large margin. These findings highlight that addressing feature noise is more critical than model complexity in this task. Our synthetic dataset offers a novel experimental environment for future noise injection studies and is publicly available on GitHub.

建筑分类文本噪声合成数据弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。