用文本嵌入分析地质矿产分布,找矿选址新方法
Geological Inference from Textual Data using Word Embeddings
- 用词嵌入提取关键词与地质文本的语义关联
- 结合降维技术提升空间分布预测准确率
- 适合地质勘探与自然语言处理交叉研究者
本研究探索利用自然语言处理技术定位地质资源,重点关注工业矿物。通过训练GloVe词嵌入模型,提取目标关键词与地质文本间的语义关系。文本经筛选仅保留具有地理意义的词汇(如城市名),并按与目标关键词的余弦相似度排序。应用主成分分析(PCA)、自编码器、变分自编码器(VAE)及带长短期记忆的变分自编码器(VAE-LSTM)进行特征提取与降维,以增强语义关系识别。基准测试中,计算前十名语义相关城市与已知矿点之间的大圆距离(haversine距离)。结果表明,结合NLP与降维技术可提供对自然资源空间分布的有意义洞察,虽预测位置与实际区域一致,但精度仍有提升空间。
原文摘要 · Abstract (English)
This research explores the use of Natural Language Processing (NLP) techniques to locate geological resources, with a specific focus on industrial minerals. By using word embeddings trained with the GloVe model, we extract semantic relationships between target keywords and a corpus of geological texts. The text is filtered to retain only words with geographical significance, such as city names, which are then ranked by their cosine similarity to the target keyword. Dimensional reduction techniques, including Principal Component Analysis (PCA), Autoencoder, Variational Autoencoder (VAE), and VAE with Long Short-Term Memory (VAE-LSTM), are applied to enhance feature extraction and improve the accuracy of semantic relations. For benchmarking, we calculate the proximity between the ten cities most semantically related to the target keyword and identified mine locations using the haversine equation. The results demonstrate that combining NLP with dimensional reduction techniques provides meaningful insights into the spatial distribution of natural resources. Although the result shows to be in the same region as the supposed location, the accuracy has room for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。