arXiv:2510.00966cs.IRcs.AI2025-10被引 1

用深度学习提升阿拉伯语搜索结果聚类效果

Deep Learning-Based Approach for Improving Relational Aggregated Search

  • 结合AraBERT与堆叠自编码器提取文本特征
  • 聚类准确率显著提升,结果更相关
  • 适合做多语言信息检索优化的开发者

互联网信息爆炸背景下,亟需高效的内容聚合搜索系统。为提升聚合搜索环境中阿拉伯语文本数据的聚类效果,本研究探索了堆叠自编码器与AraBERT嵌入等先进自然语言处理技术的应用。相较于传统搜索引擎在精准度、上下文相关性和个性化方面的不足,本方法能生成更丰富、上下文感知的搜索结果表征。通过K-means聚类算法挖掘结果中的独特特征与关系,并在多个阿拉伯语查询上验证有效性。实验表明,堆叠自编码器在表示学习中适用于聚类任务,显著提升聚类质量与搜索结果的准确性和相关性。

原文摘要 · Abstract (English)

Due to an information explosion on the internet, there is a need for the development of aggregated search systems that can boost the retrieval and management of content in various formats. To further improve the clustering of Arabic text data in aggregated search environments, this research investigates the application of advanced natural language processing techniques, namely stacked autoencoders and AraBERT embeddings. By transcending the limitations of traditional search engines, which are imprecise, not contextually relevant, and not personalized, we offer more enriched, context-aware characterizations of search results, so we used a K-means clustering algorithm to discover distinctive features and relationships in these results, we then used our approach on different Arabic queries to evaluate its effectiveness. Our model illustrates that using stacked autoencoders in representation learning suits clustering tasks and can significantly improve clustering search results. It also demonstrates improved accuracy and relevance of search results.

搜索优化聚类阿拉伯语深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。