arXiv:2512.16694cs.AI2025-12被引 1

用关联规则挖掘自动分组圣训主题,助力数字伊斯兰研究

Unsupervised Thematic Clustering Of hadith Texts Using The Apriori Algorithm

  • 基于Apriori算法挖掘未标注圣训文本中的语义关联模式
  • 发现礼拜次数与启示经文、圣训故事间的显著关联关系
  • 适合数字人文与宗教文本分析领域的研究者参考

为应对伊斯兰文本数字化带来的挑战,本研究旨在自动化圣训主题聚类。基于文献综述,无监督学习中的Apriori算法在识别未标注文本中的关联模式和语义关系方面表现良好。研究采用布哈里圣训的印尼译本作为数据集,经过大小写统一、标点清理、分词、停用词去除和词干化等预处理步骤后,使用Apriori算法进行关联规则挖掘,设置支持度、置信度和提升度参数。结果揭示出礼拜次数-祈祷、经文启示-圣训故事之间的有意义关联模式,分别对应崇拜、启示和圣训叙述主题。研究表明,Apriori算法能有效自动发现隐含语义关系,推动数字伊斯兰研究发展,并为基于技术的学习系统提供支持。

原文摘要 · Abstract (English)

This research stems from the urgency to automate the thematic grouping of hadith in line with the growing digitalization of Islamic texts. Based on a literature review, the unsupervised learning approach with the Apriori algorithm has proven effective in identifying association patterns and semantic relations in unlabeled text data. The dataset used is the Indonesian Translation of the hadith of Bukhari, which first goes through preprocessing stages including case folding, punctuation cleaning, tokenization, stopword removal, and stemming. Next, an association rule mining analysis was conducted using the Apriori algorithm with support, confidence, and lift parameters. The results show the existence of meaningful association patterns such as the relationship between rakaat-prayer, verse-revelation, and hadith-story, which describe the themes of worship, revelation, and hadith narration. These findings demonstrate that the Apriori algorithm has the ability to automatically uncover latent semantic relationships, while contributing to the development of digital Islamic studies and technology-based learning systems.

文本挖掘圣训分析关联规则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。