用机器学习预测质性研究数据饱和点,让样本量选择更客观可靠。
Q-Sat AI: Machine Learning-Based Decision Support for Data Saturation in Qualitative Studies
- 基于10个关键参数的集成学习模型,量化判断数据饱和
- 最高解释力达测试R²约0.85,有效捕捉非线性关系
- 适合研究者、审稿人和导师用于科学论证样本量
质性研究中样本量确定长期依赖主观的数据饱和原则,易导致不一致并损害方法严谨性。本研究提出一种基于机器学习(ML)的系统化模型,利用来自五种基础质性研究方法(案例研究、扎根理论、现象学、叙事研究、民族志)的数据集,构建集成学习模型。评估了研究范围、信息力、研究者能力等10项关键参数,采用序数尺度作为输入特征。经预处理与异常值剔除后,对比训练多种机器学习算法。KNN、梯度提升(GB)、随机森林(RF)、XGBoost及决策树(DT)表现最佳,测试R²约0.85,能有效建模质性抽样决策中的复杂非线性关系。特征重要性分析证实研究设计类型与信息力的关键作用,为质性方法论核心假设提供量化验证。研究最后提出一个面向网络应用的概念框架,旨在为质性研究者、期刊审稿人及论文导师提供决策支持系统。该模型标志着在标准化样本量论证、提升透明度及强化质性探究认识论基础方面迈出关键一步。
原文摘要 · Abstract (English)
The determination of sample size in qualitative research has traditionally relied on the subjective and often ambiguous principle of data saturation, which can lead to inconsistencies and threaten methodological rigor. This study introduces a new, systematic model based on machine learning (ML) to make this process more objective. Utilizing a dataset derived from five fundamental qualitative research approaches - namely, Case Study, Grounded Theory, Phenomenology, Narrative Research, and Ethnographic Research - we developed an ensemble learning model. Ten critical parameters, including research scope, information power, and researcher competence, were evaluated using an ordinal scale and used as input features. After thorough preprocessing and outlier removal, multiple ML algorithms were trained and compared. The K-Nearest Neighbors (KNN), Gradient Boosting (GB), Random Forest (RF), XGBoost, and Decision Tree (DT) algorithms showed the highest explanatory power (Test R2 ~ 0.85), effectively modeling the complex, non-linear relationships involved in qualitative sampling decisions. Feature importance analysis confirmed the vital roles of research design type and information power, providing quantitative validation of key theoretical assumptions in qualitative methodology. The study concludes by proposing a conceptual framework for a web-based computational application designed to serve as a decision support system for qualitative researchers, journal reviewers, and thesis advisors. This model represents a significant step toward standardizing sample size justification, enhancing transparency, and strengthening the epistemological foundation of qualitative inquiry through evidence-based, systematic decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。