用低成本数据预测家庭用水大肠杆菌污染,助力资源匮乏地区水质监测。
Peoples Water Data: Enabling Reliable Field Data Generation and Microbial Contamination Screening in Household Drinking Water
- 基于物理化学与环境指标构建两阶段机器学习模型
- 在2207个样本上实现高可靠污染风险预测
- 结合学生实践与实时质控,提升野外数据可信度
安全饮用水仍是全球性的重大公共卫生问题,尤其在低资源地区,常规微生物监测难以普及。尽管大肠杆菌是国际公认的粪便污染指示物,但实验室检测常无法大规模开展。本研究针对印度金奈地区分散式家庭用水点,开发并评估了一种两阶段机器学习框架,利用低成本的理化参数和上下文信息预测大肠杆菌存在。数据集来自“人民水数据”计划,共采集3023份样本,经标准化、技术清洗和异常值筛选后保留2207个有效样本。该框架可作为资源受限环境中优先开展微生物检测的可扩展决策支持工具,填补了点对点污染风险评估的重要空白。此外,研究还嵌入了由AI支持的现场实施框架,结合面向学生的指导手册与实时质量控制,显著提升了协议依从性、可追溯性与数据可靠性。
原文摘要 · Abstract (English)
Unsafe drinking water remains a major public health concern globally, particularly in low-resource regions where routine microbiological surveillance is limited. Although Escherichia coli is the internationally recognized indicator of fecal contamination, laboratory-based testing is often inaccessible at scale. In this study, we developed and evaluated a two-stage machine-learning framework for predicting E. coli presence in decentralized household point-of-use drinking water in Chennai, India using low-cost physicochemical and contextual indicators. The dataset comprised 3,023 samples collected under the Peoples Water Data initiative; after harmonization, technical cleaning, and outlier screening, 2,207 valid samples were retained. This framework provides a scalable decision-support tool for prioritizing microbiological testing in resource-constrained environments and addresses an important gap in point-of-use contamination risk assessment. Beyond predictive modeling, the present study was conducted within an AI-supported field implementation framework that combined student-facing guidance and real-time QC to improve protocol adherence, traceability, and data reliability in decentralized household water monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。