Machine Learning-Based Groundwater Quality Assessment and Drinkability Prediction

Main Article Content

Onkar Nath Thakur, Santosh Kumar, Nitesh Singh Bhati

Abstract

Groundwater is a major source of drinking water for a large proportion of the global population; however, rapid urbanization, industrialization, and agricultural activities have led to a marked deterioration of groundwater quality. Thus, there is a need for efficient, reliable and rapid groundwater quality assessment techniques for identifying contamination hotspots. The typical measurement of groundwater quality involves physicochemical analysis and WQI calculation, which is a time-taking and labor-intensive method and not suitable for large-scale monitoring. This study applies the supervised machine learning method for the automated assessment of groundwater quality and prediction of its drinkability quality in terms of hydrochemical characteristics. The Central Ground Water Board (CGWB), India, contains a dataset of 1345 groundwater samples containing 13 hydrochemical parameters (pH; Electrical Conductivity (EC); Total Dissolved Solids (TDS); Total Hardness (TH); Alkalinity; Calcium; Magnesium; Sodium; Potassium; Chloride; Sulphate; Fluoride; Bicarbonate). The Water Quality Index (WQI) was computed based on standard guidelines for drinking water quality and groundwater samples were categorized into four drinkability classes viz. Excellent, Good, Poor and Bad thereby creating a multi-class classification problem. Before the development of the applied model, exploratory data analysis, correlation analysis, outlier detection, data cleansing, and feature normalization were done to improve data quality. Using accuracy, precision, recall, and F1-score six supervised machine learning algorithms such as Logistic Regression, K-Nearest Neighbors (KNN), Decision Tree, Support Vector Machine (SVM), AdaBoost and Extreme Gradient Boosting(XGBoost) were comparatively evaluated. Consistent with previous findings, experimentation shows that ensemble and tree-based learning algorithms outperform standard classifiers in groundwater quality classification. Of the algorithms trained, XGBoost performed the best with a classification accuracy of 98%, followed by Polynomial SVM, whose accuracy was 97% and decision tree with 96%. The comparison provides benchmark performance for conventional machine learning models, which show potential for rapid, low-cost, data-driven monitoring of groundwater quality and drinking water management.

Article Details

Section
Articles