This article does not contain any studies involving human participants or animals performed by the authors. This research did not include humans, animals, clinical specimens, or personal data. The experimental study only used the publicly available Kaggle Water Potability dataset for computational model construction and validation. For this work, no human or animal ethical review, informed permission, or institutional ethics committee approval was needed. The public dataset data use conditions were followed for the study.
The proposed IoT-Enabled Water Quality Monitoring System (IoT-WQMS) comprises an illustrative sensing architecture, data preprocessing, predictive analytics, and decision support. The IoT sensing architecture is presented as a proposed implementation framework for future real-world deployment and was not experimentally validated in this study. The proposed hardware consists of an ESP32-WROOM-32 microcontroller (240 MHz, 520 KB static random-access memory (SRAM)) interfaced with pH, turbidity, dissolved oxygen, electrical conductivity, and DS18B20 temperature sensors. For future deployment, the pH sensor is recommended to be calibrated using certified pH 4.00, 7.00, and 10.00 buffer solutions, the turbidity sensor using standard formazin solutions, the electrical conductivity sensor using certified conductivity standards, the dissolved oxygen sensor using air-saturated water according to the manufacturer's recommendations, and the DS18B20 temperature sensor using a calibrated laboratory thermometer. Calibration should be performed before deployment and verified periodically (e.g., monthly or whenever sensor drift exceeds the specified tolerance). Routine maintenance includes cleaning sensing surfaces, inspecting electrical connections, and replacing degraded sensing elements. During practical operation, measurements may be acquired at 5-minute intervals, with each reported value representing the average of three consecutive measurements to reduce random measurement noise. Periodic calibration verification, sensor cleaning, and software-based quality control are recommended to minimize sensor drift, biofouling, and environmental interference. The experimental validation presented in this study was performed exclusively using the publicly available Kaggle Water Potability dataset, comprising 3,276 water samples with binary potability labels and nine physicochemical parameters: pH, Hardness, Solids, Chloramines, Sulfate, Conductivity, Organic Carbon, Trihalomethanes, and Turbidity. Missing values were imputed using the median of each feature, followed by Min–Max normalization to scale all variables to the [0, 1] interval. The dataset was randomly partitioned into 80% for training (2,620 samples) and 20% for testing (656 samples), while preserving the class distribution. Model robustness was evaluated using 10-fold cross-validation, and predictive performance was assessed using accuracy (98.21%), precision (97.86%), recall (98.04%), F1-score (97.95%), and ROC-AUC (0.991). The proposed framework achieved an average inference latency of 0.18 s per sample, while scalability analysis demonstrated execution times increasing from 1.8 s to 7.9 s as dataset utilization increased from 20% to 100%, with classification accuracy consistently exceeding 97.5%. The computational workflow was implemented using Python 3.10, Jupyter Notebook, Pandas 2.2, NumPy 1.26, Scikit-learn 1.5, and Matplotlib 3.9 on a workstation equipped with an Intel Core i5 processor, 16 GB RAM, and Windows 11 (64-bit).
Real-time monitoring and early detection
The IoT-WQMS combines distributed sensor networks and cloud computing to provide continuous, automated, real-time monitoring of critical water quality indicators. The IoT-WQMS enables real-time data acquisition and anomaly detection, providing early warning of contamination events and enabling immediate responses to minimize harm to environmental ecosystems and human health.
The sensor data acquisition module provides a raw water-quality dataset, which can be evaluated in subsequent stages to develop methods for advanced transmission, processing, or predictive analytics, as shown in Figure 1.
Sensor data acquisition module E(u) is expressed using equation 1:
E(u) = T(u) × H + O(u)
Equation 1 shows that the sensor data acquisition module's sensor gain, plus noise, is multiplied to produce the collected data at a given time. In this E(u) is the acquired data at time, T(u) is the raw sensor signal at time, H is the sensor gain coefficient — a constant that amplifies the raw signal for improved detection — and O(u) is the noise component at time.
Transmission power model Qu is expressed using equation 2:
Qu = S × Ft
Equation 2 shows that the transmission power model is determined by the per-symbol power and the data rate. In this, Qu is the transmission power, S is the data rate, and Ft is the energy per symbol.
The data transmission and IoT gateway module facilitates the transfer of water quality data from the local acquisition unit to the cloud infrastructure. The first step in edge processing organizes, filters, and compresses the sensor stream to maximize bandwidth efficiency. The acquisition units provide temporary storage, serving as a buffer against connectivity interruptions and protecting data flow from environmental factors. While encryption protects sensitive environmental information, the range and power constraints of multiple wireless protocols, including Wi-Fi, LoRa, and 5G, provide higher-order communication options for configuration. Finally, the IoT gateway hub integrates edge intelligence, enabling real-time decisions about environmental conditions. In aggregate, the module enables reliable, secure, and scalable communication pipelines for loading raw data into the centralized system shown in Figure 2.
Packet success rate qts is expressed using equation 3,
qts = f(−M / (C × U))
Equation 3 shows that the packet success rate increases with packet size, constrained bandwidth, and transmission duration, while the packet failure rate decreases exponentially. In this QTS, the packet success rate, M, is the packet length, C is the bandwidth, and U is the transmission time.
Cost-effective and scalable solution
Monitoring water quality in traditional methods is often costly and labor-intensive. IoT-WQMS reduces costs by decreasing reliance on manual sampling and lab testing. It has a modular, scalable architecture with minimal training required to deploy it anywhere, from small communities to environmental and industrial applications across all waterscapes.
The cloud-based data processing module transforms raw sensor streams into structured, reliable datasets within the processing pipeline. Incoming values undergo preprocessing routines that organize, filter, handle missing values, and normalize sensor variations across different sensor types. Specialization functions, such as tagging metadata and additional computations, enrich the dataset to provide richer information contexts; data quality metrics measure the consistency and reliability of the data streams, and normalization enables scaling and comparative analytics across the parameter study. Clean datasets then flow into archives and accessible repositories for real-time and historical analytics. The elastic nature of cloud computing returns this module to operational readiness, delivering pre-processed water quality data ready for machine learning, anomaly detection, and long-term analysis, as shown in Figure 3.
Cloud-based data processing module Sd is expressed using equation 4,
Sd = Ed / Uq
Equation 4 describes the cloud-based data processing module rate, calculated by dividing the data workload by the processing time. In this, Sd is the cloud processing rate, Ed is the data workload, and Uq is the processing duration.
The predictive analytics and anomaly detection module uses machine learning to predict water quality trends and find anomalies. Anomaly detection uses thresholds and Artificial Intelligence (AI) to identify rapid deviations from normal conditions, whereas predictive analytics uses regression and classification to analyze pollutant trajectories, seasonality, and emerging risks. Pollution or sensor failure triggers alerts. Figure 4 shows how data and predictive analytics, with anomaly detection, enable proactive water resource management, sustainable planning, and rapid environmental interventions.
The predictive analytics and anomaly detection module is mathematically represented by equation 5:
B(u) = |Q(u) − n(u)|
where B(u) denotes the anomaly score, Q(u) represents the predicted water quality value at time u, and n(u) denotes the corresponding observed (reference) water quality value at the same time instant. Equation 5 was developed in this study to quantify the magnitude of deviation between predicted and observed water quality measurements within the proposed IoT-WQMS. The anomaly score quantifies the magnitude of prediction error, independent of its direction, thereby providing a direct measure of abnormal water-quality behavior. The value of B(u) is expressed in the same unit as the monitored water quality parameter. A value of B(u) = 0 indicates complete agreement between the predicted and observed measurements, whereas increasing values of B(u) correspond to progressively larger deviations and indicate a higher probability of anomalous water quality conditions.
Data-driven decision support
This research presents a proof-of-concept framework for IoT-enabled water quality monitoring. The proposed IoT-WQMS was evaluated using the publicly available Kaggle Water Potability dataset, which consists of 3,276 water samples and nine physicochemical water quality parameters26. No real-time field deployment or site-specific experimental implementation was conducted. Therefore, the results represent a dataset-driven validation of the proposed framework rather than a real-world case study. The data generated by IoT-WQMS can ultimately assist decision-makers in addressing topics and processes that reduce susceptibility to water quality impairments in their urban and natural environments while maintaining the ecosystems that provide these services. Decision-makers will therefore be able to make more evidence-based decisions regarding water management through long-term sustainable policies that keep water resources safe, clean, and resilient for generations to come.
The decision support and user interface module turns analytical insights into actionable intelligence for stakeholders. Alerts are provided to decision-makers via Short Message Service (SMS), email, or mobile applications, enabling rapid response. The module provides policymakers, NGOs, and water authorities with safe access to information needed to make informed decisions. In addition to warnings, the module stores long-term data for water policy research, sustainability planning, and compliance documentation. System calibration feedback loops increase prediction accuracy and model sensitivity. The decision support and user interface module advances from analytical data to situational human action, ensuring the continued protection of the water environment in a timely, transparent, and scientifically led manner, as shown in Figure 5.
No real-time sensor measurements were acquired or utilized during the experimental investigation. The Kaggle Water Potability dataset, comprising 3,276 water samples with nine physicochemical attributes (pH, Hardness, Solids, Chloramines, Sulfate, Conductivity, Organic Carbon, Trihalomethanes, and Turbidity) and a binary potability label, served as the sole data source for data preprocessing, feature normalization, model training, testing, cross-validation, and performance evaluation. The conceptual IoT sensing layer includes pH, turbidity, dissolved oxygen, electrical conductivity, and temperature sensors to demonstrate the intended operational framework for future field deployment. Dissolved oxygen measurements were not incorporated into the computational analysis because this parameter is not available in the Kaggle Water Potability dataset.
Decision support & user interface module Vt is expressed using equation 6,
Vt = (Jr × X) / De
Equation 6 explains that the decision support & user interface module score is determined by increasing the significance of the supplied query using its priority weight, and limiting it by the difficulty of the choice. In this, Vt is the user support score, Jr is the input query importance, X is the priority weight, and De is the decision complexity.
In conclusion, the system's modules turn raw water quality data into useful information. Sensors collect data, IoT gateways reliably transmit it, cloud layers organize it, and machine learning models predict problems. Decision-support interfaces notify and inform stakeholders. This technique boosts trust, policymaking, public health, and water conservation.