Machine Learning Models for Predicting Persistent Organic Pollutants in Water Bodies
Keywords:
Persistent Organic Pollutants (POPs); Machine Learning; Water Quality Prediction; Environmental Monitoring; Artificial Intelligence; Environmental Risk AssessmentAbstract
Persistent Organic Pollutants (POPs) are among the most hazardous environmental contaminants due to their persistence, bioaccumulative nature, long-range atmospheric transport, and adverse impacts on aquatic ecosystems and human health. Conventional monitoring techniques for POPs primarily rely on laboratory-based analytical methods, which are accurate but expensive, time-consuming, and spatially limited. Recent advances in machine learning (ML) offer a promising alternative by enabling predictive modeling using environmental, hydrological, meteorological, and anthropogenic datasets. This study proposes an integrated machine learning framework for predicting the concentration and spatial distribution of persistent organic pollutants in water bodies using supervised learning algorithms, ensemble methods, and deep learning architectures. The framework incorporates physicochemical parameters, land-use characteristics, climatic variables, industrial discharge records, and remote sensing data to improve prediction accuracy. Model performance is evaluated using statistical indicators including Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), coefficient of determination (R²), and cross-validation techniques. The proposed approach supports early contamination detection, environmental risk assessment, and evidence-based water resource management. The study demonstrates that data-driven predictive analytics can significantly enhance monitoring efficiency while reducing operational costs, thereby contributing to sustainable water quality management and informed environmental policymaking.