PHGMTCV: Physics-Guided Multi-Task Computer Vision for Solar Panel Soiling Classification, Segmentation, and Uncertainty-Aware Prototype Retrieval
Keywords:
Physics-Guided Computer Vision; Multi-Task Learning; Solar Panel Soiling; Image Formation; Monte Carlo Dropout; Epistemic Uncertainty; Prototype Retrieval; Segmentation; U-Net; ResNet18;Abstract
Solar photovoltaic systems lose 1.5–6.2% of annual energy yield per 1 g/m² of surface soiling, yet existing automated soiling detection systems apply single-task classification without exploiting the physically grounded image formation processes that govern how distinct soiling types dry particulate deposition, cohesive mud adhesion, organic bird droppings, mineral water spot residues, and their mixtures manifest as characteristic optical signatures in surface reflectance imagery. This paper presents, a Physics-Guided Multi-Task Computer Vision framework(PHGMTCV) for solar panel soiling characterization that addresses this gap through four novel contributions aligned with the International Journal of Computer Vision’s scope of mathematical, physical, and computational aspects of vision: (1) a physics-inspired image decomposition branch that explicitly extracts luminance value, chromatic saturation, inter-channel chromatic residuals, and Laplacian gradient response as structured visual cues reflecting the reflectance and scattering physics of distinct soiling deposits; (2) a multi-task U-Net segmentation decoder sharing encoder representations with the classification head, trained under a joint Dice-CE plus Sobel boundary consistency loss; (3) exponential moving average prototype memory enabling similarity-based evidence retrieval and prototype alignment regularization; and (4) Monte Carlo dropout-based epistemic uncertainty estimation providing calibrated predictive confidence for deployment. Trained and evaluated on a physics-informed synthetic soiling dataset includes 6 classes, and 60 training samples per class, The proposed PHGMTCV framework achieves exceptional performance, attaining 100.00% classification accuracy and weighted F1-score, along with 99.12% mean IoU and a Dice score of 0.9923. Additionally, prototype retrieval yields a one-versus-rest AUC of at least 0.9989, demonstrating strong discriminative capability. The method consistently outperforms all eight compared state-of-the-art approaches by a margin of at least 1.33 percentage points in accuracy. Component-wise ablation further reveals a substantial overall improvement of 17.7% in accuracy over the baseline ResNet backbone, with the physics-guided branch contributing the most significant gain, improving accuracy by 6.8% and mean IoU by 8.1% through the integration of physically grounded image formation cues into visual feature learning.