Exploring data-driven multivariate statistical models for the prediction of solar energy

Publications

Exploring data-driven multivariate statistical models for the prediction of solar energy

Year : 2024

Publisher : Elsevier

Source Title : Computer Vision and Machine Intelligence for Renewable Energy Systems

Document Type :

Abstract

The global energy demand has been increasing exponentially due to population growth, modern lifestyle, and advancement of consumer technology. Energy technologies are currently moving toward renewable energy sources to reduce the impact of global warming. Solar energy is one of the prominent energy sources widely used due to its high-power density and ubiquitous characteristics. It is adopted in a range of versatile applications, among which smart grids, Internet of Things, consumer electronics, and smart agriculture are some major applications. However, dependability of the performance of solar panels on weather conditions is still considered a major drawback in this domain. Several techniques, including machine learning and deep learning, have been implemented to predict the solar energy in the long term and in the short term in the recent past. However, implementing these frameworks in the field requires sophisticated hardware and a significant amount of power. In this chapter, the performance of several multivariate statistical models, such as vector autoregression, vector autoregressive moving average, vector error correction model, mean variance regularization, Bayesian linear regression, and light gradient boosting machine (LGBM), have been investigated to predict the output power of the solar panel. The models have been trained and tested with a publicly available dataset. Principal component analysis has been implemented as feature selection technique for selecting important features from the dataset. LGBM outperforms all other statistical models by achieving a maximum R2 score of 0.84 and a minimum mean square error of 0.15. Subsequently, artificially missing data maximum of up to 15% has been created, which are later imputed using several interpolation techniques, such as linear, cubic spline, pad, and nearest. Attempts have been made to analyze the performance of the models with missing or corrupted data to evaluate the robustness of the models to handle them in the dataset.