This study did not involve human participants or vertebrate animals, or sampling of tissue. All data used for this research were synthetically generated using physical propagation models and publicly available meteorological parameters. Hence, no ethics approval from an Institutional Review Board (IRB) or Institutional Animal Care and Use Committee (IACUC) was needed.
Dataset Generation Based on the Physical Propagation Theory. The dataset was built to mimic hourly atmospheric conditions for a free-space optical communication system for a whole calendar year (2024) for Iraqi atmospheric conditions. We created a synthetic database with 1,500 samples per hour.
First, weather conditions were randomly assigned based on regional trends: clear sky (54.3%), dust (24.9%), fog (10.5%), rain (7.4%), and snow (2.8%). Secondly, the corresponding physical attenuation model was applied to each sample based on the weather condition, i.e. the Beer-Lambert law for clear sky, the Kim model for fog, the Carbonneau theory for rain and Mie scattering theory for dust storms. Third, the parameters of the FSO system were set as follows: transmission power of 20 dBm, wavelength of 1550 nm, transmission distance of 3 km, transmission aperture of 2.5 cm, and receiving aperture of 20 cm. Fourth, attenuation was computed in dB/km for each sample. Finally, the complete dataset was randomly divided into 1,200 training samples (80%) and 300 testing samples (20%). The simulated conditions include large dust concentrations in association with sandstorms, rainstorms and temperature changes from −4.89°C to 47.99°C. The weather conditions and parameter distributions were selected based on the Iraqi climate records of the period 2020–2024. The five weather regimes (clear sky, fog, rain, dust storms and snow) were chosen because they cover the full spectrum of atmospheric conditions that affect FSO attenuation in Iraq, with dust storms being especially prevalent in the Middle East. The historical data on meteorology gathered throughout Iraqi regions was used to create the probability distribution for each weather condition. The distribution that resulted was as follows: 54.3% of the sky is clear (the predominant state), 24.9% is dust (representing Iraq's sandstorm problem), 10.5% is fog (frequent in northern Iraqi winters), 7.4% is rain (low rainfall amounts typical for Iraq), and 2.8% is snow (sometimes in northern mountainous locations). The relevant meteorological parameters were modeled using probability distributions for each weather condition as follows: temperature was modeled using a normal distribution (mean 28.55±11.18°C) between −4.89°C and 47.99°C based on Iraqi seasonal extremes; humidity was modeled using a uniform distribution (mean 42.01±25.56%) from 0% to 100%; visibility was modeled using a log-normal distribution between 0.05 km and 29.99 km (mean 13.10±10.91 km) to account for the frequent low-visibility events during dust storms; dust concentration was modeled using an exponential distribution between 0 and 4.96 mg/m3 (mean 0.74±1.30 mg/m3) with higher probabilities for low concentrations and long tails for extreme dust events.
The communication system was designed with a transmission power of 20 dBm, a wavelength of 1550 nm, a transmission distance up to 3 km, a transmission aperture of 2.5 cm, and a receiving aperture of 20 cm to compensate for the divergence loss . The parameters of the FSO system were divided into two groups: fixed parameters that did not change for all the samples, and variable parameters that were changed during the generation of the dataset. For all 1,500 samples the following parameters were fixed: transmission power (20 dBm), operating wavelength (1550 nm), transmission aperture (diameter 2.5 cm, efficiency 0.7) and receiving aperture (diameter 20 cm, efficiency 0.7). These parameters were fixed because they are the physical specifications of the FSO system hardware and they do not change with weather conditions. The data set was created with 1500 samples with the following parameters varied: temperature (−4.89°C to 47.99°C), humidity (0% to 100%), visibility (0.05 km to 29.99 km), dust concentration (0 to 4.96 mg/m3) and weather condition (clear sky, fog, rain, dust, snow). These parameters were modified according to probability distributions derived from Iraqi climate records for the years 2020–2024. For each sample the attenuation value (dB/km) was calculated using the corresponding physical attenuation model, according to the specific combination of weather conditions and variable parameters.
Physical attenuation was modeled using the Carbonneau model for rain, Beer–Lambert law for clear, Mie scattering theory for dust, and the Kim model for fog24. The Beer-Lambert law applies for clear sky conditions where the attenuation is dominated by molecular scattering and absorption, which decrease exponentially with distance25. The extinction coefficient α at 1550 nm is due to the Rayleigh scattering by air molecules and the absorption by atmospheric gasses26. The Kim model is a fog-specific model that relates attenuation to visibility through empirical coefficients derived from fog droplet size distributions. The wavelength-dependent exponent q accounts for Mie scattering27. The main parameter of the Carbonneau model is the rainfall rate R as the rain attenuation depends on the size and density of the raindrops and the coefficients are empirically derived at 1550 nm and specifically calibrated for optical wavelengths28. Mie scattering theory is applicable to dust conditions, since the dust particle size (0.1–100 μm radius) is comparable to the wavelength (1550 nm), and the complex refractive index m = 1.55–0.005i for Middle Eastern dust includes both scattering and absorption29. The following physical attenuation models were implemented with their respective equations and parameter settings.
For clear sky conditions, the Beer-Lambert law was used:
Aclear = 10×log₁₀(e(α×d)) (1)
where α is the extinction coefficient (varied using a normal distribution centered at 0.02 dB/km with ±0.005 dB/km variation at 1550 nm under clear conditions), and d is the transmission distance (fixed at 3 km). For fog conditions, the Kim model was implemented using the equation:
Afog = 10×ln(10)/V×(λ/550)−q (2)
where V is visibility in kilometers (varied from 0.05km to 10km), λ is the wavelength in nanometers (fixed at 1550nm), and q is the particle size distribution coefficient calculated as: q=1.6 for V>50 km, q=1.3 for 6<V<50 km, q=0.585×V(1/3) for 1 <V<6km, q=0 for 0.5<V<1km, and q=0.5 for V<0.5km. For rain conditions, the Carbonneau model was used:
Arain=0.023×R0.93 (3)
where R is the rain rates in mm/h (varied between 0.25 and 50mm/h as per Iraqi rain records). The extinction efficiency relationship was used for the dust storm circumstances using the Mie scattering:
Adust=10×log₁₀(e(τ×L)) (4)
where τ=∫₀^∞ πr2Qext(r,λ,m)N(r)dr, r is the particle radius (0.1–100μm according to Iraqi dust composition), Qext is the extinction efficiency calculated using Mie theory, λ=1550nm, m=1.55–0.005i is the complex refractive index for Middle Eastern dust, and N(r) is the particle size distribution modeled using a log-normal distribution with geometric mean radius of 2.5 μm and standard deviation of 2.0. The attenuation model was implemented for snow conditions as follows:
Asnow = 0.1×S0.75 (5)
where S is the snowfall rate in mm/h (0.5–15 mm/h). This empirical equation was chosen based upon the work in the literature30, where attenuation models for optical propagation through snow were developed using Mie scattering theory applied to snowflake size distributions. The equation is valid for snow fall rates between 0.5 and 15 mm/h and assumes dry snow conditions with typical snow flake diameters of 1–10 mm. The coefficient 0.1 and exponent 0.75 were obtained from curve-fitting to Mie scattering30 calculations for snow at 1550 nm. It does not adjust for moist snow or combined precipitation, which may have variable attenuation properties, even though it offers a fair estimate for dry snow. Because the approach is computationally effective, often referenced in the FSO publications, and appropriate for predicted snow circumstances in northern Iraq (Kurdistan region in January and February), it was selected for this investigation. Using Numpy for numerical calculations, all models were implemented in Python 3.9. The matching model was used to the randomly chosen weather condition and sampled ambient data to calculate the attenuation value for every sample. The obtained weather distribution included 814 clear-sky conditions (54.27%), 375 dust events (25.00%), 157 fog events (10.47%), 111 rain events (7.40%) and 43 snow events (2.87%).
The examination of historical meteorological information gathered from Iraqi meteorological stations in several regions (Baghdad, Basra, Mosul, and Ramadi) between 2020 and 2024 was used to establish the proportions of weather situations. The Iraqi Ministry of Transportation and the Iraqi Meteorological Organization and Seismology (IMOS) supplied the original data. Daily weather records that recorded the current atmospheric conditions for each day were included in the data. Temperature (daily minimum, maximum, and mean), relative humidity, visibility, amount of rainfall, and dust storm occurrences were among the specific characteristics extracted from these records. The Iraqi government's open data portal (https://www.motrans.gov.iq/) provides access to a portion of the IMOS data; however, the particular records utilized in this study are not publicly stored in a central repository. The climate information used to calculate the percentages of weather conditions and parameter values is summarized in Table 1. Clear sky days were defined as days with no precipitation, visibility greater than 10 km, and no dust activity, making up 54.27% of the 1,825 recorded days. Dust storm days (including full dust storm (visibility < 1 km) and suspended dust (visibility 1–5 km)) accounted for 25.00% of days, indicating the high frequency of sandstorm events in Iraq’s arid and semi-arid climate. Days with visibility less than 1 km caused by the suspension of water droplets (excluding dust-induced visibility reduction) were classified as fog days. The percentage of fog days was 10.47%, and fog days were mainly in winter in northern Iraqi regions. Rain days, days with measurable precipitation >0.1 mm, were 7.40%, consistent with the low average of annual rainfall in Iraq of 150–200 mm per year. Snow days (days with accumulation of frozen precipitation) made up 2.87% of days and were confined to mountainous northern areas (Kurdistan region) in January and February. These proportions were later used as probability weights for random sampling in the dataset generation. Thus, the synthetic dataset reflects the real-world frequency of each weather condition in the Iraqi environment.
Bias considerations in synthetic data generation
To reduce possible bias several steps were taken:
(1) Select the distribution: The statistical properties of the source climate data were used to select probability distributions. Temperature was normally distributed with a mean and standard deviation as recorded by the IMOS. The humidity was uniformly distributed over the entire observed range (0-100%). Visibility was assumed to follow a log-normal distribution to account for the frequent occurrence of low-visibility events during dust storms. The dust concentration followed an exponential distribution where there were higher probabilities at low concentrations and long tails at extreme dust events31. This was consistent with the observed frequency of dust events in Iraq32.
(2) Proportions of weather conditions: Analysis of IMOS records for 2020–2024, comprising 1,825 daily observations across all four regions, yielded the proportions: 54.3% clear sky, 24.9% dust, 10.5% fog, 7.4% rain and 2.8% snow. Days without precipitation, visibility >10 km and no dust activity were defined as clear sky days. Days with dust storms included both full dust storms (visibility <1 km) and suspended dust (visibility 1–5 km). A fog day was defined as a day where visibility was less than 1 km and the cause was a suspension of water droplets (not dust). Rain days were defined as days with measurable precipitation >0.1 mm. Snow days were defined as days with accumulated frozen precipitation33.
(3) Ranges of parameters: The ranges of parameters were based on the observed extremes in the IMOS records: temperature ranged from −4.89 °C (Mosul, winter) to 47.99 °C (Basra, summer), visibility ranged from 0.05 km (severe dust storms) to 29.99 km (clear conditions), and dust concentration ranged from 0 to 4.96 mg/m3 (based on the observed maximum dust concentration during severe haboob events)34.
(4) Independence Assumptions: We assumed that the environmental parameters were independently sampled, which is a simplification of real-world conditions where atmospheric variables are correlated (e.g., high dust concentration often correlates with low visibility). In order to provide a controlled simulation environment for methodical model comparison, this independence assumption was adopted 35. These presumptions ramifications are covered in the Discussion.
(5) Stratified Splitting: The train-test split was stratified based on the weather condition category (clear sky, fog, rain, dust, snow) to ensure that the proportion of each weather condition in the training and testing sets corresponded to the original dataset distribution. This way the test set is not imbalanced concerning rare weather conditions (especially snow at 2.87%)36.
Acknowledgment of deterministic target generation
It is important to point out that the good predictive performance seen here may partly be due to the model learning or approximating the deterministic physical equations used to generate the synthetic target values37. Unlike real-world experimental measurements, which contain measurement noise, instrument errors, and unmodeled physical phenomena, the synthetic dataset provides a clean, noise-free relationship between the input features and the attenuation target. This is because the attenuation values were computed directly from the physical propagation models (Beer-Lambert law, Kim model, Carbonneau model, and Mie scattering theory) based on the input parameters. Therefore, the quantitative performance metrics (R2, RMSE, MAE) represent performance on equation-derived synthetic data and should not be interpreted as expected performance on noisy observational or experimental data. The results should be viewed primarily as a comparative evaluation of modeling methodologies in a controlled simulation environment38.
Complete feature set for model training
The training dataset had 10 input features for model training:
1. Temperature (°C)
2. Humidity (%)
3. Visibility (km)
4. Dust concentration (mg/m3)
5. Rainfall rate (mm/h)
6. Snowfall rate (mm/h)
7. Wind speed (m/s)
8. Atmospheric pressure (hPa)
9. Month (numeric, 1–12)
10. Season (one-hot encoded: spring, summer, autumn, winter)
Important Clarification: Weather conditions (clear sky, fog, rain, dust, snow) were used as a categorical variable for stratification during the dataset split and were not included as input features for any model. The SHAP analysis includes only the 10 features listed above. The season variable was one-hot encoded (4 categories: spring, summer, autumn, winter), and for the SHAP analysis, the contributions of the one-hot-encoded season variables were summed across seasons to produce a single season contribution value. This combined value represents the total contribution of all season-related variables to the prediction of attenuation. Before creating the summary figure, the four one-hot-encoded season columns were identified, and their SHAP values were added together for each sample. This method guaranties that the model's usage of season as a composite categorical variable is consistent with the SHAP analysis.
The main environmental factors that directly affected optical attenuation through physical mechanisms were features 1–6. The addition of features 7 and 8 (wind speed and pressure) as supplementary meteorological factors may have an indirect impact on attenuation by affecting air stability and aerosol dispersion. In order to account for seasonal variations in atmospheric conditions, features 9–10 (month and season) were included as temporal descriptors. The attenuation value (dB/km) was used as the target variable for all models. Key dataset statistics included temperature (28.55°C ± 11.18°C), humidity (42.01% ± 25.56%), visibility (13.10 ± 10.91 km; range: 0.05–29.99 km), dust concentration (0.74 ± 1.30 mg/m3; maximum: 4.96 mg/m3), attenuation (4.80 ± 7.20 dB/km; range: 0.09–50.93 dB/km), operating range (5.74 ± 1.97 km), and signal-to-noise ratio (64.88 ± 15.07 dB). The Operating Range and SNR were calculated from the attenuation values using standard FSO link budget equations.
Operating range calculation
The Operating Range (in km) was calculated using the link budget equation:
Prx=Ptx×Gt×Gr×(λ/(4πR))2×10(−A×R/10) (6)
where: Prx = received power (set to minimum sensitivity of −30 dBm); Ptx = transmission power (fixed at 20 dBm); Gt = transmitter gain (calculated from aperture sizes); Gr = receiver gain (calculated from aperture sizes); λ = wavelength (1550 nm); R = range in km; A = atmospheric attenuation in dB/km (calculated from the physical models).
Transmitter and Receiver Gains: The transmitter gain (Gt) was calculated as: Gt = 10×log₁₀[0.7×(π×0.025/1.55×10⁻6)2] ≈ 44.2 dBi. The receiver gain (Gr) was calculated as: Gr = 10×log₁₀[0.7×(π×0.20/1.55×10⁻6)2] ≈ 62.3 dBi. The transmission aperture was 2.5 cm in diameter, with an efficiency of 0.7. The receiving aperture was 20 cm in diameter, with an efficiency of 0.7. The equation was solved iteratively for R to determine the maximum achievable link distance for each attenuation value.
Signal-to-noise ratio calculation
The SNR (Signal-to-Noise Ratio) in dB was calculated using the equation:
SNR=Prx−10×log₁₀(kTB)−NF (7)
where: Prx = received power in dBm (calculated from the link budget); k = 1.38×10⁻23 J/K (Boltzmann’s constant); T = 290 K (receiver temperature); B = 109 Hz (receiver bandwidth, 1 GHz); NF = 3 dB (receiver noise figure). The noise floor was calculated as:
10 × log10(kTB) ≈ −84 dBm (8)
For each sample, after calculating the attenuation A using the appropriate physical model, the Operating Range was derived by solving the link budget for R, and the SNR was calculated from the resulting received power Prx at that range.
Weather-Specific Operating Range Values: The operating range varied across weather conditions: clear sky (7.12 ± 1.85 km), fog (5.81 ± 1.92 km), snow (5.42 ± 1.56 km), rain (3.81 ± 0.98 km), and dust (3.72 ± 1.08 km). A 3 dB margin was not applied in the current calculations; the operating range represents the theoretical maximum range without a system margin. The reported operating range (5.74 ± 1.97 km) is the overall mean across all weather conditions39.
Fixed Transmission Distance: The transmission distance in the physical attenuation models was set to 3 km. This is the link distance for which the attenuation calculations were performed. The reported operating range is the theoretical maximum distance computed using the link budget equation, which may differ from the fixed 3 km transmission distance. Weather-specific attenuation values were recorded for clear conditions (0.27±0.06 dB/km), fog (1.88±1.92 dB/km), snow (6.45±2.54 dB/km), rain (13.58±6.32 dB/km), and dust (13.10±7.32 dB/km). All quantitative values reported in this manuscript are presented as mean ± standard deviation (SD) unless otherwise specified40.
Cross-validation R2 for Random Forest is reported as 0.960±0.007. In certain cases, such as temperature (−4.89 to 47.99°C), visibility (0.05 to 29.99km), dust concentration (0 to 4.96 mg/m3), and attenuation (0.09 to 50.93dB/km), the range (lowest to highest) is given verbally. The dataset was split into subgroups for testing (300 samples; 20%) and training (1,200 samples; 80%). Stratified random sampling was used to carry out the train-test split. To make sure that the percentage of each weather condition in the training set (80%) and testing set (20%) matched the original dataset distribution, stratification was used based on the weather condition category (clear sky, fog, rain, dust, and snow). In particular, 1,200 (80%) of the 1,500 samples were allocated to the learning set and 300 (20%) to the test set. Samples were chosen at random for each category of meteorological conditions while preserving the original proportions: from the 814 clear sky samples (54.27%), 651 were assigned to training and 163 to testing; from the 375 dust samples (25.00%), 300 to training and 75 to testing; from the 157 fog samples (10.47%), 126 to training and 31 to testing; from the 111 rain samples (7.40%), 89 to training and 22 to testing; from the 43 snow samples (2.87%), 34 to training and 9 to testing. Random sampling within each stratum was performed using a random seed of 42 to ensure reproducibility. This stratified approach was chosen to prevent imbalanced representation of rare weather conditions (particularly snow at 2.87%) in the testing set, which could otherwise lead to unreliable performance evaluation for those conditions.
Machine learning model evaluation
Six machine learning methods were evaluated, including Support Vector Regression (SVR) with a radial basis function kernel (C = 100), K-Nearest Neighbors (KNN; k = 10, distance-weighted), RF (200 trees, maximum depth = 20), Extreme Gradient Boosting (XGBoost; 200 estimators, maximum depth = 10, learning rate = 0.1), Light Gradient Boosting Machine (LightGBM; 200 estimators, maximum depth = 10, learning rate = 0.1), and baseline Linear Regression. For all machine learning and deep learning models, hyperparameter tuning was performed for the most critical parameters, while default values were retained for unspecified parameters. For machine learning models, the following parameters were explicitly tuned using grid search with 5-fold cross-validation on the training set: 1) Random Forest: number of trees (tested: 50, 100, 150, 200, 250) and maximum depth (tested: 10, 15, 20, 25, no limit), with optimal values of 200 trees and depth 20 selected. 2) XGBoost: number of estimators (tested: 100, 150, 200, 250), maximum depth (tested: 6, 8, 10, 12), and learning rate (tested: 0.05, 0.1, 0.2), with optimal values of 200 estimators, depth 10, and learning rate 0.1. 3) LightGBM: identical tuning ranges were used, resulting in 200 estimators, depth 10, and learning rate 0.1. 4) SVR: the regularization parameter C (tested: 1, 10, 50, 100) and kernel coefficient gamma (tested: ‘scale’, ‘auto’, 0.1, 0.01) were tuned, with optimal C = 100 and RBF kernel. 5) KNN: the number of neighbors k (tested: 3, 5, 7, 10, 15) was tuned, with optimal k = 10 and distance-weighted voting enabled.
All other parameters for these models were left at their default values as defined in scikit-learn (see Table of Materials for version; e.g., Random Forest: bootstrap=True, min_samples_split=2, min_samples_leaf=1; XGBoost: subsample=1.0, colsample_bytree=1.0, gamma=0). For deep learning models, the architecture (number of layers and units per layer) and dropout rate (20%) were manually tuned through iterative experimentation on the validation set, while the optimizer (Adam), initial learning rate (0.001), early stopping patience (20 epochs), and learning rate reduction parameters (factor 0.5, patience 10) were set based on standard practices in the literature and kept fixed across all deep learning experiments.
Climate data sources
The historical meteorological information gathered from Iraqi weather stations in several places (Baghdad, Basra, Mosul, and Ramadi) between 2020 and 2024 was used to calculate the weather state proportions and variable distributions. The Iraqi Ministry of Transportation and the Iraqi Meteorological Organization and Seismology (IMOS) provided the raw data. Daily weather records detailing the prevalent atmospheric state for each day were incorporated in the data. Specific variables obtained from these records included temperature (daily minimum, maximum, and mean), relative humidity, visibility, rainfall amount, and dust storm occurrences. The IMOS data are partially available through the Iraqi government’s open data portal (https://www.motrans.gov.iq/), though the specific records used in this study are not publicly archived in a centralized repository. A summary of the climate data used to determine the proportions of weather conditions and parameter ranges is provided in Table 1.
Fivefold cross-validation was used on the training set (1,200 samples) to tune hyperparameters and estimate performance for all machine learning models. All input variables (temperature, humidity, visibility, dust concentration, rainfall rate, snowfall rate, wind speed, pressure) were feature scaled by standardization (Z-score normalization): x_scaled = (x − μ)/σ, where μ and σ are the mean and standard deviation of the training set. We standardized within each cross-validation fold using only statistics from the training fold to avoid data leakage. Tree-based models (Random Forest, XGBoost, LightGBM) are scale-invariant, but the same standardization was applied for consistency across all machine learning models. Min-max normalization was used for deep learning models: x_scaled = (x−x_min)/(x_max−x_min), which scales features to [0, 1] based on min/max values from the training set. Bounded inputs lead to faster convergence of neural nets, which is why they were chosen. The test set was scaled using the parameters obtained from the training set, and was not used for model selection or hyperparameter tuning.
Complete performance metrics were recorded, including test coefficient of determination (R2), root mean square error (RMSE), mean absolute error (MAE), cross-validation R2, and training time. The training times for all machine learning and deep learning models are reported in seconds (s) for the faster models (Linear Regression, KNN, SVR, Random Forest, XGBoost, LightGBM) and in minutes (min) for the slower models (deep learning architectures). All models were trained on the same computational environment to ensure fair comparison41.
The training time was measured using the Python time module, i.e., the elapsed wall-clock time from the start to the end of the model-fitting function, excluding the time required for data loading and preprocessing. Training time for a deep learning model is the time taken to complete all epochs until early stopping. This includes forward propagation, backward propagation, and validation checks. All experiments were performed with the system running no other computationally intensive processes to obtain consistent timing measurements. The times are the average of 5 independent runs (standard deviations)42.
Deep learning model evaluation
Six deep learning architectures were evaluated using GPU acceleration, including a Multilayer Perceptron (MLP; 64-32-16), a Deep Neural Network (DNN) with batch normalization (128-64-32-16), a Long Short-Term Memory network (LSTM; 64-32 units, sequence length = 10), a one-dimensional convolutional neural network (1D-CNN), a CNN–LSTM hybrid model, and an attention-based network. All deep learning models were implemented using TensorFlow with the Keras API and run with GPU acceleration (see Table of Materials for hardware/software versions). The 1D-CNN architecture consisted of three convolutional layers (64, 128, and 256 filters, kernel size 3, ReLU activation, padding=’same’), two MaxPooling1D layers (pool size 2), a GlobalAveragePooling1D layer, a Dense layer with 128 units and ReLU activation, a Dropout layer (0.2), and a Dense output layer (1 unit, linear activation), totaling approximately 245,000 trainable parameters. The CNN-LSTM hybrid architecture accepted input sequences of 10 time steps with 5 features, using two Conv1D layers (64 and 128 filters, kernel size 3, ReLU, padding=’same’), a MaxPooling1D layer (pool size 2), two LSTM layers (64 and 32 units, return_sequences=False), Dropout layers (0.2), a Dense layer (32 units, ReLU), and a Dense output layer (1 unit, linear activation), totaling approximately 198,000 trainable parameters. The attention-based network used a multi-head attention mechanism with 4 heads (key and value dimensions of 64), where the input was projected to 64 dimensions, followed by scaled dot-product attention (formula: Attention(Q, K, V) = softmax(QKT/√d_k)V), residual connections, layer normalization, a feedforward network (128→64 units), global average pooling, Dropout (0.2), a Dense layer (32 units, ReLU), and a Dense output layer (1 unit, linear activation), totaling approximately 167,000 trainable parameters43.
All models used early stopping (patience = 20), learning-rate reduction (factor = 0.5, patience = 10), dropout (20%), and the Adam optimizer (learning rate = 0.001). For all deep learning models, the batch size was set to 32 samples, the maximum number of training epochs was 200 with early stopping (patience = 20, restoring best weights), and the loss function was mean squared error (MSE). The train–validation split was as follows: from the original 1,200 training samples (after the 80/20 train-test split), 80% (960 samples) were used for training and 20% (240 samples) for validation. We stratified the train-validation split by weather condition to preserve the distribution. The validation set was used only for early stopping, learning rate decay, and monitoring overfitting; it was never used for model selection or hyperparameter tuning beyond these automated procedures. We did not hold out a separate validation set for the machine learning models; instead, we used five-fold cross-validation on the 1,200 training samples to tune hyperparameters and estimate performance44.
Justification for LSTM and CNN–LSTM architecture evaluation
The main dataset consists of independently generated weather samples, but we also tested LSTM and CNN–LSTM architectures for the following reasons: (1) real-world atmospheric conditions are temporally autocorrelated and testing sequence-based models allows us to determine whether capturing such dependencies could improve prediction accuracy; (2) recent research in atmospheric prediction has demonstrated the potential value of sequential architectures for modeling the temporal evolution of meteorological parameters34; (3) testing a diverse range of architectures ensures a comprehensive comparison of methodological approaches, which is a key contribution of this study; and (4) the CNN–LSTM hybrid architecture combines spatial feature extraction with temporal modeling, which may be beneficial for capturing the complex interactions among multiple atmospheric variables45.
Data formatting for sequential model input
For the sequential architectures (LSTM and CNN–LSTM), the input data was restructured from independent samples into pseudo-sequences using a sliding-window approach. In particular, the 1,200 training samples were first grouped into weather condition categories to preserve physical consistency. Within each weather category, the samples were ordered by their generated timestamps (simulated hourly observations for calendar year 2024). Then a sliding window of length 10 was applied to produce input sequences of 10 consecutive time steps (each with 5 features: temperature, humidity, visibility, dust concentration, and rainfall rate) to predict the attenuation at the 11th time step. This method preserves the temporal ordering of the simulated observations, while allowing sequential models to learn temporal dependencies. The structure of the test set was the same except for the same window size and feature set. We recognize that this pseudo-sequential structuring is a methodological simplification and does not reflect real-world temporal dynamics. We have acknowledged this as a limitation in the Discussion section.
Hybrid approaches evaluation
Three hybrid approaches were examined. The first approach was a Voting Ensemble that averaged predictions from the Random Forest, XGBoost, and Deep Neural Network models using equal weights (each model assigned a weight of 1/3), with the final prediction calculated as:
ŷensemble=(1/3)ŷRF+(1/3)ŷXGB+(1/3)ŷDNN (9)
Equal weighting was chosen to avoid introducing additional hyperparameters and to evaluate the baseline ensemble performance without bias toward any individual model. The second approach employed Ridge meta-learner stacking. The base learners were Random Forest, XGBoost, and a Deep Neural Network (Attention-based). The stacking procedure involved two stages: first, each base learner was trained on the full training set of 1,200 samples using 5-fold cross-validation to generate out-of-fold predictions, creating a new meta-feature matrix of size 1,200×3 (one prediction per base model per sample). Second, a Ridge regression meta-learner (L2 regularization parameter alpha=1.0) was trained on these meta-features, using the original attenuation values as the target, to learn optimal combination weights for the base learners. The final stacking prediction was:
ŷstacking=wRF×ŷRF+wXGB×ŷXGB+wDNN×ŷDNN (10)
where the weights w were learned by the Ridge meta-learner. The third approach was a Physics-Informed Neural Network that combined 70% of the neural network’s predictions with 30% from the Kim model for fog-condition samples. The combination was performed by fixed weighted averaging using the following formula:
ŷhybrid=0.7×ŷneural+0.3×ŷKim (11)
where ŷneural is the output of the Attention-based neural network, and ŷKim is the attenuation calculated from the Kim fog model based on visibility input. For non-fog samples, the physical part was set to 0, and the model was run as a pure neural network. The weights (70% neural and 30% physics) were fixed based on preliminary experimentation on the validation set (not the test set) where we tested the weight combinations of 90:10, 80:20, 70:30, 60:40, and 50:50. The 70/30 split was chosen as it provided the best validation R2 and still maintained enough physical constraint from the Kim model to regularize predictions and avoid physically implausible outputs, especially in fog where the Kim model provides established theoretical attenuation bounds.
Feature importance and interpretability analysis
The Random Forest model with impurity-based feature importance (variance reduction) was used to extract all 10 input feature importance rankings. The analysis showed that dust concentration (67.3%) and visibility (21.2%) were the most important predictors, together explaining 88.5% of the total predictive importance. The third most important feature was rain rate (6.0%), followed by wind speed (2.1%), temperature (1.5%), humidity (0.9%), month (0.5%), season (0.3%), snowfall rate (0.1%), and atmospheric pressure (0.1%). The low importance scores for the temporal features (month and season) indicate that seasonal variation in atmospheric attenuation is captured primarily by underlying environmental parameters rather than by time-based patterns alone.
A SHAP (Shapley Additive exPlanations) analysis was performed to evaluate the relationships between environmental factors and attenuation. The SHAP implementation used was the TreeExplainer module from the SHAP library, which is specifically optimized for tree-based models, including Random Forest, XGBoost, and LightGBM (see Table of Materials for version). The configuration for the SHAP analysis was as follows: the trained Random Forest model was passed to the TreeExplainer, which computed SHAP values using the interventional (marginal) feature attribution approach based on the conditional expectation of the model’s output. SHAP values were computed for all 300 test set samples, making a matrix of size 300 × 10 (one SHAP value per feature per sample). For each feature, the SHAP value represented its contribution to the prediction relative to the baseline (the average model prediction). Negative SHAP values showed a downward shift; whilst positive SHAP values showed that the feature boosted the attenuation prediction. The strength of the contribution was indicated by the SHAP value's magnitude. The distribution of SHAP numbers for each feature (using beeswarm plots), the direction of influence (the correlation between feature values and SHAP values), and feature importance rankings were all visualized using summary plots. The built-in plotting functions of the SHAP library—shap.summary_plot() for the beeswarm plot and shap.bar_plot() for global feature importance—were used to create all SHAP visualizations.
Handling of One-Hot Encoded Variables: Four binary columns (spring, summer, autumn, and winter) were first used to encode the season variable. In order to create a single "season" contribution value per sample for the SHAP analysis, the contributions of these four one-hot-encoded variables were combined by adding the SHAP values for each season category. To accomplish this grouping, all columns that corresponded to the one-hot-encoded season groups were found, their SHAP values were extracted for each sample, and they were then summed element-wise. The season's overall contribution to the attenuation prediction is represented by the ensuing combined SHAP values. This method permits a single "season" row in the SHAP summary plot and guaranties consistency with the model's usage of season as a composite categorical variable. Since the combined number offers a more comprehensible depiction of the season's total contribution, the season SHAP values were not shown separately for each season category.