Three-step framework for PV generation and losses

A three-step framework was developed to estimate facility-level PV energy generation and its associated major losses from clouds and aerosols globally. This framework introduces two main advances over previous methods: it combines multiple sources to assemble a more complete global PV dataset and it improves accuracy by decoupling the initial identification of PV sites from the precise extraction of panel footprints.

A key challenge is that global-scale detection is inherently difficult because PV installations must be distinguished from diverse non-PV surfaces, often leading to boundary errors. For this reason, separating the initial detection from the final segmentation improves accuracy over single-stage models that perform both tasks simultaneously. Our process therefore begins with Step 1: identifying candidate PV locations using a combination of existing inventories10,32,33, crowd-sourced records and a convolutional neural network (CNN) classifier trained on global-scale Sentinel-2 imagery.

This is followed by Step 2, where, in contrast to broad detection, segmentation applied to satellite imagery with identified PV sites is more tractable, as panels show distinct visual and spectral contrast with their surroundings. Therefore, high-resolution imagery from confirmed sites was processed using the SAM35, developed by Meta, to extract precise panel boundaries using a ‘few-shot’ learning approach. To evaluate the spatial accuracy of the resulting dataset, comparisons were then made with two existing global PV inventories, as described in Supplementary Table 1.

In Step 3, these high-accuracy panel footprints were integrated with atmospheric data from the Modern-Era Retrospective Analysis for Research and Applications (MERRA-2) re-analysis in a validated PV model. This allowed the estimation of a time series of facility-level electricity generation, and the distinct losses caused by clouds and aerosols for each of the 140,945 facilities in our database.

Identification of PV facilities from multiple data sources

We first identified potential PV facilities using OpenStreetMap (OSM), a global crowd-sourced geospatial digital database. While OSM data are extensively used for mapping and spatial analysis, its data quality and completeness vary by region. In particular, PV facilities may be mapped at the construction area level, often overestimating panel coverage and including planned rather than operational sites. Candidate PV facilities were extracted by querying OSM with the attributes: power=plant, plant:source=solar and plant:method=photovoltaic. Global queries were regionally partitioned and run in parallel on UK’s JASMIN supercomputer, which provided the necessary computational power for large-scale geospatial processing.

To supplement OSM, we integrated publicly available PV datasets generated by machine learning-based analyses of satellite imagery. One such resource is the global inventory by Kruitwagen et al.10, which used a U-Net model trained on Sentinel-2 imagery to detect PV installations worldwide as of 2019. Their dataset consists of polygons outlining each identified PV site. From these polygons, we extracted the central coordinates and bounding box of each potential site to guide the next stage of our analysis.

For China, Feng et al.33 employed a random forest classifier on Google Earth Engine using 2020 Sentinel-2 imagery, training province-specific models to improve detection accuracy. This dataset captured older installations missing from global inventories but introduced more false positives. After clustering the 10-m resolution GeoTIFF PV masks into individual facilities, we extracted each PV facility’s central coordinates and bounding information. For India, we incorporated the PV dataset from Ortiz et al.32, which combined a U-Net detection model with hard negative mining to improve detection accuracy. This database includes 1,363 validated and grouped PV installations updated to 2020.

To identify new or previously unrecorded PV installations, we trained a U-Net model with a VGG1657 backbone on Sentinel-2 imagery. A total of 12,000 PV installations were manually annotated to build a globally representative training dataset. To improve detection across diverse environments, the model was trained on PV sites situated in deserts, croplands, mountains, floating systems and built environments. Further details on model architecture, training data preparation and deployment are provided in Supplementary Note 1.

By integrating the OSM records, published PV inventories and detections from our custom model, we ultimately identified 326,423 potential PV polygons globally.

PV panel segmentation with SAM

Estimating PV generation from satellite imagery requires accurately extracting the panel area that receives solar irradiance. This, in turn, depends on the precise extent of the PV panels themselves, excluding non-panel features and empty spaces often included within the facility polygons derived above. Including these non-PV areas leads to overestimates of panel surface area. Applying uniform scaling factors to correct such errors is unreliable because these ratios are neither constant nor reliably correlated with location or facility size.

The SAM35 is a foundation model trained on 11 million images and 1.1 billion masks. Designed to generalize effectively across diverse image datasets, it can be adapted to new imagery and unfamiliar objects with little or no additional training. It generalizes to new imagery using prompting techniques, where user-provided inputs (such as text descriptions or spatial information) guide the model’s segmentation. SAM consists of an image encoder that computes embeddings, a prompt encoder that processes user inputs and a lightweight mask decoder that generates segmentation masks. SAM has been applied across diverse domains, including three-dimensional object segmentation, medical imaging and crop disease identification58. Here, SAM was used to segment PV arrays from pre-identified locations on satellite images, using spatial prompts in the form of bounding boxes and individual points to initiate the segmentation process (Supplementary Fig. 3).

Red, green and blue (RGB) composites were generated from Sentinel-2’s 12-band imagery. At this local scale, PV panels and their immediate surroundings are readily distinguishable in an RGB composite. Areas of persistent cloud cover were filled by incorporating PlanetScope RGB imagery in those areas. With daily global coverage at a spatial resolution of 3 m, PlanetScope data offered more opportunities for cloud-free views and improved segmentation accuracy, especially for smaller PV arrays. High-resolution Google imagery was used where it was available and up to date.

Spatial prompts included bounding boxes and point annotations placed on PV panels. Segmented PV facilities were manually reviewed to ensure accurate boundaries. False detections and non-PV features were corrected using additional prompts if necessary. This two-stage approach ultimately produced segmented PV polygons for all sites previously identified. To consolidate closely located polygons into coherent PV facilities, those within 50 m of one another were merged. This final aggregation resulted in a comprehensive global database of 140,945 distinct PV facilities.

PV installation time from time series classification

To determine the installation time of each PV facility, we applied a time series classification approach using the sktime machine learning framework59. Monthly composites of Sentinel-2 top of atmosphere reflectance from 2017 to 2024 were generated, and for each PV polygon, the mean reflectance across all pixels and spectral bands was extracted. This resulted in a multivariate time series representing the temporal reflectance signature of each facility. A manually labelled training dataset was assembled using PV sites with known installation times, identified from high-resolution satellite imagery. A supervised classifier was trained to detect installation events on the basis of changes in reflectance patterns over time. The model was optimized using a large and diverse training set to distinguish true PV installation signals from common pre-installation changes, including grading, land clearing and temporary construction activities. The trained model was applied to all detected PV polygons to estimate installation time. As full data were not available until 2017, all PV facilities installed before 2017 were assigned an installation time of 2017. Example reflectance trajectories used to support classification are shown in Supplementary Fig. 4, with corresponding visual examples of detected facility expansion shown in Supplementary Fig. 5.

Estimation of PV power-generation potential

The theoretical electricity generation of each identified PV facility was estimated using the following equation:

$$E\,=\,I\,\times \,A\,\times \,k\,\times \,C,$$

(1)

where E is the potential alternating current (a.c.) power output (kW), I is the available solar irradiance at the PV surface (kW m−2), \(A\) is the effective PV panel area (m2) derived from our database, k is the coefficient accounting for system-level energy losses and C is the module conversion efficiency from solar radiation to electricity.

The parameter I depends on both atmospheric conditions and the facility’s geographic orientation and tilt. Atmospheric conditions, including aerosol loading and cloud cover, can alter the balance of direct and diffuse solar radiation that reach the Earth’s surface, influencing the global horizontal irradiance (GHI). The orientation and tilt angle of PV panels are optimized to maximize solar exposure over seasonal and diurnal cycles.

GHI was estimated using MERRA-260, the latest atmospheric reanalysis produced by NASA’s Global Modeling and Assimilation Office. MERRA-2 provides an hourly, globally gridded dataset (0.5° × 0.625°) of atmospheric and surface variables, including radiation, temperature, relative humidity and wind speeds. MERRA-2 assimilates aerosol measurements from spaceborne observations, representing their interactions with other physical processes. For this analysis, the surface incoming shortwave flux (SWGDN) from MERRA-2 M2T1NXRAD data products was used. SWGDN represents GHI, including both direct and diffuse components.

For a PV facility with fixed orientation and tilt angle, the optimal tilt was estimated from its latitude61. To translate GHI on a horizontal surface into the solar irradiance available on a tilted panel, GHI was decomposed into direct normal irradiance (DNI) and diffuse horizontal irradiance (DHI)62

$$\mathrm{GHI}\,=\,\mathrm{DNI}\,\times \,\cos \,\theta \,+\,\mathrm{DHI},$$

(2)

where \(\theta\) is the solar zenith angle. The diffuse fraction (\(\mathrm{DHI}/\mathrm{GHI}\)) is determined using the ratio of global to extraterrestrial irradiance on a horizontal plane, as implemented in the open-source pvlib library63. Using GHI components and PV geometry, irradiance on optimally tilted panels was derived using pvlib.

The effective panel area \(A\) of a PV depends on both the overall facility footprint and the spacing between individual panel arrays. This relationship is commonly expressed using a ground coverage ratio, defined as the ratio of panel-covered area to total facility land area: \(A\,=\,{A}_{{\rm{L}}}\,\times \,r\), where \({A}_{{\rm{L}}}\) is the total land area and \(r\) is the ground coverage ratio. Applying this calculation directly to PV polygons derived from traditional machine learning methods can be problematic, as satellite-based detections often overestimate panel areas by including non-PV features. By contrast, this two-stage search and extraction approach substantially improves detection accuracy by excluding non-PV elements during the extraction phase. This improved detection enables the use of a uniform ground coverage ratio \(r\) globally to estimate the active panel area.

The coefficient \(k\), which accounts for system-level losses, follows a model validated by Saxena et al.64 on a large sample of operational PV panels. This factor includes losses from suboptimal PV orientation (0.5%), mismatch in power points (0.3%), d.c. cabling (2%), a.c. cabling (0.5%), d.c.-to-a.c. conversion (2.2%), module downtime (0.5%), soiling by dust and snow (3.5%), and transformer inefficiencies (0.9%). Collectively, these factors amount to a total loss of 10.09%.

The conversion efficiency \(C\) represents the fraction of incident solar energy converted into electricity by the solar cells. For conventional silicon-based technologies, this efficiency typically ranges between 18% and 22%. A fixed value of 20% was used in this study.

To isolate the direct climate impacts that reduce PV power generation, MERRA-2’s SWGDN estimated under clear-sky and clean-sky conditions is used. SWGDNCLR represents shortwave flux under clear-sky (cloud-free) conditions. Under such conditions, one can estimate facility-level PV power output solely for clear-sky scenarios:

$${E}_{\mathrm{clear}}\,=\,\mathrm{SWGDNCLR}\,\times \,A\,\times \,k\,\times \,C.$$

(3)

Although MERRA-2 does not directly provide incoming shortwave flux under clean-sky (aerosol-free) conditions, it can be approximated by scaling the total SWGDN with the ratio of clean-sky net shortwave flux (SWGNTCLN) to full-sky net shortwave flux (SWGNT). This approach removes aerosol contributions and yields \({E}_{\mathrm{clean}}\), the estimated PV output in the absence of aerosols

$${E}_{\mathrm{clean}}\,=\,\frac{\mathrm{SWGNTCLN}}{\mathrm{SWGNT}}\,\times \,\mathrm{SWGDN}\,\times \,A\,\times \,k\,\times \,C.$$

(4)

By comparing the baseline output \(E\) (from all-sky conditions) with \({E}_{\mathrm{clear}}\) (no clouds) and \({E}_{\mathrm{clean}}\) (no aerosols), cloud and aerosol impacts can be separately quantified. Finally, the total climate-induced loss \(L\) can be computed using the following equation:

$$L\,=\,\left(\frac{{E}_{\mathrm{clean}}\,-\,E}{E}\,+\,\frac{{E}_{\mathrm{clear}}\,-\,E}{E}\right)\,\times \,100 \% .$$

(5)

To illustrate the structure of the resulting dataset, Supplementary Fig. 9 shows histograms of estimated annual generation per facility in 2023 for the globe and major regions (China, the USA, Europe, including the UK, India and the rest of the world).

Calculation of PV-effective AOD and gas concentrations

To evaluate the impact of atmospheric pollution on PV generation, aerosol conditions were quantified by calculating effective AOD for each PV facility. Major anthropogenic aerosol emissions originate from fossil-fuel combustion, particularly from coal-fired and other thermal power plants, industrial processes, vehicular transport and residential heating, as well as from biomass burning. Natural sources include desert dust, sea salt and volcanic activity65. AOD was obtained from the MERRA-2 reanalysis (M2T1NXAER collection), which provides hourly AOD at 550 nm on a 0.625° × 0.5° grid. The accuracy of the MERRA-2 AOD product has been extensively validated in previous global and regional studies66. These evaluations demonstrate close agreement with AERONET and MODIS observations, with typical correlations of 0.7–0.9 and root mean square error values below 0.2, confirming that MERRA-2 AOD provides reliable and stable estimates suitable for global and regional analyses. To further assess consistency in our study, we performed an effective-AOD trend analysis for all PV facilities in China using the 1-km Multi-Angle Implementation of Atmospheric Correction (MAIAC) AOD product67 (Extended Data Fig. 2). The MAIAC-based results show a comparable declining trend to that derived from MERRA-2, with a reduction rate of −0.0085 yr−1 (−2.9% yr−1) during 2013–2023. This difference can be partly attributed to MAIAC’s finer spatial resolution, which better captures local aerosol gradients around coal plants and aligns with the PV–coal distance distribution shown in Supplementary Fig. 8, where most PV facilities are located 20–30 km from the nearest coal power plant, below the MERRA-2 grid spacing. These findings indicate that the AOD decline and its implications for PV performance are robust across independent datasets.

To represent trace-gas pollutants, we further derived tropospheric NO2 and PBL SO2 column densities from the spaceborne Ozone Monitoring Instrument (OMI). For NO2, we used the OMI_MINDS_NO2d product and extracted the ‘ColumnAmountNO2TropCloudScreened’ variable, which represents daily tropospheric vertical column densities (molecules per square centimetre) filtered for high-quality observations with cloud fractions below 0.3 and solar zenith angles under 85°. For SO2, we used the OMSO2e product, which provides daily global estimates of SO2 column density in the PBL at 0.25° × 0.25° resolution.

We computed effective pollutant values to reflect their contribution to generation reductions at each PV facility. Two weighting factors were applied: (1) clear-sky irradiance weighting, which assigns greater weight to pollutants present during high-irradiance conditions (for example, daytime and summer), and (2) PV area weighting, which accounts for the relative size of each facility. The effective pollutant value was computed as

$${P}_{\mathrm{eff}}\,=\,\frac{{\sum }_{i\,}\left({P}_{i}\,\times \,{I}_{i}\,\times \,{A}_{i}\right)}{{\sum }_{i\,}\left({I}_{i}\,\times \,{A}_{{\rm{i}}}\right)},$$

(6)

where Pi is the pollutant value (AOD, tropospheric NO2 or PBL SO2) in pixel i, Ii is the clear-sky irradiance and Ai is the PV panel area.

Aerosol source attribution using GEOS-Chem

Aerosol source attribution was carried out using GEOS-Chem version 14.6.368 in the aerosol-only configuration, which simulates the transport and removal of aerosol species from anthropogenic and natural sources (sulfate, nitrate, ammonium, black carbon, organic carbon, dust and sea salt) using archived oxidant fields. Anthropogenic emissions used in the GEOS-Chem simulations are taken from the Community Emissions Data System (CEDS) inventory (version 2025-04), which provides an internally consistent global emissions dataset widely used in chemistry-transport modelling. China-specific inventories such as the Multi-Resolution Emission Inventory for China (MEIC) are also available and may offer more detailed representations of recent emission changes associated with national clean-air policies69. Simulations were conducted at 0.5° × 0.625° horizontal resolution with 47 vertical layers extending to 0.01 hPa and were driven by MERRA-2 meteorological fields provided at the same native resolution (0.5° × 0.625°, 72 levels). The model was run for the year 2023, following a 6-month spin-up for the global (2° × 2.5°) simulation and an additional 6-month spin-up for the nested China simulation (0.5° × 0.625°). Lateral boundary conditions for the nested domain were updated every 3 h from the corresponding 2° × 2.5° global run.

Two categories of simulations were performed (Supplementary Fig. 10). The base simulation included all anthropogenic and natural emissions and represents total atmospheric aerosol loading. To isolate the contribution of coal-related and other energy-sector emissions, a set of sector-specific sensitivity simulations was conducted in which emissions from individual power-sector sources were activated within China while all other emissions remained switched off. These simulations used identical emission inventories and model configurations for the global and nested domains to ensure consistency of sources across spatial scales (Supplementary Table 6).

To relate the GEOS-Chem attribution to PV energy losses, monthly baseline AOD and energy-sector AOD were computed for all MERRA-2 grid cells in China for 2023. The monthly fraction of total AOD attributable to energy-sector emissions was then aggregated across all PV facilities using a weighted mean, where the weight for each facility equals its aerosol-induced energy reduction for that month. This weighting accounts for site-specific capacity, monthly irradiance and effective AOD, providing a direct link between sector-attributed aerosol loading and PV energy losses.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.