PEcAn × ILAMB

GSoC 2026

2026Google Summer of Code, PEcAn Project

Benchmarking PEcAn’s Carbon Reanalysis Against TRENDY and CMIP Using ILAMB

A 100-member North American carbon reanalysis benchmarked against established model ensembles and tested to determine whether its ensemble uncertainty is calibrated.

Sites
~8,000 SDA
Ensemble
100 members
Grid
1 km
Period
2012–2024
Variables
Biomass, soil carbon, LAI
Begin

01The first question

The straightforward comparison came first.

PEcAn’s North American reanalysis assimilates satellite and ground-based observations into the SIPNET process model, producing 1 km, 100-member estimates of carbon and water pools across the continent. But before this project, it had never been benchmarked head-to-head against established model ensembles using ILAMB.

So the work began with that comparison, end to end. Convert the reanalysis into ILAMB-compatible fields, regrid CMIP6 and TRENDY onto the same grid, then score every model and all 100 PEcAn members individually against the same observational benchmarks.

The three benchmarked quantities span different parts of the terrestrial carbon system. Leaf area index describes the canopy, aboveground biomass captures carbon stored in vegetation, and soil carbon captures carbon stored belowground. Every ratio quoted later in this page belongs to one of these three.

Framework
ILAMB, from Collier et al. 2018
Reference ensembles
CMIP6 historical and historical plus ssp245, alongside TRENDY from the Global Carbon Budget 2024
Evaluation windows
2012 to 2014, and 2015 to 2023
Scored quantities
Aboveground biomass, soil carbon, leaf area index

02Scoring the mean

PEcAn’s members score well, and they score alike.

  1. 1.PEcAn scores above TRENDY on all three variables.
  2. 2.It scores above both ensembles on leaf area index, in every window.
  3. 3.It is comparable to CMIP6 on biomass and on soil carbon.
  4. 4.Its 100 members vary far less in skill than the models inside either ensemble.
Three panels of ILAMB skill scores for biomass, leaf area index and soil carbon. In each panel PEcAn's 100 members form a tight blue cluster at a high score, while individual CMIP6 models in red and TRENDY models in green scatter across a much wider and generally lower range.
FIG. 01
  • 100 PEcAn members
  • CMIP6 models
  • TRENDY models
  • Ensemble-mean field

ILAMB score per variable, higher is better: The left group covers 2012 to 2014 and the right group covers 2015 to 2023, and the short horizontal bar marks each group’s median. This measures how well each member scores, not how wide its uncertainty is.

Explore the live ILAMB scorecard The full scorecard, with bias, RMSE, seasonal cycle and spatial distribution scored for every variable and every model.

03Where the project changed

The benchmark showed strong deterministic skill. The next question was whether the ensemble spread actually represented its error.

The standard ILAMB scorecard evaluates deterministic model performance. It does not test whether PEcAn’s between-member spread is calibrated.

04Agreement is not calibration

Members can agree tightly and still fail to contain the observation.

What the diagram shows: Each mark is one of the 100 data-assimilation members at a single site. They begin apart and settle into a narrow cluster. The benchmark value lies outside the ensemble’s narrow central range.

Diagram. One hundred ensemble members converge into a narrow cluster while the observation remains far outside the ensemble’s central 90% range, so the ensemble spread is much smaller than its error.

  1. 01One hundred dispersed members
  2. 02They settle into a narrow range
  3. 03The observation sits outside it
  4. 04The error is much larger than the spread

Schematic: The spread is the width of the ensemble, and the error is the distance from the ensemble to the observation. Calibration asks whether ensemble spread is commensurate with realized error across many observations.

05Spread divided by error

A calibrated ensemble scores 1.00, and none of the three variables comes close.

For this diagnostic, a spread-to-error ratio near 1.00 indicates that ensemble spread is appropriately sized relative to realized error. Values far below 1.00 indicate underdispersion. Measured across the whole domain against the benchmarks used here, PEcAn’s three ratios are 0.07 for biomass, 0.15 for soil carbon and 0.20 for leaf area index.

  1. 0.07
  2. 0.15
  3. 0.20
1.00 calibrated
  1. Biomassagainst XuSaatchi and ESA-CCI
  2. Soil carbonagainst the reference soil product
  3. LAIagainst the reference LAI product

Read the gap to 1.0, not just the bar height: It is the distance each variable falls short of a calibrated ensemble. These are bias-removed spread-to-error ratios measured across the whole domain. The second diagnostic, 90% coverage, agrees with them, because far fewer than the expected 90% of observations fall inside PEcAn’s central 90% interval for every variable.

06The measurement

Both calibration diagnostics, against both reference ensembles.

The same diagnostics applied to CMIP6 and TRENDY place those ensembles substantially closer to their reference values than PEcAn. These represent different kinds of uncertainty. PEcAn is a within-model data-assimilation ensemble, while CMIP6 and TRENDY represent structural disagreement across different models. The comparison here concerns uncertainty calibration, not deterministic accuracy.

Two bar panels. Panel a, ensemble spread divided by error: PEcAn's bars are near zero for biomass, soil carbon and LAI while CMIP6 and TRENDY sit close to or above the well-calibrated line of 1.0. Panel b, fraction of observations inside the central 90% range: PEcAn is close to zero while CMIP6 and TRENDY are between 0.3 and 0.86 against the expected 0.90.
FIG. 02

Two views of ensemble calibration: Panel (a) is the bias-removed spread-to-error ratio against the well-calibrated value of 1.0, and panel (b) is the fraction of observations inside each ensemble’s central 90% range against the expected 0.90. The dashed lines are the reference values in both panels.

07Robustness

Changing the benchmark and accounting for observational uncertainty leaves the result largely unchanged.

A calibration result is only as informative as the reference it is measured against, so biomass was re-scored against a second independent product, then re-scored again with the benchmark’s own uncertainty folded in.

Two bar panels comparing biomass calibration against XuSaatchi, ESA-CCI 2020 and ESA-CCI 2024. Spread-to-error ratios are near 0.07 for all three benchmarks and 90% coverage stays near or below 0.05, far from the calibrated and expected reference lines.
FIG. 03

Biomass against three benchmarks: XuSaatchi and ESA-CCI in both 2020 and 2024 give almost identical ratios, so the result is not specific to a single observational benchmark.

Two bar panels showing 90% coverage and spread-to-error ratio for biomass under three treatments: raw, plus ESA-CCI per-pixel standard deviation, and plus XuSaatchi to ESA-CCI disagreement. Coverage rises only from about 0.02 to about 0.10 and the ratio stays near 0.07.
FIG. 04

Accounting for benchmark uncertainty: Biomass coverage rises from roughly 0.02 to roughly 0.10, still far below the expected 0.90. The spread-to-error ratio changes only modestly, so benchmark uncertainty does not account for most of the observed underdispersion.

How the two biomass benchmarks compare with each other
Scatter plot of ESA-CCI against XuSaatchi carbon density in Mg C per hectare, with a one-to-one line. Points spread widely at low densities and tighten toward the one-to-one line at higher, densely forested values.
FIG. 05

The two products disagree with each other, most of all at low carbon density: That disagreement is itself one of the observation-error terms used in Fig. 04, which is why the robustness check is strongest for biomass. The soil-carbon and LAI reference datasets used here do not carry published per-pixel uncertainty layers, so the same test cannot be run on them.

08Where it happens

Not one corner of the continent.

Rather than report a single continental number, the same calibration engine was rerun on subsets of sites, first by MODIS land-cover class and then by EPA and CEC ecoregion. Every class and every tested Level 1 ecoregion stays below the calibrated ratio. The breakdown shows where the miscalibration is most pronounced.

Map of North America with roughly eight thousand state-data-assimilation sites plotted as coloured dots, one colour per MODIS plant functional type class, from evergreen needleleaf trees through shrubs, grass and croplands.
FIG. 06

The assimilation network: Roughly 8,000 SDA sites, coloured by land-cover class.

Map of North America with Level 1 ecoregions shaded by biomass spread-to-error ratio on a red to green scale. Almost every ecoregion is red, indicating ratios near zero rather than the well-sized value of one.
FIG. 07

Biomass ratio by ecoregion: Values near 1.0 indicate well-sized spread, while lower values indicate increasing underdispersion.

Two stacked bar panels by MODIS land-cover class. Top, spread-to-error ratios for biomass, soil carbon and LAI all fall between about 0.07 and 0.32, well under the well-calibrated line at 1.0. Bottom, 90 percent coverage, where biomass coverage is zero in the four forest classes.
FIG. 08

By land cover: Every land-cover class remains below the calibrated spread-to-error ratio. Biomass coverage is effectively zero across the four forest PFT classes.

09One omitted uncertainty source

The downscaler’s own error is larger than the disagreement among members.

Every ensemble member is carried to the 1 km grid by its own random forest. The mapped ensemble does not explicitly propagate the predictive error quantified by each random forest’s out-of-bag residuals. Comparing OOB RMSE with between-member spread at the same sites and in the same units therefore provides a diagnostic of an uncertainty source that may be missing from the reported spread.

Biomass

OOB RMSE is 2.76× the between-member spread
Between-member spread
5.85 Mg/ha
Downscaling out-of-bag RMSE
16.16 Mg/ha

Soil carbon

OOB RMSE is 1.42× the between-member spread
Between-member spread
6.37 Mg/ha
Downscaling out-of-bag RMSE
9.07 Mg/ha

Both variables are shown on the same scale: Green marks the between-member ensemble spread, while the lighter segment extends to the random forest out-of-bag RMSE. For biomass, OOB RMSE is 2.76× the ensemble spread. For soil carbon, it is 1.42×. Zhang et al. note that emulator training uncertainty is not propagated into the mapped uncertainty, making downscaling one plausible omitted uncertainty source consistent with the observed underdispersion. This diagnostic does not establish that it is the only source.

10Contributions

Seven pull requests to PEcAn.

The arc runs from raw format conversion through the full calibration analysis and its regional and mechanistic breakdowns. All seven are merged into PEcAn’s develop branch.