4 LandR Biomass_speciesParameters Module

module-version-Badge

Issues-badge

4.0.0.1 Authors:

Ian Eddy [aut, cre], Eliot McIntire [aut], Ceres Barros [ctb]

This documentation is work in progress. Potential discrepancies and omissions may exist for the time being. If you find any, contact us using the “Get help” link above.

4.1 Module Overview

4.1.2 Summary

LandR Biomass_speciesParameters (hereafter Biomass_speciesParameters) calibrates species growth and mortality trait values used in Biomass_core, by matching theoretical species’ growth curves obtained with different trait values (see Simulated species data) against observed growth curves derived from Permanent Sample Plots (PSP data) across Canada (see Permanent sample plot data), to find the combination of trait values that allows a better match to the observed curves. In particular, it calibrates the growthcurve, mortalityshape, maximum biomass (maxB) and maximum aboveground net primary productivity (maxANPP) traits (see Parameter estimation/calibration), the latter two in conjunction with the module Biomass_borealDataPrep.

This module will not obtain other traits or parameters used in Biomass_core and so it is meant to be used in conjunction with another data/calibration module that does so (e.g., Biomass_borealDataPrep). However it can be used stand-alone in an initial developmental phase for easier inspection of the statistical calibration procedure employed.

As of September 21, 2026, the raw PSP data used in this module is not freely available, and data sharing agreements must be obtained from the governments of SK, AB, and BC to obtain it. However, the processed and anonymized PSP data is provided via a Google Drive folder accessed automatically by the module.

*Google Account is therefore necessary to access the data used for calibration.**

4.2 Module manual

4.2.1 General functioning

Tree cohort growth and mortality in Biomass_core are essentially determined by five parameters: growthcurve, mortalityshape, maximum biomass (maxB), maximum aboveground net primary productivity (maxANPP) and longevity.

The growthcurve and mortalityshape parameters (called ‘growth curve’ and ‘mortality shape’ in LANDIS-II Biomass Succession Extension v3.2, the base model for Biomass_core) strongly modulate the shape of species growth curves and so it is important that they are calibrated to the study area in question.

Also, the growth and mortality equations used in Biomass_core are non-linear and their resulting actual biomass accumulation curve is an emergent phenomenon due to competition effects. This means that the ideal trait/parameter values should not be estimated on pure single species growth conditions, as their resulting dynamics will be different in a multi-species context.

Biomass_speciesParameters attempts to address these issues (at least partially) using a “curve-matching” approach. It compares the best fit (according to their AIC) of three non-linear forms (Chapman-Richard’s, Gompertz, and a logistic form) fitted to permanent sample plot (PSP) data to a large collection of theoretical (i.e. simulated) species curves, each representing a different set of the five key parameters that govern biomass increment in Biomass_core: growthcurve, mortalityshape, the ratio of maxANPP to maxB, and longevity. This library of curves is produced by the Biomass_speciesFactorial module.

Biomass_speciesParameters generally follows other LandR data modules, like Biomass_boreaDataPrep, which also attempts to calibrate previously estimated spatially varying species traits such as maxB and maxANPP from the input data layers.

4.2.1.1 Permanent sample plot data

Biomass_speciesParameters can use all the PSP data available (note that it may span several thousands of kilometres), or select the data based on a shapefile (studyAreaANPP; see List of input objects).

By default, all available PSP is obtained, via the PSPdataTypes parameter. The particular data sets will depend on the version of ianmseddy/PSPclean that is installed (assuming requisite file-sharing permissions are available). This may include data from the provinces of BC, AB, SK, ON, QC, and NB, as well as the National Forest Inventory. If no file-sharing agreements are in place, the dummy option can be used. This will obtain data from BC, AB, SK, and the NFI that were previously treated for errors and standardized into a single data set with the exact location and identifying attributes anonymized. However, it should be noted that data for BC, Quebec, New Brunswick, and the NFI are freely available, thus no file-sharing agreement is necessary for these jurisdictions.

The data include individual species, diameter at breast height (DBH), and sometimes tree height measurements for each tree in a plot, as well as stand age. As part of the standardization process, dead trees were removed from the data set. Tree biomass was then per estimated species using either a DBH-only model or a DBH-height model from Lambert et al. (2005), in \(g/m^2\).

Note that the model used to calculate biomass can also be changed to Ung et al. (2008) via the P(sim)$biomassModel module parameter (see list of parameters).

4.2.1.2 Simulated species data

The Biomass_speciesFactorial module was used to create a library of theoretical species curves (biomass accumulation curves, to be more precise) to which the best non-linear model form fit to the PSP-biomass will be matched for each species and species combinations in the study area landscape. The library of curves was created by running several Biomass_core simulations with no reproduction, competition, disturbance, or dispersal effects. The species in the simulations encompassed the full factorial of traits, growing alone and in competition with one cohort of each other combination. Each simulation differed in the combination of species trait values that influence growth and mortality dynamics, namely: growthcurve, mortalityshape, longevity, maxANPP and maximum biomass (maxBiomass, not to be confused with the data-driven maxB which is later calibrated).

The values for maxANPP were explored via the mANPPproportion, the ratio of maxANPP to maxBiomass (the parameter used for theoretical curves), as it reflects their relationship. While any factorial combination is possible to run using the module Biomass_speciesFactorial, the default objects used by Biomass_speciesParameters utilized growthcurve values ranging from 0.65 to 0.85 in increments of 0.02, mortalityshape values of 20, 22, and 24, mANPPproportion values from 3.5 to 6.5 in increments of 0.5, and longevity values of 125 to 700 in increments of 25, all of which total over 6.5 million unique combinations of traits.

Results from these simulations were compiled into a table (cohortDataFactorial ; see List of input objects) that is accessed by Biomass_speciesParameters, so that the module can be run without needing to re-simulate the theoretical curves.

4.2.1.3 Parameter estimation/calibration

Biomass_speciesParameters calibrates growthcurve, mortalityshape and mANPPproportion by matching the theoretical species curves produced by Biomass_speciesFactorial (cohortDataFactorial) against observed species growth curves from the PSP data.

The parameter P(sim)$speciesFittingApproach determines which of four possible fitting approaches to use: all, single, pair-wise, or focal. The all approach combines all PSPs into a single species, thus parameterizing all species identically. This is not intended to have an ecological application. The "single" approach is the earliest method, where each species utilizes plots where 50% or more of the biomass is composed solely of that species. Non-linear models for each species are fit to these observations independently of each other, with each species possessing its own subset of plots. However, this approach is unable to accurately characterize the competitive effects that arise when multiple species occupy the same plot. Therefore it is suitable only for characterizing species that grow in pure stands, and has been retained largely for backwards compatibility.

The "pairwise" approach retains all plots where exactly two species each represent more than 20% of the total plot biomass, e.g. a plot with 40% Pinus contorta and 39% Populus tremuloides is retained only if the remaining 21% of biomass is composed of more than one species. Then, for each combination of 2 species-of-interest, separate non-linear models are fit for each species. #TODO: I believe the models are ultimately combined but this might not happen until traits are selected. Of the three approaches, the pairwise is best able to account for competition, but is most sensitive to data availability and can be confounded by combinations of species that seldom occur together.

The third approach speciesFittingApproach, "focal", is the default. This approach is similar to "pairwise", but combines the species that are not the species-of-interest (or ‘focal’ species). For example, in a plot with biomass composition of 20% Pinus contorta, 40% Populus tremuloides, and 21% Picea mariana, when fitting the P. contorta equations, the biomass of P. tremuloides and P. mariana are combined, and when fitting P. tremuloides, the biomass of P. contorta and P. mariana are combined. In comparison to the pairwise approach, this approach sacrifices some detail but is less sensitive to data quality and availability.

Before calculating the observed species growth curves (i.e., the best of three non-linear forms to match PSP data), the module subsets the PSP data to stand ages below the 95th percent quantile for all species (this can be changed via the P(sim)$quantileAgeSubset module parameter), as records for larger age classes were limited and constituted statistical outliers. In some species, changing the quantile value may improve results, however. Two examples are Pinus banksiana and Populus sp (in western Canada), for which using the 99th percent quantile improved the models, because these are short-lived species for which data at advanced ages is scarce. In addition, weights are added at the origin (age = 0 and biomass = 0) to force the intercept to be essentially at 0 age and 0 biomass.

The best fit of three non-linear forms, for each focal species, is then calculated. Focal species are defined as either 50% of dominance in the plot, or 20% if we are looking to capture the multi-species dynamics (currently the default). Three growth model forms are then fit to the observations for the focal species: a Chapman-Richard’s form [Equation (4.1); see, e.g., Coble & Lee (2006)], a Gompertz form (Equation (4.2)) and a Logistic form [Equation (4.3); see Fekedulegn et al. (1999) for a complete overview of these equations]. Multiple tries using the estimation methods from the robustbase::nlrob function for each form are used, and the best model fit is selected via Akaike Information Criterion (AIC).

\[\begin{equation} B \sim A \times (1 - e^{-k \times age})^{p} \tag{4.1} \end{equation}\] \[\begin{equation} B \sim A \times e^{-k \times e^{-p \times age}} \tag{4.2} \end{equation}\] \[\begin{equation} B \sim \frac{A}{1 + k \times e^{-p \times age}} \tag{4.3} \end{equation}\]

Species biomass (\(B\)) is estimated as a function of stand age (\(age\)), with the best values of the \(A\), \(k\) and \(p\) parameters to fit the PSP data.

It is possible that some selected species do not have enough data to allow for model convergence. In this case, Biomass_speciesParameters skips parameter calibration. Consequently the module will interpolate the mANPPproportion, growthcurve, and mortalityshape from the respective means of other species. These species are tracked via the source column in sim$species, which will be one of "interpolated" or "estimated". This mechanism ensures these species remain competitive, as the parameterized traits are often significant departures from defaults used in many LANDIS-II applications to Canada’s boreal forests.

After each species best fit is selected (using AIC), Biomass_speciesParameters compares it to the library of theoretical curves, and picks the best one based on maximum likelihood. This best theoretical curve will be associated with a given combination of growthcurve, mortalityshape and maxANPPproportion values, which are then used directly as the calibrated values, in case of growthcurve and mortalityshape, or to calibrate maxANPP in the case of maxANPPproportion (see below).

Because simulated growth curves never achieve the maximum biomass parameter (the maxBiomass parameter set to 5000 for all simulations of theoretical species curves, or the maxB parameter in Biomass_core simulations), it acts as an asymptotic limit that reflects the potential maximum biomass for a species in an ecolocation (ecological zone and land cover combination).

Biomass_speciesParameters uses the ratio between the potential maximum biomass (maxBiomass, always 5000) to the achieved maximum biomass in the theoretical curves, to rescale maxB. This ratio is called the inflationFactor and it is multiplied by maxB values previously estimated from data (e.g. by Biomass_borealDataPrep). This way, species simulated in Biomass_core are able to achieve the maximum observed biomass used to initially estimate maxB.

Finally, the module calibrates maxANPP using the mANPPproportion value from the best matching theoretical growth curve as:

\[\begin{equation} maxB \times \frac{mANPPproportion}{100} \tag{4.4} \end{equation}\]

where maxB is the already (re-)calibrated version.

4.2.2 List of input objects

The full list of input objects required by the module is presented below (Table 4.1). The input studyAreaANPP (the study area used extract the PSP data from) is optional, and therefore no default is supplied. All other input objects have internal defaults, but the user may need to request access to their online files.

Of these inputs, the following are particularly important and deserve special attention:

Spatial layers

-   `studyAreaANPP` -- shapefile. An `sf` or `SpatVector` object 
determining the geographic extent of the PSP data. 
  • Tables

    • speciesTableFactorial and cohortDataFactorial – a tables of species trait combinations and the theoretical species growth curve data (respectively)

    • PSPmeasure_sppParams, PSPplot_sppParams and PSPgis_sppParams – tree measurement, biomass growth and geographical data of the PSP datasets used to build observed species growth curves.

    • species – a table of invariant species traits that may have been produced by another module. It must contain the columns ‘species’, ‘growthcurve’ and ‘mortality shape’, whose values will be calibrated.

    • speciesEcoregion – table of spatially-varying species traits that may have been produced by another module. It must contain the columns ‘speciesCode’, ‘maxB’ and ‘maxANPP’ and ‘ecoregionGroup’ (the ecolocation ID). ‘maxB’ and ‘maxANPP’ values will be calibrated by species.

Table 4.1: List of Biomass_speciesParameters input objects and their description.
objectName objectClass desc sourceURL
cohortDataFactorial_path fs_path Path where the cohortDataFactorial object is saved as an arrow dataset. A large cohortData table ( sensu Biomass_core) with columns age, B, and speciesCode that joins with speciesTableFactorial. See PredictiveEcology/Biomass_factorial for further information. https://drive.google.com/file/d/1NH7OpAnWtLyO8JVnhwdMJakOyapBnuBH/
PSPmeasure_sppParams data.table Merged PSP and TSP individual tree measurements. Must include the following columns: MeasureID, OrigPlotID1, MeasureYear, TreeNumber, Species, DBH and PSP, where Species corresponds to species names in LandR::sppEquivalencies_CA$Latin_full. Defaults to randomized PSP data stripped of real plotIDs https://drive.google.com/file/d/1LmOaEtCZ6EBeIlAm6ttfLqBqQnQu4Ca7/view?usp=sharing
PSPplot_sppParams data.table Merged PSP and TSP plot data. Defaults to randomized PSP data stripped of real plotIDs. Must contain columns MeasureID, MeasureYear, OrigPlotID1, and baseSA, the latter being stand age at year of first measurement https://drive.google.com/file/d/1LmOaEtCZ6EBeIlAm6ttfLqBqQnQu4Ca7/view?usp=sharing
PSPgis_sppParams sf Plot location sf object. Defaults to PSP data stripped of real plotIDs/location. Must include field OrigPlotID1 for joining to PSPplot object https://drive.google.com/file/d/1LmOaEtCZ6EBeIlAm6ttfLqBqQnQu4Ca7/view?usp=sharing
species data.table A table of invariant species traits with the following trait colums: ‘species’, ‘Area’, ‘longevity’, ‘sexualmature’, ‘shadetolerance’, ‘firetolerance’, ‘seeddistance_eff’, ‘seeddistance_max’, ‘resproutprob’, ‘mortalityshape’, ‘growthcurve’, ‘resproutage_min’, ‘resproutage_max’, ‘postfireregen’, ‘wooddecayrate’, ‘leaflongevity’ ‘leafLignin’, and ‘hardsoft’. Only ‘growthcurve’, ‘hardsoft’, and ‘mortalityshape’ are used in this module. Default is from Dominic Cyr and Yan Boulanger’s applications of LANDIS-II https://raw.githubusercontent.com/dcyr/LANDIS-II_IA_generalUseFiles/master/speciesTraits.csv
speciesEcoregion data.table Table of spatially-varying species traits (maxB, maxANPP, establishprob), defined by species and ecoregionGroup). Defaults to a dummy table based on dummy data of biomass, age, ecoregion and land cover class. NA
speciesTableFactorial_path fs_path Path where the speciesTableFactorial object is saved as an arrow dataset. A large species table ( sensu Biomass_core) with all columns used by Biomass_core, e.g., longevity, growthcurve, mortalityshape, etc., when it was used to generate cohortDataFactorial. See PredictiveEcology/Biomass_factorial for futher information. https://drive.google.com/file/d/1NH7OpAnWtLyO8JVnhwdMJakOyapBnuBH/
sppEquiv data.table Table of species equivalencies - see ?LandR::sppEquivalencies_CA. Traits will be estimated for each unique entry in the sppEquivCol column. NA
sppEquivLong data.table The full table of species equivalencies - see ?LandR::sppEquivalencies_CA. Biomass will be estimated for each species based on the sp_Biomass_eq' column, which usespemisc::biomassCalculation` to derive AGB from DBH and height (based on the equations from https://doi.org/10.1139/x05-112). The full table is used to improve stand biomass estimates even if some species are not of interest. NA
studyAreaANPP sf Optional study area used to crop PSP data before building growth curves. If supplied, an ecoregion-scale object is recommended, at a minimum. NA

4.2.3 List of parameters

The full list of parameters used by the module is presented below (Table 4.2), all of which have default values specified in the module’s metadata.

Of these parameters, the following are particularly important:

Calibration parameters

-   `standAgesForFitting` -- determines the range of ages for which the fit of the growth curves is evaluated against the non-linear model. It should be a subset of the full growth curve, because there is extremely limited plot data for stands aged 0-20, as well as towards the end of a species' longevity.

-   `speciesFittingApproach` -- should the calibration take into account 
species growing in single- or multi-species context?

-   `quantileAgeSubset` -- upper quantile age value used to subset PSP data.
It can be a vector named by species, or a single numeric. It is used to limit the outsized influence of old stands at the tail end of the stand age 
distribution. However, it will have varying effects by species and PSP data.

Data processing

-   `PSPdataTypes` -- which jurisdictional plot data to use, important 
because the default will attempt to use PSP data that may be inaccessible

-   `sppEquivCol` -- the column name in `sim$sppEquiv` that ultimately determines which species are modeled, and whether any are combined. It should be complete for every row, and duplicates are merged together. For example, a table with `"Picea englemannii"` and `"Picea glauca"` in the "Latin" column, and "White spruce" in the "EN_generic_short" column would result in the two rows merged into one single growth curve if "EN_generic_short" were the sppEquivCol, but have two growth curves if `Latin` were the sppEquivCol. The estimate biomass estimation is always dependent upon the column "PSP", and missing entries will lead to the species' omission. 
Table 4.2: List of Biomass_speciesParameters parameters and their description.
paramName paramClass default min max paramDesc
biomassModel character Lambert2005 NA NA The model used to calculate biomass from DBH. Can be either ‘Lambert2005’ or ‘Ung2008’.
landis logical FALSE NA NA If TRUE, run in ‘LANDIS mode’: expose the fitted LANDIS-version growth curves (the scaled non-linear biomass-over-age curves shown in the LandR-vs-non-linear plot) per species as the output object speciesGrowthCurvesLandis, for use as inputs to LANDIS-II Biomass Succession. The per-species growth-curve parameters (growthcurve, mortalityshape, etc.) are written to species regardless. Default FALSE preserves the standard behaviour.
maxBInFactorial integer 5000 NA NA The arbitrary maximum biomass for the factorial simulations. This is a per-species maximum within a pixel
minimumPlots numeric 50 10 NA Minimum number of PSP plots per species
minDBH integer 0 0 NA Minimum diameter at breast height (DBH) in cm used to filter PSP data. Defaults to 0 cm, i.e. all tree measurements are used.
PSPdataTypes character all NA NA Which PSP datasets to source, defaulting to all. Other available options include ‘BC’, ‘AB’, ‘SK’, ‘ON’, ‘NB’, ‘NFI’, and ‘dummy’. ‘dummy’ should be used for unauthorized users.
PSPperiod numeric 1920, 2019 NA NA The years by which to subset sample plot data, if desired. Must be a vector of length 2
quantileAgeSubset numeric 99 1 100 Quantile by which to subset PSP data. As older stands are sparsely represented the oldest measurements become vastly more influential. This parameter accepts both a single value and a list of vectors, named according to sppEquivCol.
speciesFittingApproach character focal NA NA Either ‘all’, ‘pairwise’, ‘focal’ or ‘single’, indicating whether to pool all species into one fit, do pairwise species (for multiple cohort situations) do pairwise species, but using a focal species approach where all other species are pooled into ‘other’ or do one species at a time. If ‘all’, all species will have identical species-level traits.
sppEquivCol character LandR NA NA The column in sim$sppEquiv data.table that defines individual species. The names should match those in the species table.
standAgesForFitting integer 21, 91 NA NA The minimum and maximum ages of the biomass-by-age curves used in fitting. It is generally recommended to keep this param under 200, given the low data availability of stands aged 200+, with some exceptions. For a closed interval, end with a 1, e.g. c(31, 101).
useHeight logical TRUE NA NA Should height be used to calculate biomass (in addition to DBH). DBH is used by itself when height is missing.
.plots character screen NA NA Used by Plots function, which can be optionally used here
.plotInitialTime numeric 0 NA NA This describes the simulation time at which the first plot event should occur
.plotInterval numeric NA NA NA This describes the simulation time interval between plot events
.saveInitialTime numeric NA NA NA This describes the simulation time at which the first save event should occur
.saveInterval numeric NA NA NA This describes the simulation time interval between save events
.studyAreaName character NA NA NA Human-readable name for the growth curve filename. If NA, a hash of sppEquiv[[sppEquivCol]] will be used.
.useCache character .inputOb…. NA NA Should this entire module be run with caching activated? This is generally intended for data-type modules, where stochasticity and time are not relevant
.useParallel integer 2 NA NA maximum number of threads/workers to use for data.table operations; passed to data.table::setDTthreads and should be <= 4.

4.2.4 List of outputs

The module produces the following outputs (Table 4.3). Note that species and speciesEcoregion are modified versions of the inputed objects with the same name.

Tables

-   `species` and `speciesEcoregion` -- tables with calibrated trait values.

-   `speciesGAMMs` -- the fitted GAMM model objects for each species.
Table 4.3: List of Biomass_speciesParameters output objects and their description.
objectName objectClass desc
species data.table The updated invariant species traits table (see above).
speciesEcoregion data.table The updated spatially-varying species traits table (see description for this object in inputs)
speciesGrowthCurves list list containing each species’ non-linear model, model data, and the unfiltered PSP data
speciesGrowthCurvesLandis data.table An empty data.table unless P(sim)$landis is TRUE, when it holds the fitted LANDIS-version growth curves (BscaledNonLinear) by species and standAge (the scaled non-linear curves shown in the LandR_VS_NLM_growthCurves plot), for use as LANDIS-II Biomass Succession inputs.
speciesGrowthCurvesPSP data.table An empty data.table unless P(sim)$landis is TRUE, when it holds the PSP observations used to fit the growth curves (biomass by standAge and species, with OrigPlotID1 for joining to plot locations / ecoregion), a diagnostic to plot against speciesGrowthCurvesLandis.

4.2.5 Simulation flow and module events

Biomass_speciesParameters initializes itself and prepares all inputs provided there is an active internet connection and the user has access to the data (and a Google Account to do so).

We advise future users to run Biomass_speciesParameters with defaults and inspect what the objects are like before supplying their own data. The user does not need to run Biomass_speciesFactorial to generate their own theoretical curves (unless they wish to), as the module accesses pre-generated theoretical curves.

Note that this module only runs once (in one “time step”) and only executes one event (init). The general flow of Biomass_speciesParameters processes is:

  1. Preparation of all necessary data and input objects that do not require parameter fitting (e.g., the theoretical species growth curve data);

  2. Sub-setting PSP data and calculating the observed species growth curves using non-linear growth models;

  3. Finding the theoretical species growth curve that best matches the observed curve, for each species. Theoretical curves are subset to those with longevity matching the species’ longevity (in species table) and with growthcurve and mortalityshape values;

  4. Calibrating maxB and maxANPP.

4.3 Usage example

This module can be run stand-alone, but it won’t do much more than calibrate species trait values based on dummy input trait values. We provide an example of this below, since it may be of value to run the module by itself to become acquainted with the calibration process and explore the fitted non-linear models. However, we remind that to run this example you will need a Google Account, and to be granted access to the data.

A realistic usage example of this module and a few others can be found in this repository and in Barros et al. (2023a).

4.3.1 Load SpaDES and other packages.

4.3.2 Set up R libraries

tempDir <- tempdir()

pkgPath <- file.path(tempDir, "packages", version$platform, paste0(version$major,
    ".", strsplit(version$minor, "[.]")[[1]][1]))
dir.create(pkgPath, recursive = TRUE)
.libPaths(pkgPath, include.site = FALSE)

repos <- c("predictiveecology.r-universe.dev", getOption("repos"))
options(repos = repos)
install.packages("SpaDES.project")  ## gets Require too

4.3.3 Get the module and module dependencies

library(Require)

paths <- list(inputPath = normPath(file.path(tempDir, "inputs")),
    cachePath = normPath(file.path(tempDir, "cache")), modulePath = normPath(file.path(tempDir,
        "modules")), outputPath = normPath(file.path(tempDir,
        "outputs")))

SpaDES.project::getModule(modulePath = paths$modulePath, c("PredictiveEcology/Biomass_speciesParameters@main"),
    overwrite = TRUE)
Require::Require("SpaDES.core (>= 2.1.4)")
Require("googledrive")

4.3.4 Setup simulation

times <- list(start = 0, end = 1)

modules <- list("Biomass_speciesParameters")

objects <- list()
inputs <- list()
outputs <- list()
parameters <- list()
mySim <- SpaDES.core::simInitAndSpades(times = times, params = parameters,
    modules = modules, paths = paths, objects = objects)

## to inspect the fitted growth models:
mySim$speciesGrowthCurves$Pice_mar

4.4 References