<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">SOIL</journal-id><journal-title-group>
    <journal-title>SOIL</journal-title>
    <abbrev-journal-title abbrev-type="publisher">SOIL</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">SOIL</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">2199-398X</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/soil-8-559-2022</article-id><title-group><article-title>How well does digital soil mapping represent soil geography?  An investigation from the USA</article-title><alt-title>Digital soil mapping and soil geography</alt-title>
      </title-group><?xmltex \runningtitle{Digital soil mapping and soil geography}?><?xmltex \runningauthor{D.~G.~Rossiter et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff2">
          <name><surname>Rossiter</surname><given-names>David G.</given-names></name>
          <email>david.rossiter@isric.org</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Poggio</surname><given-names>Laura</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-1892-0764</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Beaudette</surname><given-names>Dylan</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff4">
          <name><surname>Libohova</surname><given-names>Zamir</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>ISRIC – World Soil Information, Postbus 353, Wageningen 6700 AJ, the Netherlands</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Section of Soil &amp; Crop Sciences, New York State College of Agriculture and Life Sciences,<?xmltex \hack{\break}?> 233 Emerson Hall, Cornell University, Ithaca, NY 14853, USA</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>USDA – NRCS, Soil and Plant Science Division, 19777 Greenley Rd, Sonora, CA 95370, USA</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>USDA – ARS, Dale Bumpers Small Farms Research Center, 6883 South State Highway 23,<?xmltex \hack{\break}?> Booneville, AR 72927, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">David G. Rossiter (david.rossiter@isric.org)</corresp></author-notes><pub-date><day>5</day><month>September</month><year>2022</year></pub-date>
      
      <volume>8</volume>
      <issue>2</issue>
      <fpage>559</fpage><lpage>586</lpage>
      <history>
        <date date-type="received"><day>23</day><month>July</month><year>2021</year></date>
           <date date-type="rev-request"><day>13</day><month>September</month><year>2021</year></date>
           <date date-type="rev-recd"><day>1</day><month>June</month><year>2022</year></date>
           <date date-type="accepted"><day>8</day><month>August</month><year>2022</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2022 David G. Rossiter et al.</copyright-statement>
        <copyright-year>2022</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022.html">This article is available from https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022.html</self-uri><self-uri xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022.pdf">The full text article is available as a PDF file from https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e133">We present methods to evaluate the spatial patterns of the geographic distribution of soil properties in the USA, as shown in gridded maps produced by digital soil mapping (DSM) at global (SoilGrids v2), national (Soil Properties and Class 100 m Grids of the USA), and regional (POLARIS soil properties) scales and compare them to spatial patterns known from detailed field surveys (gNATSGO and gSSURGO). The methods are illustrated with an example, i.e. topsoil pH for an area in central New York state. A companion report examines other areas, soil properties, and depth intervals. A set of R Markdown scripts is referenced so that readers can apply the analysis for areas of their interest. For the test case, we discover and discuss substantial discrepancies between DSM  products and large differences between the DSM products and legacy field surveys. These differences are in whole-map statistics, visually identifiable landscape features, level of detail, range and strength of spatial autocorrelation, landscape metrics (Shannon diversity and evenness, shape, aggregation, mean fractal dimension, and co-occurrence vectors), and spatial patterns of property maps classified by histogram equalization. Histograms and variogram analysis revealed the smoothing effect of machine learning models. Property class maps made by histogram equalization were substantially different, but there was no consistent trend in their landscape metrics. The model using only national points and covariates was not substantially different from the global model and, in some cases, introduced artefacts from a lithology covariate. Uncertainty (5 %–95 % confidence intervals) provided by SoilGrids and POLARIS were unrealistically wide compared to gNATSGO/gSSURGO low and high estimated values and show substantially different spatial patterns. We discuss the potential use of the DSM products as a (partial) replacement for field-based soil surveys. There is no substitute for actually examining and interpreting the soil–landscape relation, but despite the issues revealed in this study, DSM can be an important aid to the soil surveyor.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e145">Digital soil mapping (DSM) has been defined (under the earlier term predictive soil mapping) as “the development of a numerical or statistical model of the relationship among environmental variables and soil properties, which is then applied to a geographic data base to create a predictive map”  <xref ref-type="bibr" rid="bib1.bibx55" id="paren.1"/>. Since the seminal paper of <xref ref-type="bibr" rid="bib1.bibx27" id="text.2"/>, recently reviewed by <xref ref-type="bibr" rid="bib1.bibx32" id="text.3"/>, DSM has been widely applied from the field to global levels. This is in contrast to what we here call the “traditional” soil survey, in which the soil surveyor develops a mental model of the soil geography <xref ref-type="bibr" rid="bib1.bibx20" id="paren.4"/> by interpreting the landscape with the aid of air photos, purposive transects, and detailed profile descriptions at locations thought to represent the central concepts of the soil classes present in the study area <xref ref-type="bibr" rid="bib1.bibx57" id="paren.5"/>.</p>
      <p id="d1e163">A principal attraction of DSM is that it produces consistent, geometrically correct, and reproducible gridded maps over large areas, given training data (“point” observations of soil classes, properties, or conditions), a set of environmental covariates covering the entire area to be mapped at some fixed grid resolution, and a set of algorithms implemented in computer code. This removes the need for expertise in discovering and interpreting the soil–landscape relations, also known as the paradigm of soil survey <xref ref-type="bibr" rid="bib1.bibx20" id="paren.6"/>, which is vital for traditional soil survey and difficult to acquire and harmonize among surveyors. However, expertise in soil–landscape relations is still needed to ensure that DSM outputs are reasonable and to discover reasons for any discrepancies.</p>
      <p id="d1e169">Furthermore, it may be that fewer locations can be visited in order to develop reliable models, as compared to traditional survey techniques. If the relation with covariates is strong, and locations representative of the entire covariate feature space are included in the training set, it may be possible to map large areas from relatively few field observations.
This corresponds to the “homosoil” concept <xref ref-type="bibr" rid="bib1.bibx26" id="paren.7"/>, where identical environmental conditions (as represented by covariates) should result in the same soils. Maps made by DSM can include areas that are not accessible to field mappers, because of permissions or difficult access, if the available training data cover the covariate space of the inaccessible area. However, DSM requires sufficient sampling density to cover the full covariate space, since most DSM methods do not interpolate or extrapolate in soil property space, and in any case, it is inadvisable to predict too far from the coverage of the training observations. This has been studied by <xref ref-type="bibr" rid="bib1.bibx30" id="text.8"/>, who have developed a method for measuring the distance in both the covariate <xref ref-type="bibr" rid="bib1.bibx30" id="paren.9"/> and geographic <xref ref-type="bibr" rid="bib1.bibx31" id="paren.10"/> space between prediction locations and the set of training points.</p>
      <p id="d1e184">DSM avoids some well-known problems of traditional survey, namely multiple survey projects over time with inconsistent standards and mapping concepts, inconsistency among mappers, difficulties in objectively identifying boundaries, and indeed the need to identify boundaries. However, traditional soil surveyors and users of their maps are often critical of DSM products and may not understand how they were made and how they should be used <xref ref-type="bibr" rid="bib1.bibx3" id="paren.11"/>. In the USA, there is an increasing awareness of, and interest in, DSM products. Here the most important point of contention has to do with DSM resolution (pixel size), which implies a mapping scale, compared to the scale at which differences can be reliably interpreted for user needs. Criticism of DSM products is proportional to the degree to which their implied spatial precision and accuracy is over-sold.</p>
      <p id="d1e191">Another benefit of DSM methods is the quantification of uncertainty inherent in various geostatistical and machine learning approaches <xref ref-type="bibr" rid="bib1.bibx58" id="paren.12"/>. In traditional mapping, uncertainty is implicitly encoded via the mapping scale (which determines the size of the minimum delineation), map unit purity specification (e.g. complex, association, and consociation), and taxonomic precision <xref ref-type="bibr" rid="bib1.bibx57" id="paren.13"><named-content content-type="pre">e.g. soil series vs. suborder;</named-content></xref>.</p>
      <p id="d1e202">The success of DSM in reproducing known point observations (i.e. pedons described in the field and characterized in the laboratory) is typically reported by evaluation (so-called “validation”) statistics based on data splitting or by cross-validation. These evaluations are almost never based on random sampling <xref ref-type="bibr" rid="bib1.bibx8" id="paren.14"/>, and since the source point datasets are almost always biased towards certain land uses, access constraints, or landscape locations, these evaluations carry forward these biases and must be interpreted with caution.</p>
      <p id="d1e208">A more serious issue is that point evaluations of DSM products do not consider the spatial pattern of predictions. By contrast, traditional soil surveys produce polygon maps of relatively homogeneous soil bodies (represented as soil map units), with the boundary lines placed at inflection points of maximum change between them <xref ref-type="bibr" rid="bib1.bibx23" id="paren.15"/>. These maps explicitly show the surveyor's interpretation of the soil landscape as developed from a mental model of the soil-forming processes and which, when viewed as a whole, shows the pattern of the soil cover. It has long been recognized that the soil cover forms patterns at various scales <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx19" id="paren.16"/> so that the traditional soil mapper attempts to find those patterns expressed at the map design scale. Since DSM predictions are on a grid cell basis, most DSM models neither have a concept of the relatively homogeneous natural soil bodies nor of the inflection points between them. However, it might be expected that, if the values of the DSM covariates representing the soil-forming factors also cluster in a similar pattern to the soil cover, then the DSM predictions would also cluster and approximate map units from traditional survey. Convolutional neural networks <xref ref-type="bibr" rid="bib1.bibx59" id="paren.17"><named-content content-type="pre">e.g.</named-content></xref>, not represented in the methods compared in this paper, explicitly consider neighbourhoods of various size but not explicitly connectivity. The question is thus to what degree DSM products represent the actual soil landscape spatial pattern and, more importantly, the underlying pedogenetic and geomorphic processes.</p>
      <p id="d1e222">DSM maps are most commonly produced at grid cell resolutions from 1 <inline-formula><mml:math id="M1" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> to 30 <inline-formula><mml:math id="M2" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> and even to <inline-formula><mml:math id="M3" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M4" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> for precision agriculture applications. Environmental covariates are available at these resolutions so that DSM products at high resolutions can show fine details that cannot be presented at the design scale of polygon maps made by traditional methods. These have minimum legible delineations (MLDs) of 0.25 <inline-formula><mml:math id="M5" display="inline"><mml:mrow class="unit"><mml:msup><mml:mi mathvariant="normal">cm</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx63" id="paren.18"/> or 0.40 <inline-formula><mml:math id="M6" display="inline"><mml:mrow class="unit"><mml:msup><mml:mi mathvariant="normal">cm</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>  <xref ref-type="bibr" rid="bib1.bibx13" id="paren.19"/> on the published map, which is multiplied by the scale factor. For example, a polygon map at <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">24</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula>, typical of USA traditional soil surveys, can represent spatial patterns of 1.44 <xref ref-type="bibr" rid="bib1.bibx63" id="paren.20"/> to 2.3 <inline-formula><mml:math id="M8" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ha</mml:mi></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx13" id="paren.21"/> minimum-sized polygons. The <xref ref-type="bibr" rid="bib1.bibx13" id="text.22"/> criteria have been incorporated into Natural Resources Conservation Service (NRCS) soil survey standards <xref ref-type="bibr" rid="bib1.bibx53 bib1.bibx57" id="paren.23"/>. These correspond to single grid cell resolutions of 240 to 384 <inline-formula><mml:math id="M9" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>, which are coarser than higher-resolution DSM products from (30 to 100 <inline-formula><mml:math id="M10" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>). But the question remains whether this implied fine detail represents the true differences or artefacts of the mapping process – in other words, should the DSM map unit trust the apparent differences between adjacent grid cells, or are some or most of these differences due to artefacts (noise) of the DSM process? Furthermore, there is the question of how well the medium-resolution products (e.g. 250 <inline-formula><mml:math id="M11" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>) represent the soil landscape at regional extent.</p>
      <p id="d1e348">The objective of this study is to present methods with which to evaluate the landscape and detailed level spatial patterns of DSM maps. These maps have been developed for global, national, or regional spatial extents. These patterns are compared with digital soil maps based on polygon maps produced by traditional soil survey, using field study and expert soil–landscape analysis. We chose the USA as a study area because of the availability of field-based soil surveys at <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">12</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">24</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> design scale, linked to detailed descriptions of modal soil profiles, and available as a seamless digital product. These comparisons may be useful in the context of current plans <xref ref-type="bibr" rid="bib1.bibx60" id="paren.24"/> for updating and completing the USA soil survey using DSM methods and GlobalSoilMap (GSM) specifications <xref ref-type="bibr" rid="bib1.bibx2" id="paren.25"/>. They should also be useful for developing realistic expectations for what DSM can and cannot deliver <xref ref-type="bibr" rid="bib1.bibx3" id="paren.26"/>.</p>
      <p id="d1e390">To evaluate  DSM methods we apply them to selected test areas and soil properties, and comment on the results.
This paper introduces the methods and data sources, and includes an illustrative example (one area, one soil property, one depth interval), in the context of the soil geography of the selected region.
A companion ISRIC (International Soil Reference and Information Centre) Report <xref ref-type="bibr" rid="bib1.bibx52" id="paren.27"/> presents four case studies in diverse soil geographic contexts, each with different soil properties and depth intervals. We encourage readers to apply the methods to their own study areas within the USA, and to their soil properties of interest, to evaluate the utility of the several DSM products. For this, we provide our analysis scripts as R Markdown documents <xref ref-type="bibr" rid="bib1.bibx49" id="paren.28"><named-content content-type="post">see the code availability section at the end of this paper</named-content></xref>.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Products compared</title>
      <p id="d1e409">The products compared in this study differ in their primary data source (soil maps and point observations), their geographic scope, the mapping methods used to make the digital product, their resolution, depths and coordinate reference systems, and how they assess and present uncertainty. We summarize these below; see the journal articles describing each source for details.</p>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>General character of the products</title>
      <p id="d1e419">The products are of the following three kinds: (1) digital products based on traditional soil survey without any statistical modelling, (2) DSM products based on traditional soil survey products and enhanced by statistical modelling using environmental covariates, and (3) DSM products based on statistical modelling using training points and environmental covariates. This latter is the most common DSM method worldwide, especially for areas without extensive traditional soil surveys.</p>
      <p id="d1e422">The first kind of product is represented by the reference products from the Natural Resources Conservation Service (NRCS) of the United States Department of Agriculture (USDA), based on extensive field survey, air photo interpretation, thematic maps, and expert evaluation of digital elevation mapping (DEM) derivatives. This is considered to be the most accurate information, despite the occasional presence of artefacts from the overall mapping programme, as explained later in this section. There are two closely related products from the NRCS.</p>
      <p id="d1e425">At the national level, the National Soil Geographic Database, gNATSGO <xref ref-type="bibr" rid="bib1.bibx44" id="paren.29"/>, is a composite of the Soil Survey Geographic Database (SSURGO; mostly <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">24</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> scale), State Soil Geographic Database (STATSGO2; <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">250</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> scale), and the detailed Raster Soil Surveys (RSS) database, according to the most detailed product available for all areas of the USA. It is aimed at users who require multi-state or CONUS (contiguous United States) extent mapping. For each state or equivalent political unit, the SSURGO or STATSGO2 polygon maps of soil map unit (SMU) produced by traditional survey have been rasterized to a grid, with each cell keyed to a SMU <xref ref-type="bibr" rid="bib1.bibx41" id="paren.30"/>. Grid cells link to the best available (i.e. greatest detail STATSGO-SSURGO-RSS) SMU. The digital products are delivered at 30 and 90 m resolutions for the 48 contiguous states and the federal District of Columbia of the USA (abbreviated CONUS).</p>
      <p id="d1e464">At the state level, gSSURGO is also available <xref ref-type="bibr" rid="bib1.bibx43" id="paren.31"/>. This has a higher resolution (both 10 and 30 m) to minimize the degradation of the original polygon delineations and is a direct gridding of SSURGO polygons. It neither uses STATSGO2 for infilling nor RSS if available.
It thus is a gridded version of the familiar SSURGO product that is used for local applications. gSSURGO is refreshed annually for those users who do not wish to mix STATSGO or the new raster soil surveys into their analysis.</p>
      <p id="d1e471">The gridded gNATSGO and gSSURGO maps are derived from the polygons of SSURGO, which is a representation of those delineated by the field surveyors on stereopairs or orthophotos and subsequently converted to vector digital format by manual digitization. Soil surveys conducted in the last 15 years were compiled using on-screen digitization in a geographic information system (GIS). At boundaries between survey areas, polygon lines at survey limits have been matched during digitizing <xref ref-type="bibr" rid="bib1.bibx12" id="paren.32"/>. These polygons are organized in soil map units (SMUs), with one or more components (soil taxonomic units, STUs) usually named for a soil series but more specific than the parent soil series concept. Taxa above the soil series (family or subgroup) are commonly used in soil surveys of national forestland or wilderness areas. Soil series are the lowest level of Soil Taxonomy <xref ref-type="bibr" rid="bib1.bibx56" id="paren.33"/> and are described in the official series descriptions (OSDs) as modal profiles with a set of ranges for the observed morphology and laboratory measurements. The component STU in a mapped SMU varies in the observed field properties from the OSD modal description but usually fits within a soil series range. The observed field properties of soil component units are utilized for developing a set of interpretations for SSURGO polygon map units. These polygons are available from the NRCS as vector GIS layers <xref ref-type="bibr" rid="bib1.bibx34" id="paren.34"/> and in a convenient format on a geographic background as SoilWeb  <xref ref-type="bibr" rid="bib1.bibx9" id="paren.35"/>.</p>
      <p id="d1e486">The SMUs of the source maps are mappable landscape elements at the survey design scale. These almost always have multiple component STUs, with reported estimated proportion and geomorphic arrangement within the SMU when possible.
However, the locations of the STU within the SMU are not mapped due to the design scale. The STUs are linked to database tables of representative or synthetic soil profiles, with field and laboratory measurements of multiple soil properties and interpretations for soil use. To obtain values for soil properties in a gNATSGO or gSSURGO grid cell, properties of the components of the corresponding SMU are combined by area-weighted averaging. To obtain values at coarser resolutions, weighted average properties of groups of grid cells are upscaled by averaging.</p>
      <p id="d1e489">There are inherent problems with this product. First, since traditional surveys were carried out over a long time period, series names and mapping concepts may differ between adjacent survey areas. Thus, SSURGO SMU delineations and linked tabular data represent a progressive data collection and correlation effort spanning nearly 100 years. Therefore, there exist many soil survey vintages, each a snapshot in time, tied to specific land use assumptions and technological limitations. Systematic, continuous updates to the entire SSURGO database have been made since 2013 and are ongoing. Second, the transfer from unrectified photos to topographic base and the edge matching between survey areas has not always been flawless, and in addition, polygons may have been incorrectly drawn on the original survey (Fig. S1 in the Supplement). Thus, we cannot take these primary polygon maps as a completely reliable georeference.</p>
      <p id="d1e492">The second kind of product is represented by POLARIS soil properties <xref ref-type="bibr" rid="bib1.bibx10" id="paren.36"><named-content content-type="post">hereafter PSP</named-content></xref>, which is the result of harmonizing diverse SSURGO and STATSGO2 polygon data with the DSMART algorithm <xref ref-type="bibr" rid="bib1.bibx45" id="paren.37"/> to produce a probabilistic raster soil class or component map (30 <inline-formula><mml:math id="M16" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> grid resolution) and then extract property information from gSSURGO or gNATSGO grid cells (representing polygons) aggregated by component name. Despite the source data, this is not an NRCS product and was developed independently of the NRCS.</p>
      <p id="d1e511">There are two products representing the third kind of product, i.e. one for the world and one for the continental USA only. This allows us to compare globally and nationally consistent products. The global product is SoilGrids v2.0 <xref ref-type="bibr" rid="bib1.bibx21 bib1.bibx48" id="paren.38"><named-content content-type="pre">hereafter SG2;</named-content></xref>, a further development of SoilGrids1km <xref ref-type="bibr" rid="bib1.bibx15" id="paren.39"/> and SoilGrids250m <xref ref-type="bibr" rid="bib1.bibx16" id="paren.40"/>. This uses a global point dataset and environmental covariates that cover the entire world (except the high Arctic and Antarctica) and global models. It does not use any information derived from SSURGO or STATSGO map units. Its training points are extracted from the freely shareable World Soil  Information Service (WoSIS) point dataset from ISRIC–World Soil Information <xref ref-type="bibr" rid="bib1.bibx4" id="paren.41"/>. These include all profiles in the  National Soil Survey Center (NSSC) Laboratory Characterization Database. The freely shareable WoSIS points are augmented by several datasets included in WoSIS that cannot be published externally due to restrictions by the original data providers to ISRIC but which can be used in mapping. In total, <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">240</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> profiles were used in model building.</p>
      <p id="d1e542">The continental product is the Soil Properties and Class 100 m Grids of the United States <xref ref-type="bibr" rid="bib1.bibx50" id="paren.42"><named-content content-type="pre">hereafter SPCG;</named-content></xref>, which followed the methodology of <xref ref-type="bibr" rid="bib1.bibx16" id="text.43"/>, with the addition of USA-specific covariates, notably parent material and drainage classes extracted from SSURGO or STATSGO2 map units, and only used the CONUS extent of environmental covariates in model building. SPCG is similar to SG2 in that it is primarily based on point observations, but it has a richer source of these than SG2, i.e. the NSSC Laboratory Characterization Database (<inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mn mathvariant="normal">34</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">183</mml:mn></mml:mrow></mml:math></inline-formula> pedons comprising <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:mn mathvariant="normal">213</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">499</mml:mn></mml:mrow></mml:math></inline-formula> horizons), the National Soil Information System (NASIS), and the Rapid Carbon Assessment (RaCA) dataset (<inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:mn mathvariant="normal">31</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">215</mml:mn></mml:mrow></mml:math></inline-formula> pedons); the latter applies only for organic C, total N, and bulk density. It also uses SSURGO map units to derive parent material (87) and drainage (4) classes as CONUS-specific covariates.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Mapping methods</title>
      <p id="d1e594">gNATSGO and gSSURGO are based on traditional soil surveys, mostly on unrectified air photo bases until the late 1990s. The many individual survey areas prior to this time have been partially homogenized during a process of digitization and recompilation onto topographic or orthophoto bases during the 1990s <xref ref-type="bibr" rid="bib1.bibx12" id="paren.44"/> and are provided as the polygon SSURGO map. In the early 2000s, for new surveys and updates, a transition was made to on-screen digitization over orthophotos. Field methods are described in successive editions of the Soil Survey Manual <xref ref-type="bibr" rid="bib1.bibx57" id="paren.45"/> and the field book for describing and sampling soils <xref ref-type="bibr" rid="bib1.bibx53" id="paren.46"/>. Mapping is based on conceptual models of soil–landscape relations developed in each survey area <xref ref-type="bibr" rid="bib1.bibx20" id="paren.47"/> and confirmed by purposive auger and full profile descriptions to characterize the map unit composition. Component concepts are refined with any available laboratory characterization data, with (limited) new laboratory characterization performed as needed. Thus, SSURGO provides a local model of soil–landscape relations, developed in each area from the most significant soil-forming factors relevant to that area. SSURGO is progressively updated by field inspection and correlation, as problems are identified by soil surveyors or map users. Since SSURGO has been compiled from diverse surveys over many years, in some areas there are artefacts of that survey process (Fig. S2).</p>
      <p id="d1e609">The three DSM products (SG2, PSP, and SPCG)  use a large number of gridded GIS coverages as environmental covariates in their predictive models. These represent soil-forming factors and include climate, ecology, geology, land use/cover, terrain, vegetation, and hydrography (Sect. S4). PSP also uses coarse-resolution (<inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M22" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>) estimates of U, Th, and K <inline-formula><mml:math id="M23" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>-ray decay products to represent the suspected variation in parent material kind and origin.</p>
      <p id="d1e637">PSP <xref ref-type="bibr" rid="bib1.bibx10" id="paren.48"/> uses the DSMART disaggregation algorithm <xref ref-type="bibr" rid="bib1.bibx45" id="paren.49"/> to predict the most probable component (STU), along with their probability of occurrence, at each 30 <inline-formula><mml:math id="M24" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> resolution grid cell, and from the modal soil properties of the component, a probability-weighted aggregation. Disaggregation is the process of examining a coarser-resolution gridded or smaller-scale polygon product, which is known to have multiple STU, and identify the locations at a finer grid resolution where these components would be found should the original survey have been made at larger scale. This depends on fine-scale covariates that, in theory, relate to the STU within an SMU. It attempts to deal with the problems caused by multiple surveys over time, inconsistencies among mappers, and poor georeference of SMU boundaries by sampling out of mapped SMU polygons  according to declared proportions of map unit components (STU) and using these as pseudo-observations to train DSM models of STU occurrence. PSP does not use any point observations; rather, it samples pseudo-points from gSSURGO or gNATSGO and uses these as training points for the DSMART disaggregation algorithm (see below). The model is trained in overlapping tiles, each containing some set of SSURGO primary surveys, and using covariates covering just the tile. Thus, each POLARIS tile is derived from a local model in two senses. PSP provides a fine-scale map equivalent to <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">3000</mml:mn></mml:mrow></mml:math></inline-formula> design scale, i.e. from 16 to 64 times finer resolution than the original <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">12</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">24</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> surveys included in SSURGO. An obvious question is whether it is possible to map at this resolution from the SSURGO source, even with the fine-resolution covariates used by DSMART, because of the probabilistic nature of selecting pseudo-points to match with components (STU).</p>
      <p id="d1e699">The other two methods are representative of the dominant DSM method as implemented, with some differences in detail, in many countries and for many properties <xref ref-type="bibr" rid="bib1.bibx51 bib1.bibx25 bib1.bibx1" id="paren.50"><named-content content-type="pre">e.g.</named-content></xref>.</p>
      <p id="d1e708">SG2 <xref ref-type="bibr" rid="bib1.bibx48" id="paren.51"/> uses random forests implemented in the <monospace>ranger</monospace> R package, with prior covariate selection by recursive feature elimination and model tuning by the cross-validation of model hyperparameters (number of covariates at each tree split and number of trees in the forest). The model is trained for the whole world, not per country or region; thus, it is a global model. This is based on the homosoil concept <xref ref-type="bibr" rid="bib1.bibx26" id="paren.52"/> for which identical environmental conditions anywhere in the world should result in the same soils. Its use in DSM assumes that all soil-forming factors are fully specified (i.e. over their whole range and with all their possible interactions) in the model and training set. Due to the uneven distribution of training points in covariate space, and portions of covariate space with no observations, this ideal situation is not met. An obvious question is whether or not the additional information from outside the CONUS leads to an improved model for this region.</p>
      <p id="d1e720">SPCG <xref ref-type="bibr" rid="bib1.bibx50" id="paren.53"/> is an extension of the original SoilGrids approach but uses an ensemble of two tree-based machine learning methods, namely random forests (as in the original SoilGrids) and gradient boosting. The model is trained for the CONUS and not per region; thus, it is a reduced version of the homosoil concept. It is a global model in the sense of “use all information over a wide area”, although this is not the entire globe, as in SG2.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Resolution, depths, and coordinate reference systems</title>
      <p id="d1e734">About 90 % of gSSURGO is derived from polygon maps with a design scale (<inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">12</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">24</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula>, depending on the original survey) which corresponds to MLDs of 1.44 to 2.3 <inline-formula><mml:math id="M30" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ha</mml:mi></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">24</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula>) or 0.38 to 0.575 <inline-formula><mml:math id="M32" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">ha</mml:mi></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">12</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula>) polygons, depending on the definition of the MLD (see above). These correspond to single grid cell resolutions of 240 to 384 <inline-formula><mml:math id="M34" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">24</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula>) or 60 to 96 <inline-formula><mml:math id="M36" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">12</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula>). gNATSGO includes some areas surveyed at a smaller scale (<inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">250</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">000</mml:mn></mml:mrow></mml:math></inline-formula>). gNATSGO is delivered as gridded coverages at 30 m or 90 m horizontal resolution on an Albers equal area projection covering the CONUS, with standard parallels at 29.5 and 45.5<inline-formula><mml:math id="M39" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N and the central meridian at <inline-formula><mml:math id="M40" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>96<inline-formula><mml:math id="M41" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E on the NAD83 datum, which uses the GRS80 ellipsoid. We have used the 30 m resolution product. gSSURGO is delivered as gridded coverages at 10 m or 30 m horizontal resolution on the same CONUS projection.  We have used the 30 m resolution product. Property information is provided per horizon or layer, each with depth limits. Thus, to produce a prediction for a depth interval, these must be aggregated by the depth-weighted average by thickness across the depth interval. PSP predicts at 1 arcsec of longitude and latitude resolution, i.e. 0.0002777778<inline-formula><mml:math id="M42" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> on the WGS84 datum, equivalent to <inline-formula><mml:math id="M43" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> 32 <inline-formula><mml:math id="M44" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> latitude, and proportionally smaller longitude depending on latitude. Depth intervals are the standards specified by GlobalSoilMap. SPCG predicts at 100 <inline-formula><mml:math id="M45" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> resolution for seven point depths (0, 5, 15, 30, 60, 100, and 200 cm) in the same projection as gNATSGO and gSSURGO. Predictions are means of a depth interval. SG2 predicts at 250 <inline-formula><mml:math id="M46" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> resolution for the standard depth intervals specified by GlobalSoilMap on an equal area Interrupted Goode Homolosine (IGH) projection on the WGS84 datum <xref ref-type="bibr" rid="bib1.bibx33" id="paren.54"/>. Depth interval predictions are in fact point predictions at the centre of the depth interval, considered to represent that interval. Section S3 explains how these products are accessed and made compatible for comparison at regional and local scales.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Uncertainty assessment</title>
      <p id="d1e954">SG2 and PSP predict the 5 % and 95 % quantiles of the distribution of predictions. SG2 uses quantile regression forests <xref ref-type="bibr" rid="bib1.bibx29" id="paren.55"><named-content content-type="pre">QRFs;</named-content></xref>, whereas PSP's uncertainty estimates are based on the property data available for each STU predicted by POLARIS. The profile property data are used to create a depth-harmonized profile with uncertainty for each standard depth interval.</p>
      <p id="d1e962">These uncertainty limits are specified by the GlobalSoilMap consortium <xref ref-type="bibr" rid="bib1.bibx2" id="paren.56"/>, defined as “the 90 % Prediction Interval (PI) which reports the range of values within which the true value is expected to occur 9 times out of 10 … there is no assumption that this prediction interval is necessarily symmetric around the predicted value” <xref ref-type="bibr" rid="bib1.bibx54" id="paren.57"/>.</p>
      <p id="d1e971">gNATSGO and gSSURGO provide representative, upper, and lower limit values of each property of an STU per horizon or layer. The National Soil Survey Handbook, <inline-formula><mml:math id="M47" display="inline"><mml:mi mathvariant="italic">§</mml:mi></mml:math></inline-formula> 618.2 <xref ref-type="bibr" rid="bib1.bibx61" id="paren.58"/>, explains that the representative value approximates the median but that the quantiles corresponding to the low and high values can be adjusted to the percentiles which best show the spread of the property within an STU. If there are sufficient laboratory data of the sampled profiles of the STU in the National Soil Information System <xref ref-type="bibr" rid="bib1.bibx35" id="paren.59"><named-content content-type="pre">NASIS;</named-content></xref>, then these are used as the basis for establishing the range. In all cases, expert opinion is used to adjust these to represent the range that a map user can expect to find in the field. Thus these are not directly comparable to the results of QRF but do give some idea of how the field mappers, supported by laboratory observations, conceive of the spread of a property. Note that none of these assessments implies a parametric probability distribution but rather the ranges of selected quantiles only.
<xref ref-type="bibr" rid="bib1.bibx24" id="text.60"/> discuss how these estimates can be derived for USA products following the GlobalSoilMap.net specifications.</p>
      <p id="d1e992">As pointed out by <xref ref-type="bibr" rid="bib1.bibx3" id="text.61"/>, “[t]he user community requires training in, and experience with, the new digital soil map products, especially about the use of uncertainties”. It would be hoped that the uncertainties computed by different methods would be similar.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Evaluation methods</title>
      <p id="d1e1007">We compared DSM products at regional (nominal 250 <inline-formula><mml:math id="M48" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> grid cells) and local (nominal 30 <inline-formula><mml:math id="M49" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> grid cells) levels. We evaluated both qualitatively, i.e. by visual inspection followed by expert interpretation, and numerically, over a <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M51" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> tile, selected based on its diverse soil-forming factors and environments and our familiarity with its soil geography. For the pattern analysis within this area, we selected a <inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.20</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">0.20</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M53" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> subtile and projected the maps to the UTM18N grid on the WGS84 datum (ESPG code 32618).</p>
      <p id="d1e1066">To compare maps at the regional resolution (250 m), the higher-resolution maps (gSSURGO, PSP, and SPCG) were aggregated to the lower resolution by weighted averaging (resampling) of the high-resolution pixels within 1 low-resolution pixel. Thus, there is smoothing inherent in the regional comparisons.</p>
      <p id="d1e1069">To compare maps at the local resolution, we only included the two products (gSSURGO and PSP) provided at that resolution, along with the global product (SG2) as reference, with this latter product downscaled by increasing the grid resolution, without any attempt to disaggregate within the larger grid cell, over a <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.15</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">0.15</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M55" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> subtile.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Qualitative methods</title>
      <p id="d1e1099">Qualitative methods for comparing maps rely on expert judgement to identify known soil geographic patterns and evaluate to what extent they are represented on the gridded maps. The maps are displayed side by side, along with a map of their pairwise differences. Areas of disagreement are identified and discussed.</p>
      <p id="d1e1102">The DSM product can be evaluated at selected known points, typically from the field observation of test areas. The following questions are posed: is the correct soil type or property predicted? And, if not, is the error a reasonable approximation? More interesting are the patterns in the DSM product. These can be compared to patterns used in the mental model of traditional soil survey such as, for example, toposequences and sequences of contrasting parent material.</p>
      <p id="d1e1105">In both cases (points and patterns), the evaluator may be able to infer which DSM covariates would be needed to improve the map.</p><?xmltex \hack{\newpage}?>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Numerical methods – whole map</title>
      <p id="d1e1118">Numerical methods for comparing gridded maps as a whole include (1) MD, which is the mean difference (also known as the bias), i.e. the average disagreement between maps, (2) RMSD, which is root mean squared difference, and (3) RMSD adjusted for MD, i.e. the RMSD after subtracting the bias from each prediction. These take the first-listed map as reference and the second as the map to evaluate. They can be normalized by the number of grid cells or total area. In addition, all maps can be compared by their Pearson (linear) correlations. These methods are of limited interpretive value. Their main use is to characterize the bias (MD) over the entire map; they do not reveal where any discrepancies occur. For example, there can be no bias overall but a large difference in the amount and values of higher and lower differences will be seen. This will be reflected in the RMSD although not shown on a map.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Numerical methods – spatial continuity</title>
      <p id="d1e1129">Soil properties are usually spatially correlated; we expect similar values of properties in nearby grid cells. The degree of local spatial continuity can be assessed by the variogram computed over local neighbourhoods of the gridded map. We computed and modelled the variogram within a local neighbourhood, and it automatically fit with an exponential model, using the <monospace>fit.variogram</monospace> function of the <monospace>gstat</monospace> R package <xref ref-type="bibr" rid="bib1.bibx46" id="paren.62"/>. The spatial structure is characterized by the range, proportional nugget, and structural sill of the fitted variogram model. The range shows the radius over which the selected property has spatial correlation. The proportional nugget shows the variability at the prediction point at the centre of a grid cell at a scale shorter than the grid spacing. The structural sill shows the overall variability within the range. These metrics show differences in spatial continuity (range), total variability (total sill) and short-range unexplained variability (proportional nugget) between maps.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Numerical methods – patterns</title>
      <p id="d1e1149">Numerical methods for comparing patterns include (1) the V-measure method <xref ref-type="bibr" rid="bib1.bibx40" id="paren.63"><named-content content-type="pre">Sect. <xref ref-type="sec" rid="Ch1.S3.SS4.SSS1"/>;</named-content></xref>, implemented in the <monospace>sabre</monospace> (Spatial Association Between REgionalizations) R package <xref ref-type="bibr" rid="bib1.bibx38" id="paren.64"/>, and (2) landscape-level metrics <xref ref-type="bibr" rid="bib1.bibx62" id="paren.65"><named-content content-type="pre">Sect. <xref ref-type="sec" rid="Ch1.S3.SS4.SSS2"/>;</named-content></xref>, as used in ecology and derived from the FRAGSTATS computer program <xref ref-type="bibr" rid="bib1.bibx28" id="paren.66"/> and implemented in the <monospace>landscapemetrics</monospace> R package <xref ref-type="bibr" rid="bib1.bibx18" id="paren.67"/>. These include Shannon diversity and evenness, landscape shape index, and fractal dimension. Although the ecological relevance of FRAGSTATS metrics have been criticized  <xref ref-type="bibr" rid="bib1.bibx22" id="paren.68"/>, here we use them to characterize spatial patterns of soil properties and not as inputs to landscape ecology models. Most of the metrics used here have also been used by <xref ref-type="bibr" rid="bib1.bibx47" id="text.69"/> in a study of urban pedodiversity.</p>
      <p id="d1e1188">These methods must be applied to classified maps, so the continuous soil property maps must first be classified into ranges before analysis. Different choices of class limits and widths will result in different values of these measures. A somewhat objective method to choose classes is histogram equalization. The analyst determines the number of classes, and equal numbers of grid cells are in each class. To compare maps, the combined values of all maps are used to construct the histogram. For the V measure, the gridded maps must be polygonized.</p>
<sec id="Ch1.S3.SS4.SSS1">
  <label>3.4.1</label><title>V measure</title>
      <p id="d1e1198">The V-measure metrics compare different spatial partitions of the same domain, which is, in this case, maps with classified soil properties.
The intent is to reveal how similar are these partitions. There could be two  maps that have the same total areas of each class, and even the same number of polygons within each class and even the same size distribution of these polygons, and yet be completely different in how they partition space into classes.</p>
      <p id="d1e1201">The polygons of a classified map are termed regions of a regionalization in the first (reference) map and zones of a partition in the second map.
These are intersected to produce segment polygons of the combined map, which are labelled with both zone and region classes. These polygons are then used to compute two metrics of the map to be evaluated, (1) homogeneity and (2) completeness, both with respect to the regionalization of the reference map.</p>
      <p id="d1e1204">The homogeneity of the second map is a measure of the variance of the regions within a zone normalized by the variance of the regions in the entire domain of the first map. These variances are computed by the Shannon entropy based on areas of the segments. If the variance of the regions within the zones is small, then the partition is relatively homogeneous with respect to the regionalization. A perfectly homogeneous partition (with value <inline-formula><mml:math id="M56" display="inline"><mml:mn mathvariant="normal">1</mml:mn></mml:math></inline-formula>) is when each zone of the second map is within a single region of the reference map. In this case, each zone has only one reference class. A perfectly inhomogeneous partition (with value <inline-formula><mml:math id="M57" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula>) is when each zone has the same composition of regions as the entire domain of the first map, i.e. the second map's partition (to be evaluated) is essentially random with respect to the first map's regionalization.</p>
      <p id="d1e1221">The completeness of the second map is the inverse of homogeneity; it assesses the variance of the zones within a region normalized by the variance of the zones in the entire domain of the second map. It evaluates the homogeneity of regions with respect to zones and shows how well the regionalization of the reference map fits inside the partition of the map to be evaluated. A perfectly complete regionalization is when each region of the reference map is entirely within a single zone of the map to be evaluated.
In this case, a polygon of the reference map will not be split among zones.</p>
      <p id="d1e1225">These two together are combined into a single measure, the V measure, as the harmonic mean of homogeneity <inline-formula><mml:math id="M58" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula> and completeness <inline-formula><mml:math id="M59" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> (Eq. <xref ref-type="disp-formula" rid="Ch1.E1"/>). This has a range between 0 (no spatial association between the maps) and 1 (perfect association). Obviously, we prefer high association between maps produced by DSM and a reference map. We can also assess the agreement of the patterns produced by different DSM methods by selecting one as a reference.
              <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M60" display="block"><mml:mrow><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>h</mml:mi><mml:mo>×</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>+</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="Ch1.S3.SS4.SSS2">
  <label>3.4.2</label><title>Landscape metrics</title>
      <p id="d1e1279">The landscape metrics applicable to soil maps (as opposed to, e.g., maps of vegetation types) have diverse interpretations. We compare the metrics of two maps to see if they have a similar concept of the soil landscape. The <monospace>landscapemetrics</monospace> package can compute many FRAGSTAT indices. We chose several that show the landscape-level difference between maps. We did not consider metrics of individual patches, except that they contribute to landscape-level metrics. The algorithms for these can be found in the package code repository <xref ref-type="bibr" rid="bib1.bibx17" id="paren.70"/>; here we present the formulas and their interpretations.
<list list-type="bullet"><list-item>
      <p id="d1e1290">The Shannon diversity index, <monospace>shdi</monospace> (Eq. <xref ref-type="disp-formula" rid="Ch1.E2"/>), where <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the proportion of pixels of class <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>,   characterizing the landscape diversity according to two factors, i.e. number of classes and their proportions. It is widely used as a summary diversity measure, although it does not distinguish between the two factors. More classes and/or a more even distribution of proportions lead to a higher landscape diversity. This does not account for spatial contiguity; it just considers the class of each pixel, irrespective of position. In this example, the number of classes in each map will be similar, with a maximum of eight (the chosen histogram equalization classes computed over the combined range of all maps), but some maps may lack representatives of the highest or lowest classes and so will have only seven classes.<disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M63" display="block"><mml:mrow><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>ln⁡</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item><list-item>
      <p id="d1e1367">The Shannon evenness index, <monospace>shei</monospace> (Eq. <xref ref-type="disp-formula" rid="Ch1.E3"/>), is a normalization of Shannon diversity by the maximum diversity possible for the given number of classes (<inline-formula><mml:math id="M64" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>). It varies from 0 (completely uneven distribution and low landscape diversity) to 1 (all proportions are equal and high landscape diversity). It does not depend on the number of classes and thus isolates the effect of class proportion.<disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M65" display="block"><mml:mrow><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi>D</mml:mi><mml:mrow><mml:mi>ln⁡</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item><list-item>
      <p id="d1e1403">The landscape shape index, <monospace>lsi</monospace> (Eq. <xref ref-type="disp-formula" rid="Ch1.E4"/>), where <inline-formula><mml:math id="M66" display="inline"><mml:mi>A</mml:mi></mml:math></inline-formula> is the total area of the landscape, and <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msup><mml:mi>E</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> is the total length of edges, including the boundary, quantifies the internal boundary complexity of a landscape tile, with a value of 1 when the landscape consists of a single square patch, increasing without limit as the length of edges within the landscape increases. This metric characterizes the degree of compactness of the contiguous areas of the classes.<disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M68" display="block"><mml:mrow><mml:mi mathvariant="normal">LSI</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">0.25</mml:mn><mml:msup><mml:mi>E</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow><mml:msqrt><mml:mi>A</mml:mi></mml:msqrt></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item><list-item>
      <p id="d1e1454">The landscape aggregation index, <monospace>lai</monospace> (Eq. <xref ref-type="disp-formula" rid="Ch1.E5"/>), where <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the number of like adjacencies, <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mo>max⁡</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the class-wise maximum possible number of like adjacencies of class <inline-formula><mml:math id="M71" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> (i.e. if all pixels in the class were in one cluster), and <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the proportion of landscape comprised of class <inline-formula><mml:math id="M73" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> to weight the index by class prevalence. Thus, <monospace>lai</monospace> equals the number of like adjacencies divided by the theoretical maximum possible number of like adjacencies and summed over each class and over the entire landscape.   It ranges from 0 for maximally disaggregated to 100 for maximally aggregated landscapes. This metric characterizes how dispersed the classes are.<disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M74" display="block"><mml:mrow><mml:mi mathvariant="normal">AI</mml:mi><mml:mo>=</mml:mo><mml:mo mathsize="2.5em">[</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:munderover><mml:mo mathsize="2.5em">(</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>max⁡</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo mathsize="2.5em">)</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo mathsize="2.5em">]</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item><list-item>
      <p id="d1e1595">The mean fractal dimension, <monospace>frac_mn</monospace>, characterizes the complexity of the landscape as the mean of the fractal dimension of all patches in the landscape. It approaches 1 if all patches are square and 2 if all patches are irregular. It is scale independent. The patch-level fractal dimensions are computed from the patch perimeters <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> in linear units and areas <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> in square units; these are then averaged to obtain <monospace>frac_mm</monospace>.<disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M77" display="block"><mml:mrow><mml:mi mathvariant="normal">FRAC</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>⋅</mml:mo><mml:mi>ln⁡</mml:mi><mml:mo>⋅</mml:mo><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">0.25</mml:mn><mml:mo>⋅</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mi>ln⁡</mml:mi><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item><list-item>
      <p id="d1e1680">The co-occurrence vector, <monospace>cove</monospace>, proposed by  <xref ref-type="bibr" rid="bib1.bibx39" id="text.71"/> summarizes the entire adjacency structure of the map and can be used to compare map structures. This is a normalized form of the co-occurrence matrix, which counts all the pairs of the adjacent cells for each category in a local landscape in the form of a cross-classification matrix. This vector can be considered as a probability vector for the co-occurrence of different classes.   Co-occurrence vectors of different categorical maps can then be compared by computing the distance between them. Many distance measures are possible; we choose the Jensen–Shannon distance (Eq. <xref ref-type="disp-formula" rid="Ch1.E7"/>), which computes the entropy <inline-formula><mml:math id="M78" display="inline"><mml:mi>H</mml:mi></mml:math></inline-formula> of each probability vector <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and entropy of their average and, from these, the distance in entropy space between them. Increasing values indicate increasing dissimilarity in the adjacency patterns. The computation of <monospace>cove</monospace> is implemented in the <monospace>motif</monospace> R package and the Jensen–Shannon distance in the <monospace>philentropy</monospace> R package.<disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M80" display="block"><mml:mrow><mml:mi mathvariant="normal">JSD</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mfenced close="]" open="["><mml:mrow><mml:mi>H</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:mi>H</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>v</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item></list></p>
</sec>
</sec>
<sec id="Ch1.S3.SS5">
  <label>3.5</label><title>Regional patterns</title>
      <p id="d1e1803">Regional patterns are at the scale of regional trends such as lithologic units, elevation zones in mountains, and repeating patterns (e.g. basin and range and ridge and valley). The gNATSGO maps are taken as the reference, although we are well aware that they may not always correspond to ground truth. We then comment on the differences and speculate on the causes, based on our knowledge of the DSM procedures used to make each product and the nature of the soil landscape.</p>
</sec>
<sec id="Ch1.S3.SS6">
  <label>3.6</label><title>Local patterns</title>
      <p id="d1e1814">Local patterns are at the scale of geomorphic features such as hillslope catenas, fluvial terraces, outwash fans, valley trains, and drumlin fields.
This evaluation occurs within the test area for the regional patterns but examines a smaller area with a distinctive soil–landscape pattern. The gSSURGO maps are taken as the reference. This was evaluated by two methods, as follows.</p>
<sec id="Ch1.S3.SS6.SSS1">
  <label>3.6.1</label><title>Visual method</title>
      <p id="d1e1824">We produced ground overlays of key soil properties at selected depth intervals, with corresponding Keyhole Markup Language (KML) specifications, and displayed these in Google Earth as semi-transparent overlays, using the original resolution of each product, projected into WGS84 geographic coordinates as required by Google Earth. These were then compared with gNATSGO maps streamed within Google Earth by SoilWeb Earth <xref ref-type="bibr" rid="bib1.bibx9" id="paren.72"/>. This shows the mapped polygons labelled with their map unit and linked to the map unit description, which in turn is linked to the Official Series Descriptions <xref ref-type="bibr" rid="bib1.bibx42" id="paren.73"><named-content content-type="pre">OSD;</named-content></xref>, with a complete description of the soil properties modal values and ranges.</p>
</sec>
<sec id="Ch1.S3.SS6.SSS2">
  <label>3.6.2</label><title>Quantitative method</title>
      <p id="d1e1843">This follows the procedures of the regional assessment, except that V measures are not computed, due to the very fine pattern of classified polygons.</p><?xmltex \hack{\newpage}?>
</sec>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Example area and soil property</title>
      <p id="d1e1857">To illustrate the method, we selected one area familiar to the first author and an important soil property with strong spatial variability and pattern, namely pH in the 0–5 layer. We selected this property because, in our experience, this is often well-modelled by DSM methods. For example, SG2 had global cross-validation statistics of 0.78 pH median RMSD and a model efficiency coefficient (MEC; the <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> of the <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> line actual vs.  observed) of 0.67 <xref ref-type="bibr" rid="bib1.bibx48" id="paren.74"/>. We select the topmost depth interval because it is most represented by many environmental covariates, especially land cover and those derived from remote sensing. Thus, the example shown here may be the best case, where we would hope that all mapping methods should provide similar results.</p>
      <p id="d1e1886">The example area is in central New York state bounding box (<inline-formula><mml:math id="M83" display="inline"><mml:mo lspace="0mm">-</mml:mo></mml:math></inline-formula>77 to <inline-formula><mml:math id="M84" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76<inline-formula><mml:math id="M85" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E, 42–43<inline-formula><mml:math id="M86" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N). The subtile for the pattern evaluation was <inline-formula><mml:math id="M87" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76.8 to <inline-formula><mml:math id="M88" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76.6<inline-formula><mml:math id="M89" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E and (42.2–42.4<inline-formula><mml:math id="M90" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, centred at Cayuta, NY. The regional geomorphology is described by <xref ref-type="bibr" rid="bib1.bibx7" id="text.75"/>. The underlying bedrock is a sedimentary sequence from Ordovician (north) to upper Devonian (south), with a wide variety of sedimentary facies. A strip of the bedrock geology map <xref ref-type="bibr" rid="bib1.bibx36" id="paren.76"/> covering part of the study area is shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>.</p>
      <p id="d1e1962">The entire area has been glaciated, with the portion north of about 42<inline-formula><mml:math id="M91" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>15<inline-formula><mml:math id="M92" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> (Valley Heads terminal moraines) somewhat more recently than the southern portion. A fragment of the surficial geology map <xref ref-type="bibr" rid="bib1.bibx37" id="paren.77"/> is shown in Fig. <xref ref-type="fig" rid="Ch1.F2"/>. This shows strongly expressed features resulting from the most recent glaciation; these are well-known to the traditional soil surveyors. Many glacial features are present and relevant to soil geography, including ground moraine, deep glacial troughs with proglacial lake sediments, beach lines, outwash valley trains, kame terraces,  and hanging deltas. Soil reaction in the northern half is largely controlled by the limestone spread by the glacier from outcrops of the Onondaga and Tully limestones (Fig. <xref ref-type="fig" rid="Ch1.F1"/>) that decrease to the south.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e1993">Bedrock geology of central New York state, with the transect from 43  (left) to 42<inline-formula><mml:math id="M93" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N and centred on <inline-formula><mml:math id="M94" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76<inline-formula><mml:math id="M95" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>30<inline-formula><mml:math id="M96" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> E. The orientation is north (left) to south (right). The chronological and topographic sequence from Upper Silurian (N) through Upper Devonian (S) sedimentary rocks, notably the Onondaga limestone (green, Don) and Tully limestone (crosshatched red, Dt) are shown
<xref ref-type="bibr" rid="bib1.bibx36" id="paren.78"><named-content content-type="pre">Source:</named-content></xref>.</p></caption>
        <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f01.png"/>

      </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e2043">Surficial geology of central New York state near Moravia, NY. The ground moraine (pink; if stippled shallow over bedrock), proglacial lakes (brown), organic swamps (dark green), bedrock or very thin soil cover (red), till moraine (purple), kame moraines (orange), lacustrine sand (light green), and outwash sand and gravel (yellow) are shown  <xref ref-type="bibr" rid="bib1.bibx37" id="paren.79"><named-content content-type="pre">Source:</named-content></xref>.</p></caption>
        <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f02.png"/>

      </fig>

</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Regional spatial patterns</title>
<sec id="Ch1.S5.SS1">
  <label>5.1</label><title>Visual method</title>
      <p id="d1e2073">A visual inspection of a DSM product over the landscape can be useful to identify anomalies and the degree to which the DSM product captures landscape features. These are over small areas where the soil–landscape relation is known to the evaluator. This cannot be part of a systematic evaluation but can reveal areas of concern or agreement.</p>
      <p id="d1e2076">As an example, Fig. <xref ref-type="fig" rid="Ch1.F3"/> shows SSURGO map units, from a 1965 <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> k design scale soil survey, with a minimum legible delineation 1.6 ha <xref ref-type="bibr" rid="bib1.bibx11" id="paren.80"/> and draped over a ground overlay of pH (0–5 cm) from SG2, produced by the SoilWeb streaming coverage in Google Earth Pro, with a point query showing the SSURGO map unit composition (Fig. <xref ref-type="fig" rid="Ch1.F4"/>). The map unit is described by its constituent soil series and their estimated proportions. Each series can then be queried for its OSD <xref ref-type="bibr" rid="bib1.bibx42" id="paren.81"/>, which gives a typical profile, a range of properties, and a link to lab data for the series. In this case,  the pattern of properties as predicted by SG2 somewhat follows the map unit delineations but at a much coarser resolution. This is especially evident at the transition from the end moraine (map units beginning with <monospace>H</monospace>) and the steep slopes with thin till from the local bedrock (map unit <monospace>LoF</monospace>).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e2110">SoilWeb view of the SSURGO map units and a ground overlay of pH, at 0–5 cm, predicted by SG2. The colours are low (red) to yellow (high) pH. The centre is at <inline-formula><mml:math id="M98" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76<inline-formula><mml:math id="M99" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>38<inline-formula><mml:math id="M100" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>04<inline-formula><mml:math id="M101" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> E, 42<inline-formula><mml:math id="M102" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>20<inline-formula><mml:math id="M103" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>07<inline-formula><mml:math id="M104" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> N. An interactive view of SSURGO is available at <uri>https://casoilresource.lawr.ucdavis.edu/gmap/?loc=42.33215,-76.63590,z15</uri> (last access: 18 August 2022). The background is from © Google Earth.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f03.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e2193">A SSURGO map unit <monospace>HrC</monospace> composition at  <inline-formula><mml:math id="M105" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76<inline-formula><mml:math id="M106" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>38<inline-formula><mml:math id="M107" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>05<inline-formula><mml:math id="M108" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> E, 42<inline-formula><mml:math id="M109" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>19<inline-formula><mml:math id="M110" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>53<inline-formula><mml:math id="M111" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> N. An interactive map unit summary is available at <uri>https://casoilresource.lawr.ucdavis.edu/soil_web/list_components.php?mukey=295620</uri> (last access: 18 August 2022).</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f04.png"/>

        </fig>

</sec>
<sec id="Ch1.S5.SS2">
  <label>5.2</label><title>Regional maps</title>
      <p id="d1e2284">Table <xref ref-type="table" rid="Ch1.T1"/> shows the statistical differences between  gNATSGO (reference) and the DSM products. All DSM products underpredict topsoil pH with respect to gNATSGO by about 0.38–0.48 pH units. The RMSD is substantial also, of the order of 0.49–0.67 pH units, and is somewhat less than this when corrected for bias.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e2292">Statistical differences between gNATSGO and DSM products for pH <inline-formula><mml:math id="M112" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">DSM product</oasis:entry>
         <oasis:entry colname="col2">MD</oasis:entry>
         <oasis:entry colname="col3">RMSD</oasis:entry>
         <oasis:entry colname="col4">RMSD.Adjusted</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">SG2</oasis:entry>
         <oasis:entry colname="col2">3.796</oasis:entry>
         <oasis:entry colname="col3">6.111</oasis:entry>
         <oasis:entry colname="col4">4.789</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PSP</oasis:entry>
         <oasis:entry colname="col2">3.843</oasis:entry>
         <oasis:entry colname="col3">4.908</oasis:entry>
         <oasis:entry colname="col4">3.052</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SPCG</oasis:entry>
         <oasis:entry colname="col2">4.815</oasis:entry>
         <oasis:entry colname="col3">6.693</oasis:entry>
         <oasis:entry colname="col4">4.649</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e2382">Figure <xref ref-type="fig" rid="Ch1.F5"/> shows whole-map histograms.
PSP has a bimodal distribution, and predicts few pH values around pH 5.8.
This was unexpected, since this value is well-represented in the gNATSGO map.
It may be an artefact of a covariate that is influential over a wider area than this tile and results in two regional distributions from contrasting elevation or climate zones. The other distributions are fairly symmetric, although SG2 and SPCG are more even than gNATSGO, which is strongly concentrated near pH 6.2. This shows the smoothing effect of the machine learning models.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e2390">Histograms of pH <inline-formula><mml:math id="M113" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm. Note the bimodal distribution of PSP and the flatter distributions of SG2 and SPCG compared to gNATSGO.</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f05.png"/>

        </fig>

      <p id="d1e2406">Figure <xref ref-type="fig" rid="Ch1.F6"/> shows the pairwise Pearson correlations between the products. The products are overall well correlated. SG2 and SPCG are very closely correlated, since they use similar mapping methods, despite the additional covariates used by SPCG. PSP and gNATSGO are also closely correlated. These correlations do not account for bias. They do, however, show that the maps are similar in their overall pattern as they are evaluated per grid cell.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e2413">Pearson correlations between all products for  pH at 0–5 cm. Strong correlations, especially between gNATSGO vs. PSP and SG2 vs. SPCG, are shown.</p></caption>
          <?xmltex \igopts{width=213.395669pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f06.png"/>

        </fig>

      <p id="d1e2422">Figure <xref ref-type="fig" rid="Ch1.F7"/> shows gNATSGO (reference) along with the predictions of pH of the PSP products. Figure <xref ref-type="fig" rid="Ch1.F8"/> shows these as difference maps.
These figures reveal substantial differences between products. The most obvious is in the detail of the spatial pattern. Despite having been upscaled to s regional resolution, gNATSGO shows finer detail than the other products, especially when compared to PSP.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e2431">Topsoil (0–5 cm) pH <inline-formula><mml:math id="M114" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10, according to gNATSGO and DSM products. See the text for a discussion.</p></caption>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f07.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><?xmltex \def\figurename{Figure}?><label>Figure 8</label><caption><p id="d1e2450">Difference between gNATSGO and DSM products for pH <inline-formula><mml:math id="M115" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm. See the text for a discussion.</p></caption>
          <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f08.png"/>

        </fig>

      <p id="d1e2466">These figures also show the spatial distribution of the bias compared to gNATSGO (as reference). SG2 and SPCG underpredict pH in the higher hills in the NE portion of the map and in the glaciolacustrine sediments along the lakeshores. The disagreement along the lakeshores is because SG2 and SPCG do not use a surficial geology map, which would be especially useful in recently glaciated areas such as this. The disagreement in the higher hills seems to be a direct result of elevation. This is not because of extrapolation in feature space because, at these elevations, SG2 also misses the soils derived from Onondaga limestone glacial till towards the southern end of the till plain. SG2 has no information on parent material and uses global models. SPCG has very similar differences, despite using SSURGO-derived parent material as a covariate.</p>
      <p id="d1e2469">PSP predictions are closer to gNATSGO than those of SG2 are, which is not surprising since PSP also uses gNATSGO as its primary information source.
This product has removed some of the fine variation in gNATSGO. However the disaggregation by DSMART results in some discrepancies with gNATSGO. In particular, the Homer–Tully outwash valley (northeastern side of map) is underpredicted by 1 pH unit, and the surrounding hills are overpredicted by almost as much. Many of the valley trains (southern side of map, running towards the Susquehanna River) are underpredicted. This is likely due to PSP's soil series predictions, which are based on estimated map unit composition and random selection of series locations within map units for DSM calibration.</p>
</sec>
<sec id="Ch1.S5.SS3">
  <label>5.3</label><title>Uncertainty</title>
      <p id="d1e2480">The 5 %, 50 %, and 95 % prediction quantile maps are shown in Figs. <xref ref-type="fig" rid="Ch1.F9"/> (SG2) and <xref ref-type="fig" rid="Ch1.F10"/> (PSP). The low, representative, and high values from gNATSGO are shown in  Fig. <xref ref-type="fig" rid="Ch1.F11"/>. Each figure has its own stretch. gNATSGO has narrower ranges than the two DSM products and, by design, does not include unrealistic values. SG2 and PSP have unrealistically wide ranges at all locations. In addition, PSP shows a curious feature, i.e. fine patterning at the two extremes that is not present at the median prediction.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9" specific-use="star"><?xmltex \currentcnt{9}?><?xmltex \def\figurename{Figure}?><label>Figure 9</label><caption><p id="d1e2491">Quantiles of the prediction, SG2, for pH <inline-formula><mml:math id="M116" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm. Note the unrealistically wide range at all locations and consistent patterning among quantiles.</p></caption>
          <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f09.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F10" specific-use="star"><?xmltex \currentcnt{10}?><?xmltex \def\figurename{Figure}?><label>Figure 10</label><caption><p id="d1e2509">Quantiles of the prediction, PSP, for pH <inline-formula><mml:math id="M117" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm. Note the unrealistically wide range at all locations and the fine patterning at the two extremes.</p></caption>
          <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f10.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F11" specific-use="star"><?xmltex \currentcnt{11}?><?xmltex \def\figurename{Figure}?><label>Figure 11</label><caption><p id="d1e2528">Low, representative, and high values from gNATSGO for pH <inline-formula><mml:math id="M118" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm.</p></caption>
          <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f11.png"/>

        </fig>

      <p id="d1e2544">Figure <xref ref-type="fig" rid="Ch1.F12"/> shows the interquartile range (IQR) 5 %–95 % for the two DSM products, along with the low–high range for gNATSGO. SG2 has a fairly consistent IQR, mostly from about 2.5 to 3.5 pH, whereas PSP has a much wider range of uncertainties, mostly from about 1.5 to 4.5 pH, and shows much more spatial pattern. PSP has the widest ranges on the steep valley sides, especially in the Seneca Army Depot at the north interlake area, and the lowest on the broad till plains and through valleys. These are wide ranges and, although they are an honest reflection of the DSM models, should give pause to map users. This suggests that the GlobalSoilMap specifications for uncertainty <xref ref-type="bibr" rid="bib1.bibx2" id="paren.82"/> are unduly pessimistic. Sources for uncertainty assessment (SG2 for training points and global covariates and PSP for mapped soil series and national covariates) and the different machine learning methods lead to greatly different estimates of prediction uncertainty. The gNATSGO low–high range is narrower than the DSM IQR, but these are not comparable because the expert-assigned range is not based on an estimate of a 5 %–95 %  IQR.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F12" specific-use="star"><?xmltex \currentcnt{12}?><?xmltex \def\figurename{Figure}?><label>Figure 12</label><caption><p id="d1e2554">Interquartile ranges at 0.05–0.95 for pH <inline-formula><mml:math id="M119" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm. The SG2 IQR is fairly consistent, from about 2.5 to 3.5 pH. The PSP IQR has a wider range and more spatial patterning. The gNATSGO low–high range is narrower.</p></caption>
          <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f12.png"/>

        </fig>

      <p id="d1e2570">Figure <xref ref-type="fig" rid="Ch1.F13"/> shows the differences between the IQR of the DSM products and the low–high range from gNATSGO. Both DSM products have substantially wider ranges than gNATSGO almost everywhere; however, the pattern of differences is not similar. For example, the difference with SG250 is much larger in the north of the study area, whereas PSP has the larger differences in the southern hills.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F13" specific-use="star"><?xmltex \currentcnt{13}?><?xmltex \def\figurename{Figure}?><label>Figure 13</label><caption><p id="d1e2577">IQR<inline-formula><mml:math id="M120" display="inline"><mml:mo>/</mml:mo></mml:math></inline-formula>range at <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">95</mml:mn></mml:mrow></mml:math></inline-formula> % vs. low/high, POLARIS–gNATSGO <bold>(a)</bold>, and SG250–gNATSGO <bold>(b)</bold> for pH <inline-formula><mml:math id="M122" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm. Each figure has its own stretch.</p></caption>
          <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f13.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F14" specific-use="star"><?xmltex \currentcnt{14}?><?xmltex \def\figurename{Figure}?><label>Figure 14</label><caption><p id="d1e2621">Fitted variograms for pH at 0–5 cm. Semivariance units are <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">pH</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">10</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. Note the shorter range of gNATSGO and the low sill of PSP.</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f14.png"/>

        </fig>

</sec>
<sec id="Ch1.S5.SS4">
  <label>5.4</label><title>Local spatial autocorrelation</title>
      <p id="d1e2657">The local variograms and their fitted exponential models are shown in Fig. <xref ref-type="fig" rid="Ch1.F14"/>. Table <xref ref-type="table" rid="Ch1.T2"/> shows their statistics.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e2667">Fitted variogram parameters for pH at 0–5 cm. The effective range is in metres, the structural sill is in <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">pH</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">10</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, and the proportional nugget is <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.89}[.89]?><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Product</oasis:entry>
         <oasis:entry colname="col2">Effective range</oasis:entry>
         <oasis:entry colname="col3">Structural sill</oasis:entry>
         <oasis:entry colname="col4">Proportional nugget</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">gNATSGO</oasis:entry>
         <oasis:entry colname="col2">1938.00</oasis:entry>
         <oasis:entry colname="col3">10.32</oasis:entry>
         <oasis:entry colname="col4">0.00</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SG2</oasis:entry>
         <oasis:entry colname="col2">3699.00</oasis:entry>
         <oasis:entry colname="col3">12.93</oasis:entry>
         <oasis:entry colname="col4">0.00</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SPCG</oasis:entry>
         <oasis:entry colname="col2">6924.00</oasis:entry>
         <oasis:entry colname="col3">11.81</oasis:entry>
         <oasis:entry colname="col4">0.01</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PSP</oasis:entry>
         <oasis:entry colname="col2">3918.00</oasis:entry>
         <oasis:entry colname="col3">6.50</oasis:entry>
         <oasis:entry colname="col4">0.02</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table></table-wrap>

      <p id="d1e2800">gNATSGO has the shortest effective range. This indicates a fine-scale structure  at 250 <inline-formula><mml:math id="M126" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> resolution, which is of the same order as the minimum legible delineation (MLD) as a grid cell (see the Introduction). The mappers who defined the boundaries between soil classes (and thus representative property values) were able to divide the landscape at this high spatial frequency, if appropriate to the soil pattern. The DSM products have longer ranges, likely due to the longer-range spatial continuity in many of the covariates. PSP has a longer range and is lower sill than the gNATSGO from which it is derived due to the harmonization inherent in the DSMART algorithm. It has the highest proportional nugget, due to DSMART randomly assigning pixels within a gNATSGO map unit to its constituents, so that the neighbouring pixels may be contrasting at the shortest separation. The very low proportional nuggets of the other products are due to the coarse resolution.</p>
</sec>
<sec id="Ch1.S5.SS5">
  <label>5.5</label><title>Classification</title>
      <p id="d1e2820">Figure <xref ref-type="fig" rid="Ch1.F15"/> shows the topsoil pH classified into eight histogram-equalized classes in a <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.2</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M128" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> subtile. The class limits are approximately 5.01, 5.14, 5.27, 5.40, 5.54, 5.71, and 6.02 pH, with the extreme values of 4.52 and 6.96 pH. The maps show obvious spatial differences in class distribution. gNATSGO shows more areas in the highest pH class than the DSM products, which is consistent with the results from continuous property maps. The pattern of gNATSGO is the coarsest because the classified values come from minimum-area polygons, whereas the DSM products predict per grid cell. PSP shows the finest spatial pattern because of its disaggregation algorithm that randomly divides gNATSGO polygons according to component proportion. If these components are in different pH classes, there will be a fine-scale pattern within the original polygon. This is clearly the case in the large gNATSGO polygon in the northeastern portion of the map (Connecticut Hill). In this example, SPCG shows large homogeneous areas of the lowest pH class, covering the highest hills, whereas SG2 presents a more nuanced view.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F15" specific-use="star"><?xmltex \currentcnt{15}?><?xmltex \def\figurename{Figure}?><label>Figure 15</label><caption><p id="d1e2847">pH classes at 0–5 cm and central NY in detail. Most areas of gNATSGO are in higher pH classes. PSP has the finest spatial pattern due to the DSMART disaggregation algorithm.</p></caption>
          <?xmltex \igopts{width=412.564961pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f15.png"/>

        </fig>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T3"><?xmltex \currentcnt{3}?><label>Table 3</label><caption><p id="d1e2859">V-measure statistics for pH <inline-formula><mml:math id="M129" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm.</p></caption><oasis:table frame="topbot"><?xmltex \begin{scaleboxenv}{.89}[.89]?><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">DSM products</oasis:entry>
         <oasis:entry colname="col2">V measure</oasis:entry>
         <oasis:entry colname="col3">Homogeneity</oasis:entry>
         <oasis:entry colname="col4">Completeness</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">gNATSGO vs. SG2</oasis:entry>
         <oasis:entry colname="col2">0.0128</oasis:entry>
         <oasis:entry colname="col3">0.0143</oasis:entry>
         <oasis:entry colname="col4">0.0116</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">gNATSGO vs. SPCG</oasis:entry>
         <oasis:entry colname="col2">0.0258</oasis:entry>
         <oasis:entry colname="col3">0.0275</oasis:entry>
         <oasis:entry colname="col4">0.0243</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">gNATSGO vs. PSP</oasis:entry>
         <oasis:entry colname="col2">0.084</oasis:entry>
         <oasis:entry colname="col3">0.0897</oasis:entry>
         <oasis:entry colname="col4">0.079</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SPCG vs. SG2</oasis:entry>
         <oasis:entry colname="col2">0.3342</oasis:entry>
         <oasis:entry colname="col3">0.3495</oasis:entry>
         <oasis:entry colname="col4">0.3201</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup><?xmltex \end{scaleboxenv}?></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S5.SS6">
  <label>5.6</label><title>V measure</title>
      <p id="d1e2972">Table <xref ref-type="table" rid="Ch1.T3"/> shows the statistics from several V-measure comparisons, based on the histogram-equalized class maps. Only SG2 and SPCG have somewhat comparable patterns. gNATSGO is considerably different from the DSM products because of its derivation from minimum-area polygons.</p>
      <p id="d1e2977">Figure <xref ref-type="fig" rid="Ch1.F16"/> shows the inhomogeneity and incompleteness of the SG2 pH class map (the second map for the V measure), with respect to the gNATSGO pH class map (the reference map). These values are the inverse of the composite values of Table <xref ref-type="table" rid="Ch1.T3"/> because the very low values in the table correspond to high values in the figure. In the homogeneity map, the blue polygons are the most homogeneous areas of the SG2 map, i.e. where an SG2 polygon has the most homogeneous set of gNATSGO classified values and thus comes closest to the reference. In the completeness map, the blue polygons are the most complete areas of the SG2 map, i.e. where the gNATSGO reference map has the most homogeneous set of SG2 classified values. The two maps have no areas with similar patterns.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F16" specific-use="star"><?xmltex \currentcnt{16}?><?xmltex \def\figurename{Figure}?><label>Figure 16</label><caption><p id="d1e2986">Homogeneity <bold>(a)</bold> and completeness <bold>(b)</bold> measures of the SG2 pH class map, with respect to the reference gNATSGO pH class map, at 0–5 cm. The values are the inhomogeneity of each zone <bold>(a)</bold> and the incompleteness of each region <bold>(b)</bold>.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f16.png"/>

        </fig>

      <p id="d1e3008">A contrasting result is shown in Fig. <xref ref-type="fig" rid="Ch1.F17"/>, which compares the SG2 pH class map with respect to the SPCG map. These maps were made with similar methods and at the same resolution. The inhomogeneity and incompleteness are much lower than for the previous map, showing that the pattern of these two classified maps are fairly similar.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F17" specific-use="star"><?xmltex \currentcnt{17}?><?xmltex \def\figurename{Figure}?><label>Figure 17</label><caption><p id="d1e3015">Homogeneity <bold>(a)</bold> and completeness <bold>(b)</bold> measures of the SG2 pH class map, with respect to the SPCG pH class map, at 0–5 cm.
The values are the inhomogeneity of each zone <bold>(a)</bold> and the incompleteness of each region <bold>(b)</bold>.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f17.png"/>

        </fig>

</sec>
<sec id="Ch1.S5.SS7">
  <label>5.7</label><title>Landscape metrics</title>
      <p id="d1e3044">Table <xref ref-type="table" rid="Ch1.T4"/> shows the statistics from the landscape metrics calculations. The mean fractal dimensions are almost identical. There is quite some range of aggregations, with SPCG being the most aggregated, i.e. the least complex. PSP has the most complex landscape shape due to its fine-scale disaggregation of gSSURGO polygons. The Shannon diversity indices are highest for SG2, indicating that there is the most even areal division into classes. This may be an artefact of the histogram equalization.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T4"><?xmltex \currentcnt{4}?><label>Table 4</label><caption><p id="d1e3052">Landscape metrics statistics for pH at 0–5 cm. <monospace>frac_mn</monospace> is the mean fractal dimension, <monospace>lsi</monospace> is the landscape shape index, <monospace>shdi</monospace> is the Shannon diversity, <monospace>shei</monospace> is the Shannon evenness, and <monospace>ai</monospace> is the aggregation index.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Product</oasis:entry>
         <oasis:entry colname="col2"><monospace>ai</monospace></oasis:entry>
         <oasis:entry colname="col3"><monospace>frac_mn</monospace></oasis:entry>
         <oasis:entry colname="col4"><monospace>lsi</monospace></oasis:entry>
         <oasis:entry colname="col5"><monospace>shdi</monospace></oasis:entry>
         <oasis:entry colname="col6"><monospace>shei</monospace></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">gNATSGO</oasis:entry>
         <oasis:entry colname="col2">48.188</oasis:entry>
         <oasis:entry colname="col3">1.034</oasis:entry>
         <oasis:entry colname="col4">22.602</oasis:entry>
         <oasis:entry colname="col5">1.666</oasis:entry>
         <oasis:entry colname="col6">0.801</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SG2</oasis:entry>
         <oasis:entry colname="col2">50.659</oasis:entry>
         <oasis:entry colname="col3">1.034</oasis:entry>
         <oasis:entry colname="col4">21.768</oasis:entry>
         <oasis:entry colname="col5">2.06</oasis:entry>
         <oasis:entry colname="col6">0.991</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SPCG</oasis:entry>
         <oasis:entry colname="col2">58.483</oasis:entry>
         <oasis:entry colname="col3">1.041</oasis:entry>
         <oasis:entry colname="col4">18.557</oasis:entry>
         <oasis:entry colname="col5">1.887</oasis:entry>
         <oasis:entry colname="col6">0.907</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PSP</oasis:entry>
         <oasis:entry colname="col2">47.025</oasis:entry>
         <oasis:entry colname="col3">1.04</oasis:entry>
         <oasis:entry colname="col4">23.232</oasis:entry>
         <oasis:entry colname="col5">1.898</oasis:entry>
         <oasis:entry colname="col6">0.913</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e3209">Table <xref ref-type="table" rid="Ch1.T5"/> shows the Jensen–Shannon distance between co-occurrence vectors of the four products. The co-occurrence patterns of SG2 is quite similar to that of the other DSM products, whereas gNATSGO is quite different to PSP and SPCG and somewhat different to SG2. This shows that, given this histogram equalization and for the selected property and depth interval, none of the DSM products match the pattern from the traditional soil survey well.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T5"><?xmltex \currentcnt{5}?><label>Table 5</label><caption><p id="d1e3218">The Jensen–Shannon distance between co-occurrence vectors.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">gNATSGO</oasis:entry>
         <oasis:entry colname="col3">SG2</oasis:entry>
         <oasis:entry colname="col4">SPCG</oasis:entry>
         <oasis:entry colname="col5">PSP</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">gNATSGO</oasis:entry>
         <oasis:entry colname="col2">0.000</oasis:entry>
         <oasis:entry colname="col3">0.149</oasis:entry>
         <oasis:entry colname="col4">0.281</oasis:entry>
         <oasis:entry colname="col5">0.261</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SG2</oasis:entry>
         <oasis:entry colname="col2">0.149</oasis:entry>
         <oasis:entry colname="col3">0.000</oasis:entry>
         <oasis:entry colname="col4">0.067</oasis:entry>
         <oasis:entry colname="col5">0.087</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SPCG</oasis:entry>
         <oasis:entry colname="col2">0.281</oasis:entry>
         <oasis:entry colname="col3">0.067</oasis:entry>
         <oasis:entry colname="col4">0.000</oasis:entry>
         <oasis:entry colname="col5">0.111</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PSP</oasis:entry>
         <oasis:entry colname="col2">0.261</oasis:entry>
         <oasis:entry colname="col3">0.087</oasis:entry>
         <oasis:entry colname="col4">0.111</oasis:entry>
         <oasis:entry colname="col5">0.000</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T6"><?xmltex \currentcnt{6}?><label>Table 6</label><caption><p id="d1e3335">The statistical differences between gSSURGO and DSM products for pH <inline-formula><mml:math id="M130" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm. The centre of the map is at <inline-formula><mml:math id="M131" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76<inline-formula><mml:math id="M132" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>30<inline-formula><mml:math id="M133" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>30<inline-formula><mml:math id="M134" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> E, 42<inline-formula><mml:math id="M135" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>52<inline-formula><mml:math id="M136" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>30<inline-formula><mml:math id="M137" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> N.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">DSM product</oasis:entry>
         <oasis:entry colname="col2">MD</oasis:entry>
         <oasis:entry colname="col3">RMSD</oasis:entry>
         <oasis:entry colname="col4">RMSD.Adjusted</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">SG2</oasis:entry>
         <oasis:entry colname="col2">4.436</oasis:entry>
         <oasis:entry colname="col3">6.758</oasis:entry>
         <oasis:entry colname="col4">5.097</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PSP</oasis:entry>
         <oasis:entry colname="col2">3.462</oasis:entry>
         <oasis:entry colname="col3">5.625</oasis:entry>
         <oasis:entry colname="col4">4.433</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
</sec>
<sec id="Ch1.S6">
  <label>6</label><title>Local spatial patterns</title>
      <p id="d1e3487">The interest here is to see how well DSM methods at a relatively fine resolution reproduce known relations at the local geomorphic level, e.g. hillslopes, transects across valleys with multiple terrace levels, and within farms. It has been claimed that DSM at 30 m resolution is sufficient for the management of, or even within, individual farm fields. PSP is the only DSM product which predicts at this resolution.</p>
      <p id="d1e3490"><?xmltex \hack{\newpage}?>We examine this qualitatively first, i.e. by visual inspection, and then quantitatively, mostly following the methods of the regional assessment.</p>
<sec id="Ch1.S6.SS1">
  <label>6.1</label><title>Qualitative assessment</title>
      <p id="d1e3501">Here we use the silt concentration, as it reveals stronger qualitative discrepancies than pH in this test area. Figure <xref ref-type="fig" rid="Ch1.F18"/> shows the silt concentration of the 0–5 cm layer for the (top) gridded SSURGO overlain on the original polygons from which it was derived and (bottom) the disaggregated PSP grid cells in a hilly landscape near Caroline, NY.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F18"><?xmltex \currentcnt{18}?><?xmltex \def\figurename{Figure}?><label>Figure 18</label><caption><p id="d1e3508">Ground overlay from gSSURGO (top) and PSP (bottom) for silt (%) at 0–5 cm, with SSURGO polygons from SoilWeb. The centre of the image is at <inline-formula><mml:math id="M138" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76<inline-formula><mml:math id="M139" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>16<inline-formula><mml:math id="M140" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>25<inline-formula><mml:math id="M141" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> E, 42<inline-formula><mml:math id="M142" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>22<inline-formula><mml:math id="M143" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>53<inline-formula><mml:math id="M144" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> N, with a view azimuth of <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">247</mml:mn><mml:mo>∘</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>. Red colours are low silt, and in this window, alluvial fans (the <monospace>C*</monospace> map units) are shown. Pale grey colours are organic soils (the <monospace>Hk</monospace> and <monospace>Hl</monospace> map units). Light colours are high silt surface soils (the <monospace>L*</monospace>, <monospace>V*</monospace>, <monospace>B*</monospace>, and <monospace>M*</monospace> map units) from thin glacial till developed on shale and mudstone bedrock. gNATSGO polygons have only one value, and PSP disaggregates these, hence the pixelated pattern and somewhat smoothed boundaries. The background is from © Google Earth.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f18.png"/>

        </fig>

      <p id="d1e3618">The gSSURGO product follows the SSURGO lines exactly. Some of the sharp boundary lines do correspond with abrupt transitions on the ground, for example, where the steep hillsides are buried by fan alluvium. But others are not, for example, on the hilltops. These differences are because the predicted silt concentrations are taken from the official series descriptions. PSP follows the map unit lines fairly well but is much finer grained and each 30 m pixel is separately predicted. This results in some smoothing of the abrupt boundary lines from gSSURGO on the hilltops. However,  within some SSURGO map units, PSP predicts some differences in the topsoil silt concentration. These are map units with contrasting components, which PSP attempts to disaggregate according to their correlation with covariates.
For the most part, these do not seem to be related to terrain or land use.</p>
      <p id="d1e3622">For example, Fig. <xref ref-type="fig" rid="Ch1.F19"/> shows the detail of the Holly–Papakting map unit within this PSP window. This map unit has two contrasting soils in similar proportions, i.e. a mineral alluvial soil (Holly series) and an organic soil (Papakting series), and the second has a much lower silt concentration.</p>
      <p id="d1e3627">It is difficult to see the reason for the pattern within this map unit. PSP has placed the component series in their proper proportions but not according to any apparent landscape feature or covariate.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F19"><?xmltex \currentcnt{19}?><?xmltex \def\figurename{Figure}?><label>Figure 19</label><caption><p id="d1e3632">Ground overlay from PSP in the Holly–Papakting map unit and silt (%) at 0–5 cm. The centre of the image is at <inline-formula><mml:math id="M146" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76<inline-formula><mml:math id="M147" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>16<inline-formula><mml:math id="M148" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>03<inline-formula><mml:math id="M149" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> E, 42<inline-formula><mml:math id="M150" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>22<inline-formula><mml:math id="M151" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula>30<inline-formula><mml:math id="M152" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> N. The disaggregation appears to be random and not related to covariates. The background is from © Google Earth.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f19.png"/>

        </fig>

      <p id="d1e3709">Another example from this same area is shown in Sect. S5.</p>
</sec>
<sec id="Ch1.S6.SS2">
  <label>6.2</label><title>Quantitative assessment</title>
      <p id="d1e3720">To see the fine differences at this high resolution, we consider a <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.15</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">0.15</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M154" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> subtile, with the lower-right corner at <inline-formula><mml:math id="M155" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>76.30<inline-formula><mml:math id="M156" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E, 42.45<inline-formula><mml:math id="M157" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, and evaluate pH as in the regional assessment (Sect. <xref ref-type="sec" rid="Ch1.S5"/>).</p>
      <p id="d1e3770">Table <xref ref-type="table" rid="Ch1.T6"/> shows the statistical differences between  gSSURGO (reference) and the DSM products, along with the predictions of pH.
Figure <xref ref-type="fig" rid="Ch1.F20"/> shows the pairwise Pearson correlations between the maps. These results are comparable to those for the full tile at regional resolution because both SG2 and PSP underpredict pH by about 0.35–0.45 pH. The correlations are fairly strong between PSP and gSSURGO and between SG2 and PSP but weak between SG2 and gSSURGO.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F20"><?xmltex \currentcnt{20}?><?xmltex \def\figurename{Figure}?><label>Figure 20</label><caption><p id="d1e3779">Pearson correlations between local products for pH at 0–5 cm. These are moderate but weak for gSSURGO vs. SG2.</p></caption>
          <?xmltex \igopts{width=199.169291pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f20.png"/>

        </fig>

      <p id="d1e3789">Figure <xref ref-type="fig" rid="Ch1.F21"/> shows gSSURGO (reference) along with the predictions of pH by the PSP products. Figure <xref ref-type="fig" rid="Ch1.F22"/> shows these as difference maps.
Clearly, gSSURGO has overall higher values than the other two products and,  despite the fine resolution, has in general large areas of identical values.
The differentiation between map units follows sharp boundaries, even within a single landscape (e.g. the plateau towards the south of the map), and this is likely an artefact of relying on the representative profiles in the official series descriptions for property values. PSP has a finer pattern, due to disaggregation, and shows a smoother local pattern without the sharp  boundaries between map units within a landscape. PSP shows large areas of  low pH. SG2 does not follow the landscape lines well, especially the sharp boundaries between uplands and valleys, and predicts very low pH (<inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">4.5</mml:mn></mml:mrow></mml:math></inline-formula>) on the plateau. It is difficult to recognize local landscape units in this global product.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F21" specific-use="star"><?xmltex \currentcnt{21}?><?xmltex \def\figurename{Figure}?><label>Figure 21</label><caption><p id="d1e3809">Topsoil (0–5 cm) for pH <inline-formula><mml:math id="M159" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10, according to the gSSURGO and DSM products. See the text for a discussion.</p></caption>
          <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f21.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F22" specific-use="star"><?xmltex \currentcnt{22}?><?xmltex \def\figurename{Figure}?><label>Figure 22</label><caption><p id="d1e3827">The difference between gSSURGO and DSM products for pH <inline-formula><mml:math id="M160" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 10 at 0–5 cm. See the text for a discussion.</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f22.png"/>

        </fig>

<sec id="Ch1.S6.SS2.SSS1">
  <label>6.2.1</label><title>Class maps</title>
      <p id="d1e3850">Figure <xref ref-type="fig" rid="Ch1.F23"/> shows the topsoil pH classified into eight histogram-equalized classes. Class limits in this area are approximately 5.30, 5.44, 5.55, 5.61, 5.74, 5.89, and 6.15 pH, with the extreme values of 4.44 and 7.00 pH. SG2 clearly is less detailed than the other two products.
PSP shows a fine pattern that is not closely related to the fine pattern of gSSURGO. As previously noted, gSSURGO is consistently about one pH class higher than the other products.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F23" specific-use="star"><?xmltex \currentcnt{23}?><?xmltex \def\figurename{Figure}?><label>Figure 23</label><caption><p id="d1e3857">The pH classes at 0–5 cm. The coordinates are UTM18N (in metres).</p></caption>
            <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f23.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F24" specific-use="star"><?xmltex \currentcnt{24}?><?xmltex \def\figurename{Figure}?><label>Figure 24</label><caption><p id="d1e3868">Fitted variograms for pH at 0–5 cm. The semivariance units are <inline-formula><mml:math id="M161" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">pH</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">10</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>. Note the short range of gSSURGO and the low sill of SG2 and PSP.</p></caption>
            <?xmltex \igopts{width=503.61378pt}?><graphic xlink:href="https://soil.copernicus.org/articles/8/559/2022/soil-8-559-2022-f24.png"/>

          </fig>

</sec>
<sec id="Ch1.S6.SS2.SSS2">
  <label>6.2.2</label><title>Local spatial autocorrelation</title>
      <p id="d1e3904">The local variograms and their fitted exponential models are shown in Fig. <xref ref-type="fig" rid="Ch1.F24"/>. Table <xref ref-type="table" rid="Ch1.T7"/> shows their statistics. gSSURGO has the shortest effective range and highest sill. PSP has a longer range and low sill due to the harmonization from DSMART that removes some of the overall variability. SG2 has no nugget variance, a low sill, and long range, which is consistent with its regional scale.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T7" specific-use="star"><?xmltex \currentcnt{7}?><label>Table 7</label><caption><p id="d1e3914">Fitted variogram parameters for pH at 0–5 cm. The effective range is in metres, the structural sill is in <inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">pH</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">10</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, and the proportional nugget is <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Product</oasis:entry>
         <oasis:entry colname="col2">Effective range</oasis:entry>
         <oasis:entry colname="col3">Structural sill</oasis:entry>
         <oasis:entry colname="col4">Proportional nugget</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">gSSURGO</oasis:entry>
         <oasis:entry colname="col2">774.00</oasis:entry>
         <oasis:entry colname="col3">13.67</oasis:entry>
         <oasis:entry colname="col4">0.12</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SG2</oasis:entry>
         <oasis:entry colname="col2">2550.00</oasis:entry>
         <oasis:entry colname="col3">7.34</oasis:entry>
         <oasis:entry colname="col4">0.00</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PSP</oasis:entry>
         <oasis:entry colname="col2">1455.00</oasis:entry>
         <oasis:entry colname="col3">6.36</oasis:entry>
         <oasis:entry colname="col4">0.22</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S6.SS2.SSS3">
  <label>6.2.3</label><title>Landscape metrics</title>
      <p id="d1e4041">Table <xref ref-type="table" rid="Ch1.T8"/> shows the statistics from the landscape metrics calculations. The mean fractal dimensions are almost identical.
SG2 is much more aggregated, i.e. the least complex, than gSSURGO or PSP.
PSP has a higher landscape shape and Shannon diversity than the other products. Table <xref ref-type="table" rid="Ch1.T9"/> shows the Jensen–Shannon distance between the co-occurrence vectors of the four products. The co-occurrence patterns of SG2 is somewhat similar to that PSP but quite different from gSSURGO.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T8"><?xmltex \currentcnt{8}?><label>Table 8</label><caption><p id="d1e4051">Landscape metrics statistics (local) for pH at 0–5 cm. <monospace>frac_mn</monospace> is the mean fractal dimension, <monospace>lsi</monospace> is the landscape shape index, <monospace>shdi</monospace> is the Shannon diversity, <monospace>shei</monospace> is the Shannon evenness, and <monospace>ai</monospace> is the aggregation index.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Product</oasis:entry>
         <oasis:entry colname="col2"><monospace>ai</monospace></oasis:entry>
         <oasis:entry colname="col3"><monospace>frac_mn</monospace></oasis:entry>
         <oasis:entry colname="col4"><monospace>lsi</monospace></oasis:entry>
         <oasis:entry colname="col5"><monospace>shdi</monospace></oasis:entry>
         <oasis:entry colname="col6"><monospace>shei</monospace></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">gSSURGO</oasis:entry>
         <oasis:entry colname="col2">73.658</oasis:entry>
         <oasis:entry colname="col3">1.049</oasis:entry>
         <oasis:entry colname="col4">71.395</oasis:entry>
         <oasis:entry colname="col5">1.845</oasis:entry>
         <oasis:entry colname="col6">0.887</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SG2</oasis:entry>
         <oasis:entry colname="col2">87.647</oasis:entry>
         <oasis:entry colname="col3">1.106</oasis:entry>
         <oasis:entry colname="col4">34.978</oasis:entry>
         <oasis:entry colname="col5">1.941</oasis:entry>
         <oasis:entry colname="col6">0.934</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PSP</oasis:entry>
         <oasis:entry colname="col2">56.376</oasis:entry>
         <oasis:entry colname="col3">1.045</oasis:entry>
         <oasis:entry colname="col4">116.476</oasis:entry>
         <oasis:entry colname="col5">2.006</oasis:entry>
         <oasis:entry colname="col6">0.965</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
</sec>
</sec>
<sec id="Ch1.S7" sec-type="conclusions">
  <label>7</label><title>Conclusions</title>
      <p id="d1e4196">The presented methods are well able to expose differences in maps produced by different DSM mapping methods and traditional soil survey. There are also well-documented differences between maps produced by traditional survey methods. For example, <xref ref-type="bibr" rid="bib1.bibx6" id="text.83"/> compared four independent surveys of a 19 km<inline-formula><mml:math id="M164" display="inline"><mml:msup><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:math></inline-formula> area in Cyprus and found that the maps differed considerably in their map unit purity and their proportions of interclass and intraclass variability. Thus, the use of the NRCS products as reference should be seen as a basis for comparison with DSM products and not as the truth. However, it is the best available representation of the soil landscape at the given design scale and legend.</p>
      <p id="d1e4211">Our methods for comparing maps have two limitations due to the decision that they be applicable, using the supplied computer code, to any area within the USA. The first is the use of histogram equalization for the class maps which are then evaluated for the class pattern. For specific areas and properties,  it would be preferable to use established class limits relevant for land use, for example, limits from soil survey interpretation tables. The second is the choice of the exponential model for automatic variogram fitting and the somewhat arbitrary choice of empirical variogram cutoff and bin width. For each area, property and depth interval variograms could be computed and fit according to the analysts' prior knowledge.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T9"><?xmltex \currentcnt{9}?><label>Table 9</label><caption><p id="d1e4217">The Jensen–Shannon distance between co-occurrence vectors (local).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="right"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">gSSURGO</oasis:entry>
         <oasis:entry colname="col3">SG2</oasis:entry>
         <oasis:entry colname="col4">PSP</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">gSSURGO</oasis:entry>
         <oasis:entry colname="col2">0.000</oasis:entry>
         <oasis:entry colname="col3">0.218</oasis:entry>
         <oasis:entry colname="col4">0.168</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SG2</oasis:entry>
         <oasis:entry colname="col2">0.218</oasis:entry>
         <oasis:entry colname="col3">0.000</oasis:entry>
         <oasis:entry colname="col4">0.112</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">PSP</oasis:entry>
         <oasis:entry colname="col2">0.168</oasis:entry>
         <oasis:entry colname="col3">0.112</oasis:entry>
         <oasis:entry colname="col4">0.000</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e4300">A variety of metrics to compare DSM products among themselves and to the reference map are proposed in this paper. This raises several questions about their utility and possible redundancy. Since we only consider one example case here, our conclusions are tentative. Comparing the results here with those in the companion case studies report <xref ref-type="bibr" rid="bib1.bibx52" id="paren.84"/>, we know that these are context dependent, and no general conclusions can be drawn. All the metrics provide useful information based on different summaries of the maps, so none is redundant.</p>
      <p id="d1e4306">A first question is which metrics best reflect the visual differences in patterns, e.g. for the regional patterns of Fig. <xref ref-type="fig" rid="Ch1.F7"/>. For the continuous maps, the whole-map histograms (Fig. <xref ref-type="fig" rid="Ch1.F5"/>) reveal whether the feature–space distribution of the property known from gNATSGO has been distorted by the DSM method. In the example case, PSP produced a strongly bimodal distribution, so its map showed few values near pH 5.8. Much of the patterning in the strongly acid soil region is homogenized towards lower values, and there is a sharper boundary between the strongly and moderately acid areas. By contrast, both SG2 and SPCG reduced the peak modal value of pH 6 and made more predictions towards the two tails of the univariate distribution. This can be seen in the resulting maps by more areas with the colours towards the two ends of the colour ramp. The whole-map variograms (Fig. <xref ref-type="fig" rid="Ch1.F14"/>) reveal the longer-range spatial continuity of the DSM products compared to gNATSGO. This can be seen in the maps as less fine detail and larger areas with similar values.</p>
      <p id="d1e4315">For the classified maps (e.g. Fig. <xref ref-type="fig" rid="Ch1.F15"/>), the large discrepancies between them is due to the slicing from histogram equalization.
Each landscape metric (Table <xref ref-type="table" rid="Ch1.T4"/>) reveals a different aspect of the maps. For example, the aggregation index, <monospace>ai</monospace>, shows that SPCG contains much larger one-class areas, on average, than the other products, and this is clear in the figure. Consistent with this, the landscape shape index, <monospace>lsi</monospace>, shows that SPCG has a simpler overall shape.</p>
      <p id="d1e4328">A second question is which metrics best discriminate the different DSM products. The whole-map histograms and variograms clearly show which products are more similar. In the example case, SG2 and SPCG are quite close, and PSP is substantially different by both these metrics. The Jensen–Shannon distance between co-occurrence vectors (Table <xref ref-type="table" rid="Ch1.T5"/>) clearly shows dissimilarity in the adjacency patterns of classes. In the example case, again SG2 and SPCG are quite close, but here PSP is not too different.
The landscape metrics were inconsistent in this case.</p>
      <p id="d1e4333">It is clear from the best-case example presented in this paper that different DSM methods, with different training points, covariates, and algorithms, can produce quite different predictive soil maps. Thus comparing maps with point-wise evaluation from (almost always biased) field observations gives an incomplete picture of how the different methods represent the soil landscape, which is after all what dictates how the soil is used and managed.</p>
      <p id="d1e4336">The main findings from the example case are as follows:
<list list-type="order"><list-item>
      <p id="d1e4341">Although the regional products (250 m resolution) are well-correlated, the DSM products are biased and underpredict topsoil pH by about 0.38–0.48. They also differ substantially, with an RMSD adjusted for bias of the order of 0.31–0.48 pH. This is based on representative pH values of the mapped STU and not on measured values.</p></list-item><list-item>
      <p id="d1e4345">The DSM products differ substantially among themselves and with the reference product in their local spatial pattern, as revealed by empirical variograms. gNATSGO has a short effective range, but this is smoothed to a range 2 to 3.5 times as long by DSM.</p></list-item><list-item>
      <p id="d1e4349">A classification by histogram equalization reveals major differences in the spatial patterns of the produced class maps, as evaluated both by visual inspection and landscape metrics.</p></list-item><list-item>
      <p id="d1e4353">Despite using USA-specific covariates (parent material and drainage classes) derived from gNATSGO and covariates limited in geographic scope to the USA, the predictive map made by SPCG is not substantially different from that made by SG2, likely due to the similar modelling method.</p></list-item><list-item>
      <p id="d1e4357">The estimates of uncertainty provided by SG2 and PSP are substantially different, both in the width of the uncertainty interval (much narrower in SG2) and in the spatial pattern. This could be in part because SG2 is a global model, whereas PSP is based on local soil surveys and covariates restricted to one tile. The confidence intervals seem unrealistically wide compared to the expert-derived high–low value range provided by gNATSGO.</p></list-item><list-item>
      <p id="d1e4361">At the local level (30 m resolution), the disaggregation provided by PSP does not appear to correspond to landscape positions associated with STU components. PSP obscures the fine-scale details of the local spatial pattern, and SG2 is substantially more general, due to its resolution.</p></list-item></list></p>
      <p id="d1e4365">These results will differ in different soil geographic regions, for different soil properties and for different depth intervals, as shown in the companion case studies report <xref ref-type="bibr" rid="bib1.bibx52" id="paren.85"/>.</p>
      <p id="d1e4371">Why are the results from these DSM examples so poor? Why do they not approximate traditional surveys better? In the following, we present some possible reasons:
<list list-type="order"><list-item>
      <p id="d1e4376">The dominant DSM methods do not explicitly consider spatial continuity or pattern. Experiments have been started with convolutional neural networks and other methods with varying window sizes of covariates.</p></list-item><list-item>
      <p id="d1e4380">Environmental covariates to represent past soil-forming conditions (the time factor) have only been available since the satellite remote sensing age, which is very short in terms of soil formation.</p></list-item><list-item>
      <p id="d1e4384">Environmental covariates to represent soil parent material (e.g. surficial geology) are not available globally, and even for the USA, the proxy of using parent material derived from SSURGO in SPCG was not of sufficient precision to improve the predictive models.</p></list-item><list-item>
      <p id="d1e4388">Point observations were mostly placed by the soil surveyor at supposed typical or representative locations in order to characterize map units and do not capture the full range of variability along toposequences.</p></list-item><list-item>
      <p id="d1e4392">Poor georeference of legacy point observations, many from the pre-GPS era, leads to poor correlation with environmental covariates, hence to poor models and to much noise in the DSM product, which can obscure patterns.</p></list-item><list-item>
      <p id="d1e4396">Traditional soil survey is also a predictive activity. The surveyors use as covariates (i.e. non-soil environmental information related to soil geography) what can be inferred from air photos and direct landscape observations (terrain, vegetation, land use, etc.). These give a more detailed and nuanced view than possible at the resolutions used in practical DSM at regional scale, i.e. 100 m or coarser.</p></list-item></list></p>
      <p id="d1e4399">Despite the discrepancies between DSM products and field survey, DSM can be a valuable tool for soil survey. Because of the expense and difficulty of field survey, in practice, DSM is likely to be the most-used method of making or updating soil maps in areas with no resources or poorly resourced soil survey organizations. For unsurveyed areas, DSM can provide a useful pre-map for planning sampling and field survey, thereby optimizing scarce resources for field work. It has the advantage of being reproducible and objective given a set of training points, relevant environmental covariates, and a machine learning method. Many of its problematic results are due to a set of training points, often with imprecise georeference, that do not properly occupy the covariate feature space and to the lack of covariates to represent some aspects of pedogenesis over time.</p>
      <p id="d1e4402">In the USA (our study area) and in other countries with active soil survey programmes, DSM will be an important but not dominant tool in the overall survey. Soil survey, as practised by the NRCS, uses methods from DSM, applied statistical modelling, and numerical ecology, along with an active and focused field programme. For example, the supervised classification of terrain derivatives and satellite imagery has been successfully used to check the internal consistency of map unit concepts and to assist with the placement of delineations. The aim is to blend the most applicable tools from traditional field survey and applied statistical methods that are supported by pedologic theory and regional land use considerations.</p>
      <p id="d1e4405">Of course, soil survey must be based on a proper examination of the soil itself. There is no substitute for actually examining the soil and landscape for either traditional soil survey or as a reliable basis for DSM.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e4412">Source code as R Markdown documents are freely available, without restriction,  at <ext-link xlink:href="https://doi.org/10.5281/zenodo.5512626" ext-link-type="DOI">10.5281/zenodo.5512626</ext-link> (<xref ref-type="bibr" rid="bib1.bibx5" id="altparen.86"/>;  <uri>https://github.com/ncss-tech/compare-psm</uri>, last access: 25 August 2022). These can be used to (1) import all products to compare, and some others not considered in this study, (2) create ground overlays and corresponding KML files for display in Google Earth, (3) compare SG2 and PSP for <inline-formula><mml:math id="M165" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula><inline-formula><mml:math id="M166" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> tiles, (4) compare SG2 with SPCG and gNATSGO for any rectangular tile, (5) compute landscape metrics and compare them between products for any subtile of these, and (6) evaluate the success of PSP in disaggregating at 30 m resolution.</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e4447">All data used in this study are freely available and can be accessed using the scripts referenced in the code availability section.</p>
  </notes><app-group>
        <supplementary-material position="anchor"><p id="d1e4450">The supplement related to this article is available online at: <inline-supplementary-material xlink:href="https://doi.org/10.5194/soil-8-559-2022-supplement" xlink:title="pdf">https://doi.org/10.5194/soil-8-559-2022-supplement</inline-supplementary-material>.</p></supplementary-material>
        </app-group><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e4459">DGR conceptualized the approach, did most of the writing, wrote the R Markdown documents, and performed the example case study. LP provided DSM expertise and detailed knowledge of SG2. DB and ZL provided USA-specific expertise, in particular about the NRCS and its products and services. All authors collaborated on the motivation, methods, and conclusions.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e4465">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e4472">Publisher’s note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e4478">The contribution of Zamir Libohova was mostly accomplished during his tenure at the USDA-NRCS-National Soil Survey Center (100 Centennial Mall North, Room 152, Lincoln, NE 68508-3866, USA). He is a member of the Consortium GLADSOILMAP from LE STUDIUM Institute for Advanced Research (France). We benefitted from the community comments on the SOIL Discussions version of this paper by Alex McBratney and colleagues at the University of Sydney and from an anonymous colleague at the NRCS.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e4483">This paper was edited by Jacqueline Hannam and reviewed by H. Curtis Monger, Bradley Miller, and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{{Araujo-Carrillo} et~al.(2021){Araujo-Carrillo}, {Varón-Ramírez},
{Jaramillo-Barrios}, {Estupiñan-Casallas}, {Silva-Arero}, {Gómez-Latorre},
and {Martínez-Maldonado}}}?><label>Araujo-Carrillo et al.(2021)Araujo-Carrillo, Varón-Ramírez,
Jaramillo-Barrios, Estupiñan-Casallas, Silva-Arero, Gómez-Latorre,
and Martínez-Maldonado</label><?label Araujo-Carrillo.etal2021?><mixed-citation>Araujo-Carrillo, G. A., Varón-Ramírez, V. M., Jaramillo-Barrios, C. I.,
Estupiñan-Casallas, J. M., Silva-Arero, E. A., Gómez-Latorre, D. A.,
and Martínez-Maldonado, F. E.: IRAKA: The First Colombian Soil
Information System with Digital Soil Mapping Products, CATENA, 196, 104940,
<ext-link xlink:href="https://doi.org/10.1016/j.catena.2020.104940" ext-link-type="DOI">10.1016/j.catena.2020.104940</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Arrouays et~al.(2014)}}?><label>Arrouays et al.(2014)</label><?label Arrouays.etal2014?><mixed-citation>
Arrouays, D., Grundy, M. G., Hartemink, A. E., Hempel, J. W., Heuvelink, G. B.,
Hong, S. Y., Lagacherie, P., Lelyk, G., McBratney, A. B., McKenzie, N. J.,
d. L. Mendonca-Santos, M., Minasny, B., Montanarella, L., Odeh, I. O.,
Sanchez, P. A., Thompson, J. A., and Zhang, G.-L.: GlobalSoilMap: Towards
a Fine-Resolution Global Grid of Soil Properties, Adv. Agron., 125,
93–134, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Arrouays et~al.(2020)Arrouays, McBratney, Bouma, Libohova,
{Richer-de-Forges}, Morgan, Roudier, Poggio, and Mulder}}?><label>Arrouays et al.(2020)Arrouays, McBratney, Bouma, Libohova,
Richer-de-Forges, Morgan, Roudier, Poggio, and Mulder</label><?label Arrouays.etal2020?><mixed-citation>Arrouays, D., McBratney, A., Bouma, J., Libohova, Z., Richer-de-Forges,
A. C., Morgan, C. L., Roudier, P., Poggio, L., and Mulder, V. L.: Impressions
of Digital Soil Maps: The Good, the Not so Good, and Making Them Ever
Better, Geoderma Reg., 20, e00255, <ext-link xlink:href="https://doi.org/10.1016/j.geodrs.2020.e00255" ext-link-type="DOI">10.1016/j.geodrs.2020.e00255</ext-link>,
2020.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Batjes et~al.(2020)Batjes, Ribeiro, and {van
Oostrum}}}?><label>Batjes et al.(2020)Batjes, Ribeiro, and van
Oostrum</label><?label Batjes.etal2020?><mixed-citation>Batjes, N. H., Ribeiro, E., and van Oostrum, A.: Standardised Soil Profile
Data to Support Global Mapping and Modelling (WoSIS Snapshot 2019), Earth
Syst. Sci. Data, 12, 299–320, <ext-link xlink:href="https://doi.org/10.5194/essd-12-299-2020" ext-link-type="DOI">10.5194/essd-12-299-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{Beaudette(2021)}?><label>Beaudette(2021)</label><?label Beaudette2021?><mixed-citation>Beaudette, D.: ncss-tech/compare-psm: PSM Comparison Code v1.0, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.5512626" ext-link-type="DOI">10.5281/zenodo.5512626</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Bie and Beckett(1973)}}?><label>Bie and Beckett(1973)</label><?label BieComparisonfourindependent1973?><mixed-citation>
Bie, S. W. and Beckett, P. H. T.: Comparison of Four Independent Soil Surveys
by Air-Photo Interpretation, Paphos Area (Cyprus), Photogrammetria,
29, 189–202, 1973.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{Bloom(2018)}}?><label>Bloom(2018)</label><?label Bloom2018?><mixed-citation>
Bloom, A. L.: Gorges History: Landscapes and Geology of the Finger Lakes
Region, Paleontological Research Institution, Ithaca, New York, ISBN 978-0-87710-524-4, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{Brus et~al.(2011)Brus, Kempen, and
Heuvelink}}?><label>Brus et al.(2011)Brus, Kempen, and
Heuvelink</label><?label BrusSamplingvalidationdigital2011?><mixed-citation>Brus, D., Kempen, B., and Heuvelink, G.: Sampling for Validation of Digital
Soil Maps, Europ. J. Soil Sci., 62, 394–407,
<ext-link xlink:href="https://doi.org/10.1111/j.1365-2389.2011.01364.x" ext-link-type="DOI">10.1111/j.1365-2389.2011.01364.x</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{California Soil Resource
Lab(2020)}}?><label>California Soil Resource
Lab(2020)</label><?label CaliforniaSoilResourceLabSoiLWeb?><mixed-citation>California Soil Resource Lab: SoilWeb Apps,
<uri>https://casoilresource.lawr.ucdavis.edu/soilweb-apps/</uri> (last access: 18 August 2022), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Chaney et~al.(2019)Chaney, Minasny, Herman, Nauman, Brungard, Morgan,
McBratney, Wood, and Yimam}}?><label>Chaney et al.(2019)Chaney, Minasny, Herman, Nauman, Brungard, Morgan,
McBratney, Wood, and Yimam</label><?label Chaney.etal2019?><mixed-citation>Chaney, N., Minasny, B., Herman, J., Nauman, T., Brungard, C., Morgan, C.,
McBratney, A., Wood, E., and Yimam, Y.: POLARIS Soil Properties: 30-m
Probabilistic Maps of Soil Properties over the Contiguous United States,
Water Resour. Res., 55, 2916–2938, <ext-link xlink:href="https://doi.org/10.1029/2018WR022797" ext-link-type="DOI">10.1029/2018WR022797</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{Cornell University Geospatial Information Repository
(CUGIR)(2022)}}?><label>Cornell University Geospatial Information Repository
(CUGIR)(2022)</label><?label CUGIR2022?><mixed-citation>Cornell University Geospatial Information Repository (CUGIR): Soil
Survey, Tompkins County NY, 1965 (FGDC Metadata),
<uri>https://cugir-data.s3.amazonaws.com/00/74/98/fgdc.html</uri>, last access: 18 August   2022.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{D'Avelo and McLeese(1998)}}?><label>D'Avelo and McLeese(1998)</label><?label DAvelo.McLeese1998?><mixed-citation>D'Avelo, T. P. and McLeese, R. L.: Why Are Those Lines Placed Where They Are?:
An Investigation of Soil Map Recompilation Methods, Soil Survey Horizons,
39, 119–126, <ext-link xlink:href="https://doi.org/10.2136/sh1998.4.0119" ext-link-type="DOI">10.2136/sh1998.4.0119</ext-link>, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{Forbes et~al.(1982)Forbes, Rossiter, and
Van~Wambeke}}?><label>Forbes et al.(1982)Forbes, Rossiter, and
Van Wambeke</label><?label Forbes.etal1982?><mixed-citation>
Forbes, T., Rossiter, D., and Van Wambeke, A.: Guidelines for Evaluating the
Adequacy of Soil Resource Inventories, Cornell University Department of
Agronomy, Ithaca, NY, ISBN 978-0-932865-07-6, 1982.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{Fridland(1974)}}?><label>Fridland(1974)</label><?label Fridland1974?><mixed-citation>Fridland, V. M.: Structure of the Soil Mantle, Geoderma, 12, 35–42,
<ext-link xlink:href="https://doi.org/10.1016/0016-7061(74)90036-6" ext-link-type="DOI">10.1016/0016-7061(74)90036-6</ext-link>, 1974.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{Hengl et~al.(2014)Hengl, de~Jesus, MacMillan, Batjes, Heuvelink,
Ribeiro, {Samuel-Rosa}, Kempen, Leenaars, Walsh, and
Gonzalez}}?><label>Hengl et al.(2014)Hengl, de Jesus, MacMillan, Batjes, Heuvelink,
Ribeiro, Samuel-Rosa, Kempen, Leenaars, Walsh, and
Gonzalez</label><?label Hengl.etal2014?><mixed-citation>Hengl, T., de Jesus, J. M., MacMillan, R. A., Batjes, N. H., Heuvelink, G.
B. M., Ribeiro, E., Samuel-Rosa, A., Kempen, B., Leenaars, J. G. B., Walsh,
M. G., and Gonzalez, M. R.: SoilGrids1km – Global Soil Information
Based on Automated Mapping, PLOS ONE, 9, e105992,
<ext-link xlink:href="https://doi.org/10.1371/journal.pone.0105992" ext-link-type="DOI">10.1371/journal.pone.0105992</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Hengl et~al.(2017)Hengl, de~Jesus, Heuvelink, Gonzalez, Kilibarda,
Blagotić, Shangguan, Wright, Geng, {Bauer-Marschallinger}, Guevara, Vargas,
MacMillan, Batjes, Leenaars, Ribeiro, Wheeler, Mantel, and
Kempen}}?><label>Hengl et al.(2017)Hengl, de Jesus, Heuvelink, Gonzalez, Kilibarda,
Blagotić, Shangguan, Wright, Geng, Bauer-Marschallinger, Guevara, Vargas,
MacMillan, Batjes, Leenaars, Ribeiro, Wheeler, Mantel, and
Kempen</label><?label HenglSoilGrids250m?><mixed-citation>Hengl, T., de Jesus, J. M., Heuvelink, G. B. M., Gonzalez, M. R., Kilibarda,
M., Blagotić, A., Shangguan, W., Wright, M. N., Geng, X.,
Bauer-Marschallinger, B., Guevara, M. A., Vargas, R., MacMillan, R. A.,
Batjes, N. H., Leenaars, J. G. B., Ribeiro, E., Wheeler, I., Mantel, S., and
Kempen, B.: SoilGrids250m: Global Gridded Soil Information Based on
Machine Learning, PLOS ONE, 12, e0169748,
<ext-link xlink:href="https://doi.org/10.1371/journal.pone.0169748" ext-link-type="DOI">10.1371/journal.pone.0169748</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Hesselbarth(2021)}}?><label>Hesselbarth(2021)</label><?label Hesselbarth2021?><mixed-citation>Hesselbarth, M. H.: R-Spatialecology/Landscapemetrics, r-spatialecology,
<uri>https://github.com/r-spatialecology/landscapemetrics</uri> (last access: 18 August 2022), 2021.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Hesselbarth et~al.(2019)Hesselbarth, Sciaini, With, Wiegand, and
Nowosad}}?><label>Hesselbarth et al.(2019)Hesselbarth, Sciaini, With, Wiegand, and
Nowosad</label><?label Hesselbarth.etal2019?><mixed-citation>Hesselbarth, M. H., Sciaini, M., With, K. A., Wiegand, K., and Nowosad, J.:
Landscapemetrics: An Open-Source R Tool to Calculate Landscape Metrics,
Ecography, 42, 1648–1657, <ext-link xlink:href="https://doi.org/10.1111/ecog.04617" ext-link-type="DOI">10.1111/ecog.04617</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{Hole and Campbell(1985)}}?><label>Hole and Campbell(1985)</label><?label HoleSoillandscapeanalysis1985?><mixed-citation>
Hole, F. and Campbell, J.: Soil Landscape Analysis, Rowman &amp; Allanheld,
Totowa, NJ, ISBN 978-0-7102-0492-9, 1985.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Hudson(1992)}}?><label>Hudson(1992)</label><?label Hudsonsoilsurveyparadigmbased1992?><mixed-citation>Hudson, B. D.: The Soil Survey as Paradigm-Based Science, Soil Sci. Soc. Am. J., 56, 836–841,
<ext-link xlink:href="https://doi.org/10.2136/sssaj1992.03615995005600030027x" ext-link-type="DOI">10.2136/sssaj1992.03615995005600030027x</ext-link>, 1992.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{{ISRIC -- World Soil
Information}(2020)}}?><label>ISRIC – World Soil
Information(2020)</label><?label ISRIC-WorldSoilInformation2020?><mixed-citation>ISRIC – World Soil Information: SoilGrids – Global Gridded Soil
Information,  <uri>https://www.isric.org/explore/soilgrids</uri> (last access: 18 August 2022), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{Kupfer(2012)}}?><label>Kupfer(2012)</label><?label Kupfer2012?><mixed-citation>Kupfer, J. A.: Landscape Ecology and Biogeography: Rethinking Landscape
Metrics in a Post-FRAGSTATS Landscape, Prog. Phys.
Geogr.-Earth   Environ., 36, 400–420,
<ext-link xlink:href="https://doi.org/10.1177/0309133312439594" ext-link-type="DOI">10.1177/0309133312439594</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{Lagacherie et~al.(1996)Lagacherie, Andrieux, and
Bouzigues}}?><label>Lagacherie et al.(1996)Lagacherie, Andrieux, and
Bouzigues</label><?label LagacherieFuzzinessuncertaintysoil1996?><mixed-citation>
Lagacherie, P., Andrieux, P., and Bouzigues, R.: Fuzziness and Uncertainty of
Soil Boundaries: From Reality to Coding in GIS, in: Geographic Objects
with Indeterminate Boundaries, edited by: Burrough, P. A., Frank, A. U., and
Salgé, F., GISDATA 2,  275–286, Taylor &amp; Francis, London, ISBN  978-0-7484-0387-5, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Libohova et~al.(2014)Libohova, Wills, and Odgers}}?><label>Libohova et al.(2014)Libohova, Wills, and Odgers</label><?label Libohova.etal2014?><mixed-citation>
Libohova, Z., Wills, S., and Odgers, N. P.: Legacy data quality and uncertainty
estimation for United States GlobalSoilMap products, in:
GlobalSoilMap: Basis of the Global Spatial Soil Information
System, edited by: Arrouays, D., McKenzie, N., Hempel, J., DeForges, A.
C. R., and McBratney, A.,   63–68, Crc Press-Taylor &amp; Francis Group, Boca
Raton,
2014.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{Liu et~al.(2020)Liu, Rossiter, Zhang, and Li}}?><label>Liu et al.(2020)Liu, Rossiter, Zhang, and Li</label><?label Liu.etal2020e?><mixed-citation>Liu, F., Rossiter, D. G., Zhang, G.-L., and Li, D.-C.: A Soil Colour Map of
China, Geoderma, 379, 114556, <ext-link xlink:href="https://doi.org/10.1016/j.geoderma.2020.114556" ext-link-type="DOI">10.1016/j.geoderma.2020.114556</ext-link>,
2020.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{Mallavan et~al.(2010)Mallavan, Minasny, and
McBratney}}?><label>Mallavan et al.(2010)Mallavan, Minasny, and
McBratney</label><?label MallavanHomosoil2010?><mixed-citation>
Mallavan, B., Minasny, B., and McBratney, A.: Homosoil, a Methodology for
Quantitative Extrapolation of Soil Information Across the Globe, in: Digital
Soil Mapping, edited by: Boettinger, J. L., Howell, D. W., Moore, A. C.,
Hartemink, A. E., and Kienast-Brown, S.,   137–150, Springer
Netherlands, Dordrecht, ISBN  978-90-481-8862-8, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{McBratney et~al.(2003)}}?><label>McBratney et al.(2003)</label><?label McBratney.etal2003?><mixed-citation>McBratney, A. B., Mendonça Santos, M. L., and Minasny, B.:
On Digital Soil Mapping, Geoderma, 117, 3–52,
<ext-link xlink:href="https://doi.org/10.1016/S0016-7061(03)00223-4" ext-link-type="DOI">10.1016/S0016-7061(03)00223-4</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{McGarigal et~al.(2012)McGarigal, Cushman, and
Ene}}?><label>McGarigal et al.(2012)McGarigal, Cushman, and
Ene</label><?label McGarigal.etal2012?><mixed-citation>
McGarigal, K., Cushman, S. A., and Ene, E.: FRAGSTATS v4: Spatial Pattern
Analysis Program for Categorical and Continuous Maps, Tech. Rep., University
of Massachusetts, Amherst, MA,
2012.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Meinshausen(2006)}}?><label>Meinshausen(2006)</label><?label MeinshausenQuantileregressionforests2006?><mixed-citation>
Meinshausen, N.: Quantile Regression Forests, J. Mach. Learn.
Res., 7, 983–999, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{Meyer and Pebesma(2020)}}?><label>Meyer and Pebesma(2020)</label><?label Meyer.Pebesma2020?><mixed-citation>Meyer, H. and Pebesma, E.: Predicting into Unknown Space? Estimating the
Area of Applicability of Spatial Prediction Models, arXiv:2005.07939,  <uri>http://arxiv.org/abs/2005.07939</uri> (last access: 18 August 2022), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx31"><?xmltex \def\ref@label{{Meyer and Pebesma(2022)}}?><label>Meyer and Pebesma(2022)</label><?label Meyer.Pebesma2022?><mixed-citation>Meyer, H. and Pebesma, E.: Machine Learning-Based Global Maps of Ecological
Variables and the Challenge of Assessing Them, Nat. Commun., 13,
2208, <ext-link xlink:href="https://doi.org/10.1038/s41467-022-29838-9" ext-link-type="DOI">10.1038/s41467-022-29838-9</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx32"><?xmltex \def\ref@label{{Minasny and McBratney(2016)}}?><label>Minasny and McBratney(2016)</label><?label MinasnyDigitalsoilmapping2016?><mixed-citation>Minasny, B. and McBratney, A. B.: Digital Soil Mapping: A Brief History and
Some Lessons, Geoderma, 264,  301–311,
<ext-link xlink:href="https://doi.org/10.1016/j.geoderma.2015.07.017" ext-link-type="DOI">10.1016/j.geoderma.2015.07.017</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx33"><?xmltex \def\ref@label{{{Moreira de Sousa} et~al.(2019){Moreira de Sousa}, Poggio, and
Kempen}}?><label>Moreira de Sousa et al.(2019)Moreira de Sousa, Poggio, and
Kempen</label><?label MoreiradeSousa.etal2019?><mixed-citation>Moreira de Sousa, L., Poggio, L., and Kempen, B.: Comparison of FOSS4G
Supported Equal-Area Projections Using Discrete Distortion Indicatrices,
ISPRS Int.   Geo-Inf., 8, 351,
<ext-link xlink:href="https://doi.org/10.3390/ijgi8080351" ext-link-type="DOI">10.3390/ijgi8080351</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{{{Natural Resources Conservation Service}(2019)}}?><label>Natural Resources Conservation Service(2019)</label><?label WebSoilSurvey?><mixed-citation>Natural Resources Conservation Service: Web Soil Survey,
<uri>https://websoilsurvey.nrcs.usda.gov/</uri> (last access: 18 August 2022), 2019.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{{Natural Resources Conservation Service(2022)}}?><label>Natural Resources Conservation Service(2022)</label><?label NASIS?><mixed-citation>Natural Resources Conservation Service: National Soil Information System
(NASIS),
<uri>https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/survey/tools/?cid=nrcs142p2_053552</uri>, last access: 18 August 2022.</mixed-citation></ref>
      <ref id="bib1.bibx36"><?xmltex \def\ref@label{{New York State Geological
Survey(1970)}}?><label>New York State Geological
Survey(1970)</label><?label NewYorkStateGeologicalSurvey1970?><mixed-citation>New York State Geological Survey: Geologic Map of New York, New York
State Geological Survey, Albany, NY,
<uri>http://www.nysm.nysed.gov/research-collections/geology/gis</uri>  (last access: 18 August 2022),
1970.</mixed-citation></ref>
      <ref id="bib1.bibx37"><?xmltex \def\ref@label{{New York State Geological
Survey(1986)}}?><label>New York State Geological
Survey(1986)</label><?label NewYorkStateGeologicalSurvey1986?><mixed-citation>New York State Geological Survey: Surficial Geologic Map of New York,
New York State Geological Survey, Albany, NY,
<uri>http://www.nysm.nysed.gov/research-collections/geology/gis</uri>  (last access: 18 August 2022),
1986.</mixed-citation></ref>
      <ref id="bib1.bibx38"><?xmltex \def\ref@label{{Nowosad(2020)}}?><label>Nowosad(2020)</label><?label Nowosad2020?><mixed-citation>Nowosad, J.: sabre: Spatial Association Between Regionalizations,
<uri>https://nowosad.github.io/sabre/</uri>  (last access: 18 August 2022), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx39"><?xmltex \def\ref@label{{Nowosad(2021)}}?><label>Nowosad(2021)</label><?label Nowosad2021?><mixed-citation>Nowosad, J.: Motif: An Open-Source R Tool for Pattern-Based Spatial
Analysis, Landscape Ecol., 36, 29–43, <ext-link xlink:href="https://doi.org/10.1007/s10980-020-01135-0" ext-link-type="DOI">10.1007/s10980-020-01135-0</ext-link>,
2021.</mixed-citation></ref>
      <ref id="bib1.bibx40"><?xmltex \def\ref@label{{Nowosad and Stepinski(2018)}}?><label>Nowosad and Stepinski(2018)</label><?label Nowosad.Stepinski2018?><mixed-citation>Nowosad, J. and Stepinski, T. F.: Spatial Association between Regionalizations
Using the Information-Theoretical V-Measure, Int. J.
Geogr. Inf. Sci., 32, 2386–2401,
<ext-link xlink:href="https://doi.org/10.1080/13658816.2018.1511794" ext-link-type="DOI">10.1080/13658816.2018.1511794</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx41"><?xmltex \def\ref@label{{NRCS Soils(2020{\natexlab{a}})}}?><label>NRCS Soils(2020a)</label><?label NRCSSoils2020a?><mixed-citation>NRCS Soils: Soils,  <uri>https://nrcs.app.box.com/v/soils</uri>  (last access: 18 August 2022),
2020a.</mixed-citation></ref>
      <ref id="bib1.bibx42"><?xmltex \def\ref@label{{NRCS Soils(2020{\natexlab{b}})}}?><label>NRCS Soils(2020b)</label><?label NRCSSoils2020b?><mixed-citation>NRCS Soils: Official Soil Series Descriptions,
<uri>https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/survey/class/?cid=nrcs142p2_053587</uri> (last access: 18 August 2022),
2020b.</mixed-citation></ref>
      <ref id="bib1.bibx43"><?xmltex \def\ref@label{{NRCS Soils(2022{\natexlab{a}})}}?><label>NRCS Soils(2022a)</label><?label NRCSSoils2022g?><mixed-citation>NRCS Soils: Description of Gridded Soil Survey Geographic (gSSURGO)
Database,
<uri>https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/home/?cid=nrcs142p2_053628</uri> (last access: 18 August 2022),
2022a.</mixed-citation></ref>
      <ref id="bib1.bibx44"><?xmltex \def\ref@label{{NRCS Soils(2022{\natexlab{b}})}}?><label>NRCS Soils(2022b)</label><?label NRCSSoils2022n?><mixed-citation>NRCS Soils: Gridded National Soil Survey Geographic Database
(gNATSGO),
<uri>https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/survey/geo/?cid=nrcseprd1464625</uri> (last access: 18 August 2022),
2022b.</mixed-citation></ref>
      <ref id="bib1.bibx45"><?xmltex \def\ref@label{{Odgers et~al.(2014)Odgers, McBratney, Minasny, Sun, and
Clifford}}?><label>Odgers et al.(2014)Odgers, McBratney, Minasny, Sun, and
Clifford</label><?label OdgersDSMART2014?><mixed-citation>
Odgers, N. P., McBratney, A. B., Minasny, B., Sun, W., and Clifford, D.:
DSMART: An Algorithm to Spatially Disaggregate Soil Map Units, in:
GlobalSoilMap: Basis of the Global Spatial Soil Information System,
edited by: Arrouays, D., McKenzie, N., Hempel, J., DeForges, A. C. R., and
McBratney, A.,   261–266, CRC Press-Taylor &amp; Francis Group, Boca
Raton,  CRC Press, ISBN 978-1-138-00119-0, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx46"><?xmltex \def\ref@label{{Pebesma(2004)}}?><label>Pebesma(2004)</label><?label PebesmaMultivariablegeostatisticsgstat2004?><mixed-citation>Pebesma, E. J.: Multivariable Geostatistics in S: The Gstat Package,
Comput. Geosci., 30, 683–691, <ext-link xlink:href="https://doi.org/10.1016/j.cageo.2004.03.012" ext-link-type="DOI">10.1016/j.cageo.2004.03.012</ext-link>,
2004.</mixed-citation></ref>
      <ref id="bib1.bibx47"><?xmltex \def\ref@label{{Pindral et~al.(2020)Pindral, Kot, Hulisz, and
Charzyński}}?><label>Pindral et al.(2020)Pindral, Kot, Hulisz, and
Charzyński</label><?label Pindral.etal2020?><mixed-citation>Pindral, S., Kot, R., Hulisz, P., and Charzyński, P.: Landscape Metrics as a
Tool for Analysis of Urban Pedodiversity, Land Degrad. Dev.,
31, 2281–2294, <ext-link xlink:href="https://doi.org/10.1002/ldr.3601" ext-link-type="DOI">10.1002/ldr.3601</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx48"><?xmltex \def\ref@label{{Poggio et~al.(2021)Poggio, {de Sousa}, Batjes, Heuvelink, Kempen,
Ribeiro, and Rossiter}}?><label>Poggio et al.(2021)Poggio, de Sousa, Batjes, Heuvelink, Kempen,
Ribeiro, and Rossiter</label><?label Poggio.etal2021a?><mixed-citation>Poggio, L., de Sousa, L. M., Batjes, N. H., Heuvelink, G. B. M., Kempen, B.,
Ribeiro, E., and Rossiter, D.: SoilGrids 2.0: Producing Soil Information
for the Globe with Quantified Spatial Uncertainty, SOIL, 7, 217–240,
<ext-link xlink:href="https://doi.org/10.5194/soil-7-217-2021" ext-link-type="DOI">10.5194/soil-7-217-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx49"><?xmltex \def\ref@label{{{R Studio}(2020)}}?><label>R Studio(2020)</label><?label RStudio2020?><mixed-citation>R Studio: R Markdown, <uri>https://rmarkdown.rstudio.com/</uri>  (last access: 18 August 2022),
2020.</mixed-citation></ref>
      <ref id="bib1.bibx50"><?xmltex \def\ref@label{{Ramcharan et~al.(2018)Ramcharan, Hengl, Nauman, Brungard, Waltman,
Wills, and Thompson}}?><label>Ramcharan et al.(2018)Ramcharan, Hengl, Nauman, Brungard, Waltman,
Wills, and Thompson</label><?label Ramcharan.etal2018?><mixed-citation>Ramcharan, A., Hengl, T., Nauman, T., Brungard, C., Waltman, S., Wills, S., and
Thompson, J.: Soil Property and Class Maps of the Conterminous United
States at 100-Meter Spatial Resolution, Soil Sci. Soc. Am.
J., 82, 186–201, <ext-link xlink:href="https://doi.org/10.2136/sssaj2017.04.0122" ext-link-type="DOI">10.2136/sssaj2017.04.0122</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx51"><?xmltex \def\ref@label{{Reddy et~al.(2021)Reddy, Chakraborty, Roy, Singh, Minasny, McBratney,
Biswas, and Das}}?><label>Reddy et al.(2021)Reddy, Chakraborty, Roy, Singh, Minasny, McBratney,
Biswas, and Das</label><?label Reddy.etal2021?><mixed-citation>Reddy, N. N., Chakraborty, P., Roy, S., Singh, K., Minasny, B., McBratney,
A. B., Biswas, A., and Das, B. S.: Legacy Data-Based National-Scale Digital
Mapping of Key Soil Properties in India, Geoderma, 381, 114684,
<ext-link xlink:href="https://doi.org/10.1016/j.geoderma.2020.114684" ext-link-type="DOI">10.1016/j.geoderma.2020.114684</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx52"><?xmltex \def\ref@label{{Rossiter et~al.(2021)Rossiter, Poggio, Beaudette, and
Libohova}}?><label>Rossiter et al.(2021)Rossiter, Poggio, Beaudette, and
Libohova</label><?label Rossiter.etalPSMCases?><mixed-citation>Rossiter, D. G., Poggio, L., Beaudette, D., and Libohova, Z.: How Well Does
Predictive Soil Mapping Represent Soil Geography? An Investigation
from the USA, Case Studies, ISRIC Report 2016-004, ISRIC-World
Soil Information, ISRIC-World Soil Information, ISRIC-World Soil Information,
<ext-link xlink:href="https://doi.org/10.17027/isric-wdcsoils.20160004" ext-link-type="DOI">10.17027/isric-wdcsoils.20160004</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx53"><?xmltex \def\ref@label{{Schoeneberger et~al.(2012)Schoeneberger, Wysocki, Benham, and {Soil
Survey Staff}}}?><label>Schoeneberger et al.(2012)Schoeneberger, Wysocki, Benham, and Soil
Survey Staff</label><?label Schoeneberger.etal2012?><mixed-citation>Schoeneberger, P. J., Wysocki, D. A., Benham, E. C., and Soil Survey Staff:
Field Book for Describing and Sampling Soils, USDA Natural Resources
Conservation Service, Lincoln, NE, 3.0 Edn.,
<uri>https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/research/guide/?cid=nrcs142p2_054184</uri> (last access: 22 August 2022), 2012.</mixed-citation></ref>
      <ref id="bib1.bibx54"><?xmltex \def\ref@label{{{Science Committee}(2012)}}?><label>Science Committee(2012)</label><?label ScienceCommittee2012?><mixed-citation>Science Committee: Specifications: Tiered GlobalSoilMap.Net Products;
Release 2.3, Tech. Rep., GlobalSoilMap.net,
<uri>http://www.ozdsm.com.au/resources/GlobalSoilMap%20specs%20version%202point3.pdf</uri> (last access: 18 August 2022),
2012.</mixed-citation></ref>
      <ref id="bib1.bibx55"><?xmltex \def\ref@label{{Scull et~al.(2003)Scull, Franklin, Chadwick, and
McArthur}}?><label>Scull et al.(2003)Scull, Franklin, Chadwick, and
McArthur</label><?label Scull.etal2003?><mixed-citation>Scull, P., Franklin, J., Chadwick, O., and McArthur, D.: Predictive Soil
Mapping: A Review, Prog. Phys. Geogr., 27, 171–197,
<ext-link xlink:href="https://doi.org/10.1191/0309133303pp366ra" ext-link-type="DOI">10.1191/0309133303pp366ra</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx56"><?xmltex \def\ref@label{{Soil Survey Division
Staff(2014)}}?><label>Soil Survey Division
Staff(2014)</label><?label SoilSurveyStaffKeysSoilTaxonomy2014?><mixed-citation>Soil Survey Division Staff: Keys to Soil Taxonomy, US Government
Printing Office, Washington, DC, 12th Edn.,
<uri>https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/survey/class/</uri> (last access: 18 August 2022),
2014.</mixed-citation></ref>
      <ref id="bib1.bibx57"><?xmltex \def\ref@label{{Soil Survey Division Staff(2017)}}?><label>Soil Survey Division Staff(2017)</label><?label SSM2017?><mixed-citation>Soil Survey Division Staff: Soil Survey Manual, no. 18 in USDA
Handbook, Government Printing Office, Washington, DC,
<uri>http://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/planners/?cid=nrcs142p2_054262</uri> (last access: 18 August 2022),
2017.</mixed-citation></ref>
      <ref id="bib1.bibx58"><?xmltex \def\ref@label{{Szatmári and
Pásztor(2018)}}?><label>Szatmári and
Pásztor(2018)</label><?label SzatmariComparisonvariousuncertainty2018?><mixed-citation>Szatmári, G. and Pásztor, L.: Comparison of Various Uncertainty Modelling
Approaches Based on Geostatistics and Machine Learning Algorithms, Geoderma, 337, 1329–1340,
<ext-link xlink:href="https://doi.org/10.1016/j.geoderma.2018.09.008" ext-link-type="DOI">10.1016/j.geoderma.2018.09.008</ext-link>, 2018.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx59"><?xmltex \def\ref@label{{Taghizadeh-Mehrjardi et~al.(2020){Taghizadeh-Mehrjardi},
Mahdianpari, Mohammadimanesh, Behrens, Toomanian, Scholten, and
Schmidt}}?><label>Taghizadeh-Mehrjardi et al.(2020)Taghizadeh-Mehrjardi,
Mahdianpari, Mohammadimanesh, Behrens, Toomanian, Scholten, and
Schmidt</label><?label Taghizadeh-Mehrjardi.etal2020?><mixed-citation>Taghizadeh-Mehrjardi, R., Mahdianpari, M., Mohammadimanesh, F., Behrens, T.,
Toomanian, N., Scholten, T., and Schmidt, K.: Multi-Task Convolutional Neural
Networks Outperformed Random Forest for Mapping Soil Particle Size Fractions
in Central Iran, Geoderma, 376, 114552,
<ext-link xlink:href="https://doi.org/10.1016/j.geoderma.2020.114552" ext-link-type="DOI">10.1016/j.geoderma.2020.114552</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx60"><?xmltex \def\ref@label{{Thompson et~al.(2020)Thompson, {Kienast-Brown}, D'Avello, Philippe,
and Brungard}}?><label>Thompson et al.(2020)Thompson, Kienast-Brown, D'Avello, Philippe,
and Brungard</label><?label Thompson.etal2020a?><mixed-citation>Thompson, J. A., Kienast-Brown, S., D'Avello, T., Philippe, J., and Brungard,
C.: Soils2026 and Digital Soil Mapping – A Foundation for the Future of
Soils Information in the United States, Geoderma Reg., 22, e00294,
<ext-link xlink:href="https://doi.org/10.1016/j.geodrs.2020.e00294" ext-link-type="DOI">10.1016/j.geodrs.2020.e00294</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx61"><?xmltex \def\ref@label{{United States Department of Agriculture, Natural Resources
Conservation Service(2022)}}?><label>United States Department of Agriculture, Natural Resources
Conservation Service(2022)</label><?label SSH?><mixed-citation>United States Department of Agriculture, Natural Resources Conservation
Service: National Soil Survey Handbook, United States Department of
Agriculture, Natural Resources Conservation Service, Washington, DC,
<uri>https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/home/?cid=nrcs142p2_054242</uri>, last access: 22 August
2022.</mixed-citation></ref>
      <ref id="bib1.bibx62"><?xmltex \def\ref@label{{Uuemaa et~al.(2013)Uuemaa, Mander, and Marja}}?><label>Uuemaa et al.(2013)Uuemaa, Mander, and Marja</label><?label Uuemaa.etal2013?><mixed-citation>Uuemaa, E., Mander, U., and Marja, R.: Trends in the Use of Landscape Spatial
Metrics as Landscape Indicators: A Review, Ecol. Indic., 28,
100–106, <ext-link xlink:href="https://doi.org/10.1016/j.ecolind.2012.07.018" ext-link-type="DOI">10.1016/j.ecolind.2012.07.018</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx63"><?xmltex \def\ref@label{{Vink(1975)}}?><label>Vink(1975)</label><?label Vink1975?><mixed-citation>
Vink, A.: Land Use in Advancing Agriculture, no. 1 in Advanced Series in
Agricultural Sciences, Springer-Verlag, New York, ISBN 978-0-387-07091-9, 1975.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>How well does digital soil mapping represent soil geography?  An investigation from the USA</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Araujo-Carrillo et al.(2021)Araujo-Carrillo, Varón-Ramírez,
Jaramillo-Barrios, Estupiñan-Casallas, Silva-Arero, Gómez-Latorre,
and Martínez-Maldonado</label><mixed-citation>
Araujo-Carrillo, G. A., Varón-Ramírez, V. M., Jaramillo-Barrios, C. I.,
Estupiñan-Casallas, J. M., Silva-Arero, E. A., Gómez-Latorre, D. A.,
and Martínez-Maldonado, F. E.: IRAKA: The First Colombian Soil
Information System with Digital Soil Mapping Products, CATENA, 196, 104940,
<a href="https://doi.org/10.1016/j.catena.2020.104940" target="_blank">https://doi.org/10.1016/j.catena.2020.104940</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Arrouays et al.(2014)</label><mixed-citation>
Arrouays, D., Grundy, M. G., Hartemink, A. E., Hempel, J. W., Heuvelink, G. B.,
Hong, S. Y., Lagacherie, P., Lelyk, G., McBratney, A. B., McKenzie, N. J.,
d. L. Mendonca-Santos, M., Minasny, B., Montanarella, L., Odeh, I. O.,
Sanchez, P. A., Thompson, J. A., and Zhang, G.-L.: GlobalSoilMap: Towards
a Fine-Resolution Global Grid of Soil Properties, Adv. Agron., 125,
93–134, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Arrouays et al.(2020)Arrouays, McBratney, Bouma, Libohova,
Richer-de-Forges, Morgan, Roudier, Poggio, and Mulder</label><mixed-citation>
Arrouays, D., McBratney, A., Bouma, J., Libohova, Z., Richer-de-Forges,
A. C., Morgan, C. L., Roudier, P., Poggio, L., and Mulder, V. L.: Impressions
of Digital Soil Maps: The Good, the Not so Good, and Making Them Ever
Better, Geoderma Reg., 20, e00255, <a href="https://doi.org/10.1016/j.geodrs.2020.e00255" target="_blank">https://doi.org/10.1016/j.geodrs.2020.e00255</a>,
2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Batjes et al.(2020)Batjes, Ribeiro, and van
Oostrum</label><mixed-citation>
Batjes, N. H., Ribeiro, E., and van Oostrum, A.: Standardised Soil Profile
Data to Support Global Mapping and Modelling (WoSIS Snapshot 2019), Earth
Syst. Sci. Data, 12, 299–320, <a href="https://doi.org/10.5194/essd-12-299-2020" target="_blank">https://doi.org/10.5194/essd-12-299-2020</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Beaudette(2021)</label><mixed-citation>
Beaudette, D.: ncss-tech/compare-psm: PSM Comparison Code v1.0, Zenodo [code], <a href="https://doi.org/10.5281/zenodo.5512626" target="_blank">https://doi.org/10.5281/zenodo.5512626</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Bie and Beckett(1973)</label><mixed-citation>
Bie, S. W. and Beckett, P. H. T.: Comparison of Four Independent Soil Surveys
by Air-Photo Interpretation, Paphos Area (Cyprus), Photogrammetria,
29, 189–202, 1973.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Bloom(2018)</label><mixed-citation>
Bloom, A. L.: Gorges History: Landscapes and Geology of the Finger Lakes
Region, Paleontological Research Institution, Ithaca, New York, ISBN 978-0-87710-524-4, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Brus et al.(2011)Brus, Kempen, and
Heuvelink</label><mixed-citation>
Brus, D., Kempen, B., and Heuvelink, G.: Sampling for Validation of Digital
Soil Maps, Europ. J. Soil Sci., 62, 394–407,
<a href="https://doi.org/10.1111/j.1365-2389.2011.01364.x" target="_blank">https://doi.org/10.1111/j.1365-2389.2011.01364.x</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>California Soil Resource
Lab(2020)</label><mixed-citation>
California Soil Resource Lab: SoilWeb Apps,
<a href="https://casoilresource.lawr.ucdavis.edu/soilweb-apps/" target="_blank"/> (last access: 18 August 2022), 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Chaney et al.(2019)Chaney, Minasny, Herman, Nauman, Brungard, Morgan,
McBratney, Wood, and Yimam</label><mixed-citation>
Chaney, N., Minasny, B., Herman, J., Nauman, T., Brungard, C., Morgan, C.,
McBratney, A., Wood, E., and Yimam, Y.: POLARIS Soil Properties: 30-m
Probabilistic Maps of Soil Properties over the Contiguous United States,
Water Resour. Res., 55, 2916–2938, <a href="https://doi.org/10.1029/2018WR022797" target="_blank">https://doi.org/10.1029/2018WR022797</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Cornell University Geospatial Information Repository
(CUGIR)(2022)</label><mixed-citation>
Cornell University Geospatial Information Repository (CUGIR): Soil
Survey, Tompkins County NY, 1965 (FGDC Metadata),
<a href="https://cugir-data.s3.amazonaws.com/00/74/98/fgdc.html" target="_blank"/>, last access: 18 August   2022.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>D'Avelo and McLeese(1998)</label><mixed-citation>
D'Avelo, T. P. and McLeese, R. L.: Why Are Those Lines Placed Where They Are?:
An Investigation of Soil Map Recompilation Methods, Soil Survey Horizons,
39, 119–126, <a href="https://doi.org/10.2136/sh1998.4.0119" target="_blank">https://doi.org/10.2136/sh1998.4.0119</a>, 1998.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Forbes et al.(1982)Forbes, Rossiter, and
Van Wambeke</label><mixed-citation>
Forbes, T., Rossiter, D., and Van Wambeke, A.: Guidelines for Evaluating the
Adequacy of Soil Resource Inventories, Cornell University Department of
Agronomy, Ithaca, NY, ISBN 978-0-932865-07-6, 1982.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Fridland(1974)</label><mixed-citation>
Fridland, V. M.: Structure of the Soil Mantle, Geoderma, 12, 35–42,
<a href="https://doi.org/10.1016/0016-7061(74)90036-6" target="_blank">https://doi.org/10.1016/0016-7061(74)90036-6</a>, 1974.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Hengl et al.(2014)Hengl, de Jesus, MacMillan, Batjes, Heuvelink,
Ribeiro, Samuel-Rosa, Kempen, Leenaars, Walsh, and
Gonzalez</label><mixed-citation>
Hengl, T., de Jesus, J. M., MacMillan, R. A., Batjes, N. H., Heuvelink, G.
B. M., Ribeiro, E., Samuel-Rosa, A., Kempen, B., Leenaars, J. G. B., Walsh,
M. G., and Gonzalez, M. R.: SoilGrids1km – Global Soil Information
Based on Automated Mapping, PLOS ONE, 9, e105992,
<a href="https://doi.org/10.1371/journal.pone.0105992" target="_blank">https://doi.org/10.1371/journal.pone.0105992</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Hengl et al.(2017)Hengl, de Jesus, Heuvelink, Gonzalez, Kilibarda,
Blagotić, Shangguan, Wright, Geng, Bauer-Marschallinger, Guevara, Vargas,
MacMillan, Batjes, Leenaars, Ribeiro, Wheeler, Mantel, and
Kempen</label><mixed-citation>
Hengl, T., de Jesus, J. M., Heuvelink, G. B. M., Gonzalez, M. R., Kilibarda,
M., Blagotić, A., Shangguan, W., Wright, M. N., Geng, X.,
Bauer-Marschallinger, B., Guevara, M. A., Vargas, R., MacMillan, R. A.,
Batjes, N. H., Leenaars, J. G. B., Ribeiro, E., Wheeler, I., Mantel, S., and
Kempen, B.: SoilGrids250m: Global Gridded Soil Information Based on
Machine Learning, PLOS ONE, 12, e0169748,
<a href="https://doi.org/10.1371/journal.pone.0169748" target="_blank">https://doi.org/10.1371/journal.pone.0169748</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Hesselbarth(2021)</label><mixed-citation>
Hesselbarth, M. H.: R-Spatialecology/Landscapemetrics, r-spatialecology,
<a href="https://github.com/r-spatialecology/landscapemetrics" target="_blank"/> (last access: 18 August 2022), 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Hesselbarth et al.(2019)Hesselbarth, Sciaini, With, Wiegand, and
Nowosad</label><mixed-citation>
Hesselbarth, M. H., Sciaini, M., With, K. A., Wiegand, K., and Nowosad, J.:
Landscapemetrics: An Open-Source R Tool to Calculate Landscape Metrics,
Ecography, 42, 1648–1657, <a href="https://doi.org/10.1111/ecog.04617" target="_blank">https://doi.org/10.1111/ecog.04617</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Hole and Campbell(1985)</label><mixed-citation>
Hole, F. and Campbell, J.: Soil Landscape Analysis, Rowman &amp; Allanheld,
Totowa, NJ, ISBN 978-0-7102-0492-9, 1985.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Hudson(1992)</label><mixed-citation>
Hudson, B. D.: The Soil Survey as Paradigm-Based Science, Soil Sci. Soc. Am. J., 56, 836–841,
<a href="https://doi.org/10.2136/sssaj1992.03615995005600030027x" target="_blank">https://doi.org/10.2136/sssaj1992.03615995005600030027x</a>, 1992.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>ISRIC – World Soil
Information(2020)</label><mixed-citation>
ISRIC – World Soil Information: SoilGrids – Global Gridded Soil
Information,  <a href="https://www.isric.org/explore/soilgrids" target="_blank"/> (last access: 18 August 2022), 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Kupfer(2012)</label><mixed-citation>
Kupfer, J. A.: Landscape Ecology and Biogeography: Rethinking Landscape
Metrics in a Post-FRAGSTATS Landscape, Prog. Phys.
Geogr.-Earth   Environ., 36, 400–420,
<a href="https://doi.org/10.1177/0309133312439594" target="_blank">https://doi.org/10.1177/0309133312439594</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Lagacherie et al.(1996)Lagacherie, Andrieux, and
Bouzigues</label><mixed-citation>
Lagacherie, P., Andrieux, P., and Bouzigues, R.: Fuzziness and Uncertainty of
Soil Boundaries: From Reality to Coding in GIS, in: Geographic Objects
with Indeterminate Boundaries, edited by: Burrough, P. A., Frank, A. U., and
Salgé, F., GISDATA 2,  275–286, Taylor &amp; Francis, London, ISBN  978-0-7484-0387-5, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Libohova et al.(2014)Libohova, Wills, and Odgers</label><mixed-citation>
Libohova, Z., Wills, S., and Odgers, N. P.: Legacy data quality and uncertainty
estimation for United States GlobalSoilMap products, in:
GlobalSoilMap: Basis of the Global Spatial Soil Information
System, edited by: Arrouays, D., McKenzie, N., Hempel, J., DeForges, A.
C. R., and McBratney, A.,   63–68, Crc Press-Taylor &amp; Francis Group, Boca
Raton,
2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Liu et al.(2020)Liu, Rossiter, Zhang, and Li</label><mixed-citation>
Liu, F., Rossiter, D. G., Zhang, G.-L., and Li, D.-C.: A Soil Colour Map of
China, Geoderma, 379, 114556, <a href="https://doi.org/10.1016/j.geoderma.2020.114556" target="_blank">https://doi.org/10.1016/j.geoderma.2020.114556</a>,
2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Mallavan et al.(2010)Mallavan, Minasny, and
McBratney</label><mixed-citation>
Mallavan, B., Minasny, B., and McBratney, A.: Homosoil, a Methodology for
Quantitative Extrapolation of Soil Information Across the Globe, in: Digital
Soil Mapping, edited by: Boettinger, J. L., Howell, D. W., Moore, A. C.,
Hartemink, A. E., and Kienast-Brown, S.,   137–150, Springer
Netherlands, Dordrecht, ISBN  978-90-481-8862-8, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>McBratney et al.(2003)</label><mixed-citation>
McBratney, A. B., Mendonça Santos, M. L., and Minasny, B.:
On Digital Soil Mapping, Geoderma, 117, 3–52,
<a href="https://doi.org/10.1016/S0016-7061(03)00223-4" target="_blank">https://doi.org/10.1016/S0016-7061(03)00223-4</a>, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>McGarigal et al.(2012)McGarigal, Cushman, and
Ene</label><mixed-citation>
McGarigal, K., Cushman, S. A., and Ene, E.: FRAGSTATS v4: Spatial Pattern
Analysis Program for Categorical and Continuous Maps, Tech. Rep., University
of Massachusetts, Amherst, MA,
2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Meinshausen(2006)</label><mixed-citation>
Meinshausen, N.: Quantile Regression Forests, J. Mach. Learn.
Res., 7, 983–999, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Meyer and Pebesma(2020)</label><mixed-citation>
Meyer, H. and Pebesma, E.: Predicting into Unknown Space? Estimating the
Area of Applicability of Spatial Prediction Models, arXiv:2005.07939,  <a href="http://arxiv.org/abs/2005.07939" target="_blank"/> (last access: 18 August 2022), 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Meyer and Pebesma(2022)</label><mixed-citation>
Meyer, H. and Pebesma, E.: Machine Learning-Based Global Maps of Ecological
Variables and the Challenge of Assessing Them, Nat. Commun., 13,
2208, <a href="https://doi.org/10.1038/s41467-022-29838-9" target="_blank">https://doi.org/10.1038/s41467-022-29838-9</a>, 2022.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Minasny and McBratney(2016)</label><mixed-citation>
Minasny, B. and McBratney, A. B.: Digital Soil Mapping: A Brief History and
Some Lessons, Geoderma, 264,  301–311,
<a href="https://doi.org/10.1016/j.geoderma.2015.07.017" target="_blank">https://doi.org/10.1016/j.geoderma.2015.07.017</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Moreira de Sousa et al.(2019)Moreira de Sousa, Poggio, and
Kempen</label><mixed-citation>
Moreira de Sousa, L., Poggio, L., and Kempen, B.: Comparison of FOSS4G
Supported Equal-Area Projections Using Discrete Distortion Indicatrices,
ISPRS Int.   Geo-Inf., 8, 351,
<a href="https://doi.org/10.3390/ijgi8080351" target="_blank">https://doi.org/10.3390/ijgi8080351</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Natural Resources Conservation Service(2019)</label><mixed-citation>
Natural Resources Conservation Service: Web Soil Survey,
<a href="https://websoilsurvey.nrcs.usda.gov/" target="_blank"/> (last access: 18 August 2022), 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Natural Resources Conservation Service(2022)</label><mixed-citation>
Natural Resources Conservation Service: National Soil Information System
(NASIS),
<a href="https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/survey/tools/?cid=nrcs142p2_053552" target="_blank"/>, last access: 18 August 2022.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>New York State Geological
Survey(1970)</label><mixed-citation>
New York State Geological Survey: Geologic Map of New York, New York
State Geological Survey, Albany, NY,
<a href="http://www.nysm.nysed.gov/research-collections/geology/gis" target="_blank"/>  (last access: 18 August 2022),
1970.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>New York State Geological
Survey(1986)</label><mixed-citation>
New York State Geological Survey: Surficial Geologic Map of New York,
New York State Geological Survey, Albany, NY,
<a href="http://www.nysm.nysed.gov/research-collections/geology/gis" target="_blank"/>  (last access: 18 August 2022),
1986.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Nowosad(2020)</label><mixed-citation>
Nowosad, J.: sabre: Spatial Association Between Regionalizations,
<a href="https://nowosad.github.io/sabre/" target="_blank"/>  (last access: 18 August 2022), 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Nowosad(2021)</label><mixed-citation>
Nowosad, J.: Motif: An Open-Source R Tool for Pattern-Based Spatial
Analysis, Landscape Ecol., 36, 29–43, <a href="https://doi.org/10.1007/s10980-020-01135-0" target="_blank">https://doi.org/10.1007/s10980-020-01135-0</a>,
2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Nowosad and Stepinski(2018)</label><mixed-citation>
Nowosad, J. and Stepinski, T. F.: Spatial Association between Regionalizations
Using the Information-Theoretical V-Measure, Int. J.
Geogr. Inf. Sci., 32, 2386–2401,
<a href="https://doi.org/10.1080/13658816.2018.1511794" target="_blank">https://doi.org/10.1080/13658816.2018.1511794</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>NRCS Soils(2020a)</label><mixed-citation>
NRCS Soils: Soils,  <a href="https://nrcs.app.box.com/v/soils" target="_blank"/>  (last access: 18 August 2022),
2020a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>NRCS Soils(2020b)</label><mixed-citation>
NRCS Soils: Official Soil Series Descriptions,
<a href="https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/survey/class/?cid=nrcs142p2_053587" target="_blank"/> (last access: 18 August 2022),
2020b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>NRCS Soils(2022a)</label><mixed-citation>
NRCS Soils: Description of Gridded Soil Survey Geographic (gSSURGO)
Database,
<a href="https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/home/?cid=nrcs142p2_053628" target="_blank"/> (last access: 18 August 2022),
2022a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>NRCS Soils(2022b)</label><mixed-citation>
NRCS Soils: Gridded National Soil Survey Geographic Database
(gNATSGO),
<a href="https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/survey/geo/?cid=nrcseprd1464625" target="_blank"/> (last access: 18 August 2022),
2022b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Odgers et al.(2014)Odgers, McBratney, Minasny, Sun, and
Clifford</label><mixed-citation>
Odgers, N. P., McBratney, A. B., Minasny, B., Sun, W., and Clifford, D.:
DSMART: An Algorithm to Spatially Disaggregate Soil Map Units, in:
GlobalSoilMap: Basis of the Global Spatial Soil Information System,
edited by: Arrouays, D., McKenzie, N., Hempel, J., DeForges, A. C. R., and
McBratney, A.,   261–266, CRC Press-Taylor &amp; Francis Group, Boca
Raton,  CRC Press, ISBN 978-1-138-00119-0, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Pebesma(2004)</label><mixed-citation>
Pebesma, E. J.: Multivariable Geostatistics in S: The Gstat Package,
Comput. Geosci., 30, 683–691, <a href="https://doi.org/10.1016/j.cageo.2004.03.012" target="_blank">https://doi.org/10.1016/j.cageo.2004.03.012</a>,
2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Pindral et al.(2020)Pindral, Kot, Hulisz, and
Charzyński</label><mixed-citation>
Pindral, S., Kot, R., Hulisz, P., and Charzyński, P.: Landscape Metrics as a
Tool for Analysis of Urban Pedodiversity, Land Degrad. Dev.,
31, 2281–2294, <a href="https://doi.org/10.1002/ldr.3601" target="_blank">https://doi.org/10.1002/ldr.3601</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Poggio et al.(2021)Poggio, de Sousa, Batjes, Heuvelink, Kempen,
Ribeiro, and Rossiter</label><mixed-citation>
Poggio, L., de Sousa, L. M., Batjes, N. H., Heuvelink, G. B. M., Kempen, B.,
Ribeiro, E., and Rossiter, D.: SoilGrids 2.0: Producing Soil Information
for the Globe with Quantified Spatial Uncertainty, SOIL, 7, 217–240,
<a href="https://doi.org/10.5194/soil-7-217-2021" target="_blank">https://doi.org/10.5194/soil-7-217-2021</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>R Studio(2020)</label><mixed-citation>
R Studio: R Markdown, <a href="https://rmarkdown.rstudio.com/" target="_blank"/>  (last access: 18 August 2022),
2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Ramcharan et al.(2018)Ramcharan, Hengl, Nauman, Brungard, Waltman,
Wills, and Thompson</label><mixed-citation>
Ramcharan, A., Hengl, T., Nauman, T., Brungard, C., Waltman, S., Wills, S., and
Thompson, J.: Soil Property and Class Maps of the Conterminous United
States at 100-Meter Spatial Resolution, Soil Sci. Soc. Am.
J., 82, 186–201, <a href="https://doi.org/10.2136/sssaj2017.04.0122" target="_blank">https://doi.org/10.2136/sssaj2017.04.0122</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Reddy et al.(2021)Reddy, Chakraborty, Roy, Singh, Minasny, McBratney,
Biswas, and Das</label><mixed-citation>
Reddy, N. N., Chakraborty, P., Roy, S., Singh, K., Minasny, B., McBratney,
A. B., Biswas, A., and Das, B. S.: Legacy Data-Based National-Scale Digital
Mapping of Key Soil Properties in India, Geoderma, 381, 114684,
<a href="https://doi.org/10.1016/j.geoderma.2020.114684" target="_blank">https://doi.org/10.1016/j.geoderma.2020.114684</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Rossiter et al.(2021)Rossiter, Poggio, Beaudette, and
Libohova</label><mixed-citation>
Rossiter, D. G., Poggio, L., Beaudette, D., and Libohova, Z.: How Well Does
Predictive Soil Mapping Represent Soil Geography? An Investigation
from the USA, Case Studies, ISRIC Report 2016-004, ISRIC-World
Soil Information, ISRIC-World Soil Information, ISRIC-World Soil Information,
<a href="https://doi.org/10.17027/isric-wdcsoils.20160004" target="_blank">https://doi.org/10.17027/isric-wdcsoils.20160004</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Schoeneberger et al.(2012)Schoeneberger, Wysocki, Benham, and Soil
Survey Staff</label><mixed-citation>
Schoeneberger, P. J., Wysocki, D. A., Benham, E. C., and Soil Survey Staff:
Field Book for Describing and Sampling Soils, USDA Natural Resources
Conservation Service, Lincoln, NE, 3.0 Edn.,
<a href="https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/research/guide/?cid=nrcs142p2_054184" target="_blank"/> (last access: 22 August 2022), 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Science Committee(2012)</label><mixed-citation>
Science Committee: Specifications: Tiered GlobalSoilMap.Net Products;
Release 2.3, Tech. Rep., GlobalSoilMap.net,
<a href="http://www.ozdsm.com.au/resources/GlobalSoilMap%20specs%20version%202point3.pdf" target="_blank"/> (last access: 18 August 2022),
2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Scull et al.(2003)Scull, Franklin, Chadwick, and
McArthur</label><mixed-citation>
Scull, P., Franklin, J., Chadwick, O., and McArthur, D.: Predictive Soil
Mapping: A Review, Prog. Phys. Geogr., 27, 171–197,
<a href="https://doi.org/10.1191/0309133303pp366ra" target="_blank">https://doi.org/10.1191/0309133303pp366ra</a>, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Soil Survey Division
Staff(2014)</label><mixed-citation>
Soil Survey Division Staff: Keys to Soil Taxonomy, US Government
Printing Office, Washington, DC, 12th Edn.,
<a href="https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/survey/class/" target="_blank"/> (last access: 18 August 2022),
2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Soil Survey Division Staff(2017)</label><mixed-citation>
Soil Survey Division Staff: Soil Survey Manual, no. 18 in USDA
Handbook, Government Printing Office, Washington, DC,
<a href="http://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/planners/?cid=nrcs142p2_054262" target="_blank"/> (last access: 18 August 2022),
2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Szatmári and
Pásztor(2018)</label><mixed-citation>
Szatmári, G. and Pásztor, L.: Comparison of Various Uncertainty Modelling
Approaches Based on Geostatistics and Machine Learning Algorithms, Geoderma, 337, 1329–1340,
<a href="https://doi.org/10.1016/j.geoderma.2018.09.008" target="_blank">https://doi.org/10.1016/j.geoderma.2018.09.008</a>, 2018.

</mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>Taghizadeh-Mehrjardi et al.(2020)Taghizadeh-Mehrjardi,
Mahdianpari, Mohammadimanesh, Behrens, Toomanian, Scholten, and
Schmidt</label><mixed-citation>
Taghizadeh-Mehrjardi, R., Mahdianpari, M., Mohammadimanesh, F., Behrens, T.,
Toomanian, N., Scholten, T., and Schmidt, K.: Multi-Task Convolutional Neural
Networks Outperformed Random Forest for Mapping Soil Particle Size Fractions
in Central Iran, Geoderma, 376, 114552,
<a href="https://doi.org/10.1016/j.geoderma.2020.114552" target="_blank">https://doi.org/10.1016/j.geoderma.2020.114552</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>Thompson et al.(2020)Thompson, Kienast-Brown, D'Avello, Philippe,
and Brungard</label><mixed-citation>
Thompson, J. A., Kienast-Brown, S., D'Avello, T., Philippe, J., and Brungard,
C.: Soils2026 and Digital Soil Mapping – A Foundation for the Future of
Soils Information in the United States, Geoderma Reg., 22, e00294,
<a href="https://doi.org/10.1016/j.geodrs.2020.e00294" target="_blank">https://doi.org/10.1016/j.geodrs.2020.e00294</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>United States Department of Agriculture, Natural Resources
Conservation Service(2022)</label><mixed-citation>
United States Department of Agriculture, Natural Resources Conservation
Service: National Soil Survey Handbook, United States Department of
Agriculture, Natural Resources Conservation Service, Washington, DC,
<a href="https://www.nrcs.usda.gov/wps/portal/nrcs/detail/soils/home/?cid=nrcs142p2_054242" target="_blank"/>, last access: 22 August
2022.
</mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>Uuemaa et al.(2013)Uuemaa, Mander, and Marja</label><mixed-citation>
Uuemaa, E., Mander, U., and Marja, R.: Trends in the Use of Landscape Spatial
Metrics as Landscape Indicators: A Review, Ecol. Indic., 28,
100–106, <a href="https://doi.org/10.1016/j.ecolind.2012.07.018" target="_blank">https://doi.org/10.1016/j.ecolind.2012.07.018</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>Vink(1975)</label><mixed-citation>
Vink, A.: Land Use in Advancing Agriculture, no. 1 in Advanced Series in
Agricultural Sciences, Springer-Verlag, New York, ISBN 978-0-387-07091-9, 1975.
</mixed-citation></ref-html>--></article>
