Publicly available geospatial, remote-sensing, planning, and aggregated socioeconomic datasets were used in this study. No human participants, animal subjects, clinical materials, or identifiable personal information were involved. Institutional ethics approval was therefore not required. The specific software environments, computational packages (including spatial econometric and high-resolution visualization libraries), dataset Digital Object Identifiers (DOIs), and source URLs for all resources necessary to reproduce this workflow are detailed in the comprehensive Table of Materials.
1. Study area and analytical framework
The analysis was conducted in the Xinjiang Uygur Autonomous Region, northwestern China, with a specific analytical focus on the delineated urban extents of key cities (Figure 1) to accurately capture localized expansion dynamics. This boundary was used consistently for raster clipping, land-use simulation, carbon-stock assessment, and administrative-unit aggregation. The analytical sample size (n) comprised 105 county-level administrative units analyzed across three discrete temporal nodes (2000, 2010, and 2020), with land-use simulations incorporating 50 stochastic replicates to capture spatial uncertainty.
A multi-scale analytical framework integrating grid-level and administrative-unit analyses was employed. At the grid scale, land-use/land-cover data, nighttime light intensity, impervious surface coverage, vegetation indices, topographic variables, and ecological planning constraints were used to identify land urbanization patterns, simulate future land-use scenarios, and evaluate carbon stock distribution. At the administrative-unit scale, socioeconomic, transportation, and environmental variables were incorporated into spatial econometric analyses to assess the direct and spillover effects of land urbanization and ecological planning on carbon-stock dynamics.
The analytical framework comprised five components. First, multi-temporal remote-sensing and land-use data were used to characterize spatiotemporal patterns of land urbanization, including built-up land expansion, development intensity, and land-use restructuring. Second, ecological protection, cropland protection, and low-carbon planning scenarios were developed to simulate future land-use patterns. Third, the InVEST carbon-storage module was used to quantify carbon-stock distribution and change under historical and future land-use conditions. Fourth, global and local spatial autocorrelation analyses were conducted to identify clustering patterns and hotspot regions associated with carbon-stock variation. Fifth, spatial econometric models were used to evaluate the magnitude, direction, and spillover effects of land urbanization, ecological-planning intensity, and environmental and socioeconomic drivers on carbon-stock dynamics.
Figure 1 presents the overall analytical workflow. The workflow links study-area definition, multi-source data integration, land-urbanization identification, ecological-planning scenario development, land-use simulation, carbon-stock assessment, spatial autocorrelation analysis, and spatial econometric modeling. Multi-source spatial data served as the input layer. Land-urbanization identification and historical land-use analysis formed the pattern-characterization layer. Ecological planning scenarios and carbon stock assessment formed the prediction and impact evaluation layer. Spatial autocorrelation and econometric analyses were used to identify spatial mechanisms and support policy interpretation.

Figure 1: Study area, data integration, and technical workflow. (A) Multi-scalar spatial context and precise urban analytical boundaries for key cities in the Xinjiang Uygur Autonomous Region, northwestern China. (B) Integrated technical workflow linking multi-source spatial data, remote sensing-based land urbanization identification, ecological planning scenario design, PLUS land-use simulation, InVEST carbon stock assessment, spatial autocorrelation, hotspot analysis, and spatial econometric modeling. ND = natural development; EP = ecological protection; CP = cropland protection; LC = low-carbon optimization; OLS = ordinary least squares; SAR = spatial autoregressive model; SEM = spatial error model; SDM = spatial Durbin model. Please click here to view a larger version of this figure.
2. Data sources and variable system
A multi-source dataset was compiled to support land-urbanization identification, ecological-planning scenario simulation, carbon-stock assessment, and spatial econometric analysis. The dataset included land-use/land-cover data, impervious-surface fraction, nighttime light intensity, normalized difference vegetation index (NDVI), topographic variables, transportation accessibility indicators, hydrological variables, population density, gross domestic product (GDP) density, planning constraints, climate variables, and carbon pool parameters. Table 1 summarizes the dataset category, variable description, unit, spatial and temporal resolution, source, and analytical application. Key datasets were obtained for 2000, 2010, and 2020. Land-use and major remote-sensing products were harmonized to resolutions ranging from 30 m to 1 km, depending on analytical requirements.
Table 1: Data sources and variable system. The table reports dataset category, variable, description, unit, spatial/temporal resolution, source, and analytical use for land urbanization identification, scenario simulation, InVEST carbon stock accounting, and spatial econometric modeling. NDVI = normalized difference vegetation index; DEM = digital elevation model; LULC = land use/land cover. Please click here to download this Table.
All spatial layers were projected to a common coordinate system, clipped to the study-area boundary shown in Figure 1, and resampled or aggregated to the required grid or administrative-unit scale. Missing values were screened before spatial overlay analysis. Raster datasets used in PLUS and InVEST analyses were aligned at the pixel level, whereas socioeconomic variables were aggregated to administrative units for spatial regression. This preprocessing workflow ensured consistent spatial units and comparable time points across temporal analyses, scenario simulations, carbon-stock calculations, and econometric modeling.
Land-use/land-cover data served as the primary input for historical change detection and future land-use simulation. Original land-use categories were reclassified into six classes: arable land, forest land, grassland, water bodies, built-up land, and unused land. This classification scheme was used consistently for land-use change analysis, scenario development, and InVEST carbon-stock calculations. The built-up land proportion, impervious surface coverage, and nighttime light intensity were used as indicators of land urbanization. Impervious-surface coverage and nighttime-light intensity served as remote-sensing proxies for development intensity and human activity. NDVI was included as an ecological indicator in subsequent spatial econometric analyses.
Topographic and location variables included elevation, slope, distance to major roads, road density, and proximity to major rivers. These variables were used as drivers in land-use simulation and as control variables in spatial econometric models. Together, they represented terrain constraints, transportation accessibility, and hydrological connectivity.
Socioeconomic variables included population density and GDP density, supplemented by road-network density where available. These variables were used to characterize development intensity and human activity pressures associated with changes in carbon stocks. The variable framework integrated indicators of land development, socioeconomic pressure, and ecological context to represent the multiple factors influencing carbon-stock dynamics.
An ecological-planning intensity index was developed using ecological redlines, nature reserves, water-body buffer zones, slope-restriction areas, and other environmentally sensitive regions. This index served both as a constraint layer in future land-use simulations and as an explanatory variable in spatial econometric analyses.
For InVEST carbon-stock assessment, four carbon-pool parameters were compiled for each land-use class: aboveground biomass carbon, belowground biomass carbon, soil organic carbon, and dead organic matter carbon. Parameter values were obtained from published regional studies, InVEST guidance documents, and local land-cover characteristics and were mapped to the unified land-use classification system. Carbon stock refers to the total estimated carbon contained within these four pools, whereas carbon-stock density refers to carbon stock per unit area.
Variables were grouped into three categories. The first category comprised land-urbanization indicators, including the proportion of built-up land, impervious-surface coverage, and nighttime light intensity. The second category included ecological and planning variables, such as NDVI, ecological-planning intensity, elevation, slope, and hydrological proximity. The third category comprised socioeconomic drivers, including population density, GDP density, and road density. These variables were used to evaluate the relationships among land urbanization, ecological planning, and carbon-stock dynamics.
3. Remote sensing-based identification of land urbanization
The spatiotemporal patterns of land urbanization were identified using three complementary indicators: built-up land expansion, impervious-surface coverage, and nighttime light intensity. Built-up land expansion was extracted from sequential land-use maps to delineate the physical extent and progression of urban development from core built-up areas to surrounding regions. Impervious-surface coverage was calculated to quantify development intensity at the grid level, while nighttime light intensity raster data were processed and normalized to represent human activity and functional concentration. These spatial layers were subsequently integrated to construct the historical land-urbanization dataset. The combined indicators were used to identify areas of sustained development intensification, transition zones, and relatively stable regions, and to evaluate the spatial correspondence between physical land development and functional urbanization.
A land-use transition matrix was constructed to quantify the magnitude and direction of land-use conversion among arable land, forest land, grassland, water bodies, construction land, and unused land. Particular attention was given to transitions from ecological and agricultural land classes to construction land. The matrix was used to identify dominant conversion pathways and the principal source land classes contributing to urban expansion.
Landscape-pattern analysis was conducted to evaluate structural changes associated with land urbanization. Metrics included patch density, edge density, landscape shape index, and fragmentation-related indicators. These metrics were calculated for each study period and used to quantify changes in landscape configuration, spatial continuity, and fragmentation associated with the expansion of built-up land.
4. Ecological planning scenario design and land-use simulation
Future land-use patterns were simulated using the Patch-generating Land Use Simulation (PLUS) model under four planning scenarios: natural development (ND), ecological protection (EP), cropland protection (CP), and low-carbon optimization (LC). Scenario assumptions, land-conversion rules, restricted land types, and expected carbon-stock outcomes are summarized in Table 2.
Table 2: Scenario control rules and transition constraints. The table defines land conversion probabilities and spatial constraints for natural development (ND), ecological protection (EP), cropland protection (CP), and low-carbon optimization (LC) scenarios. Notes: Scenario abbreviations strictly correspond to those utilized in the PLUS modeling and Results sections. The scenario settings define policy-oriented conversion rules and restricted land types used in the simulations; all statutory exclusion zones remained non-convertible in the final spatial layers. Please click here to download this Table.
Historical land-use maps and spatial driver variables were integrated into the PLUS model to estimate baseline land-transition probabilities using the cellular automata (CA) module. For the ecological protection (EP), cropland protection (CP), and low-carbon optimization (LC) scenarios, ecological redlines, permanent basic cropland, water-body buffers, and other planning constraints were incorporated as spatial restriction layers to limit land conversion in accordance with the predefined scenario rules. Target land-demand quantities for 2030 were subsequently defined for each scenario, and the CA module was executed using 50 stochastic replicates to generate the final land-use projections.
Land-use simulation was performed using the PLUS model and its cellular automata (CA) framework with multi-type stochastic patch generation. Historical land-use maps and environmental and socioeconomic driver variables were used to estimate land-expansion probabilities for each land-use class. Future land-demand quantities were then specified according to the requirements of each planning scenario, and corresponding land-use maps were generated for the target simulation period.
Model performance was evaluated through historical back-casting prior to future simulation. Earlier land-use maps and associated driver variables were used to simulate a later observed land-use map. Agreement between simulated and observed land-use distributions was assessed using overall accuracy (OA), the Kappa coefficient, and the Figure of Merit (FoM). Validation was conducted at both the overall level and by major land-use class. Historical back-casting from 2010 to 2020 yielded an overall accuracy (OA) of 93.4%, a Kappa coefficient of 0.89, and a Figure of Merit (FoM) of 0.26, indicating a highly reliable capacity for spatial projection in subsequent multi-scenario simulations.
Uncertainty and sensitivity analyses were conducted to evaluate the robustness of simulation outcomes. PLUS sensitivity analyses examined the effects of alternative transition-resistance settings and neighborhood-weight parameters for major land-use classes. InVEST sensitivity analyses evaluated the influence of variation in carbon-pool coefficients across land-use types. Spatial econometric sensitivity analyses compared alternative spatial-weight matrix specifications. These analyses were used to assess whether scenario rankings and the direction of the principal urbanization and ecological-planning effects remained consistent under alternative parameter settings. Specifically, outcome stability was rigorously confirmed: scenario rankings and the negative spatial spillover effects of land urbanization remained invariant when transition-resistance parameters and carbon-pool coefficients were perturbed by ±15%.
5. Carbon stock assessment
Carbon stock was assessed using the InVEST carbon model. Baseline parameters for the four carbon pools (aboveground biomass, belowground biomass, soil organic carbon, and dead organic matter) were assigned to each reclassified land-use type using the biophysical values summarized in Table 3. Historical land-use maps (2000–2020) and PLUS-simulated future land-use raster datasets were then imported into the model and integrated with the corresponding carbon-density parameters. The model was subsequently executed to estimate total regional carbon stock (Tg C), grid-level carbon-stock density (Mg C/ha), and spatial maps of carbon-stock change (ΔC) for the historical and future scenarios.
Reclassified land-use maps were linked to the corresponding baseline carbon-pool parameter values reported in Table 3. The sensitivity of carbon stock estimates to parameter uncertainties was tested by adjusting these baseline values by ±15%, with detailed sensitivity outcomes. The model was used to calculate total carbon stock, carbon-stock density, and carbon-stock change for each land-use class under historical conditions and future planning scenarios.
Three categories of outputs were evaluated. First, total regional carbon stock and temporal trends were calculated to quantify the magnitude and direction of carbon-stock change over time. Second, spatial distributions of carbon stock and carbon-stock change were mapped to identify areas of carbon retention and carbon loss. Third, carbon-stock estimates were compared across planning scenarios to evaluate the relative effects of ecological protection, cropland protection, and low-carbon optimization strategies on carbon-stock conservation.
Table 3: Baseline carbon-pool parameters for different land-use types. The table reports aboveground biomass carbon, belowground biomass carbon, soil organic carbon, dead organic matter carbon, and total carbon density used in the InVEST carbon stock module. Units are Mg C/ha. Notes: Values represent baseline parameters utilized in the InVEST model. Total carbon density equals the sum of the four carbon pools. Sensitivity analyses perturbing these baseline values by ±15% are reported in Table 6. Please click here to download this Table.
6. Spatial autocorrelation and spatial econometric analysis
Spatial autocorrelation analysis was conducted to determine whether carbon stock and carbon-stock change exhibited significant spatial dependence. Global Moran's I was calculated to evaluate the overall degree of spatial clustering in carbon-stock distribution and carbon-stock change across the study area. Local Moran's I was then used to identify local spatial association patterns, including high-high, low-low, high-low, and low-high clusters. Hotspot analysis was performed to identify areas of concentrated carbon loss and carbon-stock retention.
Spatial econometric models were used to examine the relationships between land urbanization, ecological planning, environmental conditions, socioeconomic factors, and carbon-stock dynamics. Carbon-stock density or carbon-stock change served as the dependent variable. Explanatory variables included the land urbanization index, built-up land proportion, ecological planning intensity, normalized difference vegetation index (NDVI), population density, GDP density, road density, elevation, slope, mean annual precipitation, and mean annual temperature.
Ordinary least squares (OLS) regression was used as the baseline model. Residual spatial dependence was evaluated before estimating the spatial autoregressive (SAR), spatial error (SEM), and spatial Durbin (SDM) models. Model performance and coefficient estimates were compared across specifications.
A row-standardized spatial-weight matrix was constructed to represent neighborhood relationships among administrative units. The primary specification was based on spatial contiguity, and robustness analyses compared alternative distance-based and nearest-neighbor spatial-weight matrices where available.
Direct, indirect, and total effects were calculated from the SDM to evaluate local and spatial spillover relationships. Terrain, accessibility, vegetation, climate, and socioeconomic variables were included as control variables to reduce omitted-variable bias. Model coefficients were interpreted as conditional spatial associations rather than definitive causal effects. Variables with p-values exceeding conventional significance thresholds were interpreted as weak or suggestive evidence and were not treated as statistically robust effects.