$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
1. Experimental materials
C. oleifera 'Dabieshan 1' leaves were collected at three key stages of development: the juvenile stage, at around 2 years of age; the subject progresses to an early fruiting phase at approximately 6 years old, eventually reaching its peak fruiting period when it is about 10 years old. All the C. oleifera plants at the above three developmental stages were planted using container-grown seedlings of the same specification of C. oleifera 'Dabieshan 1'. The physical and chemical properties of the soil in C. oleifera forest land were shown in Table 1.
2. Collection site overview
Experimental materials were collected from Heping Town, located in the heart of Anhui Province. Shucheng County lies within the province's central region, belonging to the peripheral area of the Dabie Mountains. The coordinates are 31°11 ′N, 116°47 ′E, with a mean elevation of 90 m. The soil is classified as yellow-brown soil. The region has a humid subtropical monsoon climatic pattern, which maintains an average annual temperature of 15.6 °C and typically sees around 1,100 mm of rainfall each year, enjoys approximately 1,969 h of sunshine each year, and has a frost-free period lasting 224 days.
3. Sample collection and pretreatment
Samples were selected during the pre-flowering period (late September). For each developmental stage, the C. oleifera plants with similar, consistent, and strong growth were selected. From each plant, the third leaf was removed from top to bottom, which was moderately growing spring shoots from the east, west, south, and north directions of the outer part of the C. oleifera crown. The samples were marked with serial letters and numbers: YAT1, YAS2, and YAT3 for C. oleifera leaves at the juvenile stage, the subject progresses to an early fruiting phase, and the peak fruiting period, respectively. These were collected and stored in liquid nitrogen. Five random sample trees were selected from each stage as one biological replicate, and three biological replicates were set up in all stages. The selected specimens were brought back the same day to the laboratory and preserved at -80 °C.
4. Physiological index determination
The soluble sugar content was determined by the anthrone colorimetric method21. Organic carbon content was determined by the potassium dichromate oxidation - ferrous sulfate titration method22. The soluble protein content was obtained by the Coomassie brilliant blue method23. The total nitrogen concentration wasquantified by the Kjeldahl method22.
5. Isolation of complete RNA was conducted, followed by the construction of a complementary DNA (cDNA) library and subsequent sequencing
Total RNA was extracted from C. oleifera leaves at different stages using the plant RNA extraction kit. Total RNA concentration was determined by a full-spectrum ultraviolet-visible spectrophotometer. The integrity and clarity of RNA were determined by a biological analyzer. After constructing and validating the cDNA library, transcriptomic sequencing was conducted utilizing a sequence-by-sequencing high-throughput sequencing platform. The sequencing procedure produced individual reads, each spanning 100 base pairs24,25,26. The sequence process alignment results underwent statistical analysis using the RSEM software27. Differential expression genes (DEGs) at various developmental stages of C. oleifera were identified using the DESeq differential analysis software28,29.
6. De novo assembly and UniGene annotation
SeqPrep http://github.com/jstjohn/SeqPrep) and Sickle (http://github.com/najoshi/sickle) were used to filter the raw data, including removing low-quality sequences and contaminated adapters. Then the clean reads obtained from each sample were assembled using Trinity software30,31. All the assembled transcripts were subjected to BLASTX alignment against the Swiss-Prot, Clusters of orthologous groups for eukaryotic complete genomes, and KEGG databases in sequential order. A typical Cut-off E-value (E-value <10−5) was set to retrieve their function annotations. The KEGG was applied to analyze related metabolic pathways32.
7. Functional annotation of genes and differentially expressed genes analysis
Gene and isoform abundances were statistically analyzed using the RSEM software27. A cut-off (P < 0.05, FDR < 0.01, |log2FC|>1) was set to identify differentially expressed genes (DEGs). The expected number of fragments per kb of transcript per million fragments mapped (FPKM) is currently the most widely used method for estimating gene expression levels28. The FPKM values of all genes from the three different growth and developmental stages (YAT1, YAS2, and YAT3) were calculated. Hierarchical clustering analysis was conducted to examine the differential gene expression levels. Meanwhile, the KEGG pathway was prepared to identify DEGs significantly enriched in metabolic pathways at Bonferroni-corrected P ≤ 0.05 compared to the whole transcriptome using KOBAS29.
8. Analytical methods
The data were presented using Excel to calculate mean values and standard deviations. Analysis of variance (ANOVA) was used for statistical analysis by SPSS software, and Duncan's multiple range test was used to assess significant differences in nutrient content among the three stages. Then, significant differences between groups were determined using post-hoc multiple comparisons by the Least Significant Difference (LSD) method. Graphs and charts were generated using Origin software to visualize experimental and analytical findings. No statistical significance was denoted by 'ns' (P > 0.05). Statistically significant differences were denoted by '*' (P < 0.05), '**' (P < 0.01), '***' (P < 0.001), and '****' (P < 0.0001), respectively.