$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Genome sequencing is a powerful tool to dissect the underlying genetic basis of important plant traits. Most genome-sequencing studies focus on the nuclear genome content, as the majority of genes are located in the nucleus. However, organellar genomes, including the mitochondria (across eukaryotes) and plastids (in plants; the specialized form, the chloroplast, works in photosynthesis) contribute significant genetic information essential to organismal development, stress response, and overall fitness1. Organellar genomes are typically included in total DNA extractions intended for nuclear genome sequencing, although methods to reduce organelle numbers prior to DNA extraction are also employed2. Many studies have used sequencing results from total gDNA extractions to assemble organellar genomes3,4,5,6,7. However, when the target of the study is to focus on organellar genomes, using the total gDNA increases the sequencing costs because many reads are "lost" to the nuclear DNA sequences, particularly in plants with large nuclear genomes. Moreover, due to the duplication and transfer of organellar sequences into the nuclear genome and between organelles, resolving the correct mapping position of sequencing reads to the proper genome is bioinformatically challenging2,8. The purification of organellar genomes from the nuclear genome is one strategy to reduce these problems. Further bioinformatics strategies may be used to separate reads that map to regions of homology between the mitochondria and chloroplasts.
While the organellar genomes from many plant species have been sequenced, little is known about the breadth of organellar genome diversity available in wild populations or in cultivated breeding pools. Organellar genomes are also known to be dynamic molecules that undergo significant structural rearrangement due to recombination between repeat sequences9. Moreover, multiple copies of the organellar genome are contained within each organelle, and multiple organelles are contained within each cell. Not all copies of these genomes are identical, which is known as heteroplasmy. In contrast to the canonical picture of "master circles," there is now growing evidence for a more complex picture of organellar genome structures, including sub-genomic circles, linear chromosomes, linear concatamers, and branched structures10. The assembly of plant organellar genomes is further complicated by their relatively large sizes and substantial inverted and direct repeats.
Traditional protocols for organellar isolation, DNA purification, and subsequent genome sequencing are often cumbersome and require large volumes of tissue input, with several grams to upwards of hundreds of grams of young leaf tissue necessary as a starting point11,12,13,14,15,16,17. This makes organellar genome sequencing inaccessible when tissue is limited. In some situations, seed amounts are limited, such as when it is necessary to sequence on a generational basis or in male sterile lines that have to be maintained via crossing. In these situations, organellar DNA can be purified and then subjected to whole-genome amplification. However, whole-genome amplification can introduce significant sequencing bias, which is a particular problem when assessing structural variation, sub-genomic structures, and heteroplasmy levels18. Recent advances in library preparation for short-read sequencing technologies have overcome low-input barriers to avoid whole-genome amplification. For example, the Illumina Nextera XT library preparation kit allows for as little as 1 ng of DNA to be used as input19. However, standard library preparations for long-read sequencing applications, such as PacBio or Oxford Nanopore sequencing technologies, still require a relatively high amount of input DNA, which can pose a challenge for organellar genome sequencing. Recently, new user-made, long-read sequencing protocols have been developed to reduce the input amounts and to help facilitate genome sequencing in samples where obtaining microgram-quantities of DNA is difficult20,21. However, obtaining high-molecular weight, pure organellar fractions to feed into these library preparations remains a challenge.
We sought to compare and optimize organellar DNA enrichment and isolation methods suitable for NGS without the need of whole-genome amplification. Specifically, our goal was to determine best practices to enrich for high-molecular weight organellar DNA from limited starting materials, such as a subsample of a leaf. This work presents a comparative analysis of methods to enrich for organellar DNA: (1) a modified, traditional differential centrifugation protocol versus (2) a DNA fractionation protocol based on the use of a commercially available DNA CpG-methyl-binding domain protein pulldown approach22 applied to plant tissue23. We recommend best practices for the isolation of organellar DNA from wheat leaf tissue, which may be readily extended to other plants and tissue types.