This modified virus purification method provided an enrichment of virus DNAs useful for identifying two virus species by NGS and bioinformatics. After the homogenate was centrifuged at 40,000 x g for 2.5 h, there was a green pellet at the bottom of the tube and a white pellet along the length. The green pellet was resuspended into one microcentrifuge tube and the white pellet was resuspended into two microcentrifuge tubes. PCR was carried out using standard CaYMV PCR diagnostic primers, and products were detected in the solubilized white pellet and not the green pellet (Figure 1A). A sample of the crude preparation was examined by transmission electron microscopy and we observed bacilliform particles measuring 124-133 nm in length (Figure 1B). This is within the predicted modal length of most badnaviruses. DNA was extracted from the white and green pellets and resuspended separately. In Figure 1C, we loaded 5 µL of DNA extracted from the green and white pellet sample (1.6 µg of DNA for the green fraction and 3.1 µg of DNA for the white fraction) to 0.8% agarose gel electrophoresis and analyzed the DNA following ethidium bromide staining. The green fraction contained low molecular weight DNA whereas the white fraction produced two bands of higher molecular weight DNA, as well as the lower molecular weight DNA (Figure 1C). The gel presented in Figure 1C was run for 40 min at 100 V and the smear in lane 3 suggests that the gel voltage should be lowered to produce clearer bands. These data suggest that the white pellet was enriched for virions. The DNA (0.6 µg/mL) concentration extracted from the white sample was low, but adequate for NGS, which requires a minimum of 10 ng of DNA to proceed. Fragmented DNAs were used to prepare a library for NGS.
In parallel, RNA was extracted from infected canna plants (Figure 1D) for high-throughput RNA-seq. A standard workflow was carried out for library preparation, NGS, creating contigs, and identifying viral genome sequences (Figure 1E).The output results from using DNA and RNA as starting materials were compared.
We obtained 188,626 raw DNA reads by NGS using DNA isolated from crude virus preparation. Reads were assembled into 13,269 contigs and BLASTn was used to search the NCBI dataset of nucleotide sequences (using Viridplantae TaxID: 33090 and Virus TaxID: 10239 as the limiting organisms) (Figure 1E). The NCBI-BLASTn results revealed that 93% of de novo assembled contigs were cellular sequences, 22% were unknown, and 0.3% were virus contigs (Figure 2A). The majority of contigs categorized as cellular sequences were identified as mitochondrial or chloroplast DNA. Within the dataset of virus contigs, 32% of the virus contigs were related to members of Caulimoviridae (that were not Badnavirus sequences) and 58% of these were related to Badnavirus. Of the virus contigs, 29% were highly similar (e < 1 x 10-30) to CaYMV isolate V17 ORF3 gene (EF189148.1), Sugarcane bacilliform virus isolate Batavia D, complete genome (FJ439817.1), and Banana streak CA virus complete genome (KJ013511). Within this population, there were long contigs that resembled two full length genomes.
High-throughput RNA-seq produced 153,488 cleaned individual sequence reads with an average read length of < 500 bp. Contig assembly reduced this to 8,243 contigs. These were submitted to NCBI-BLASTn (using Viridplantae TaxID: 33090 and Virus TaxID: 10239 as the limiting organisms) and the outputs placed 76% of the contigs in a category of plant cellular sequences, 23% were unknown, and 0.1% were categorized as virus contigs (Figure 2B). Closer examination of the population of the 0.1% population of virus contigs determined that 68% of these were assigned to Caulimoviridae (Figure 2B). Three large contigs within this population were identified with high similarity (e < 1 X 10-30) to CaYMV isolate V17 ORF3 gene (EF189148.1), Sugarcane bacilliform virus isolate Batavia D, complete genome (FJ439817.1) and Banana streak CA virus complete genome (KJ013511). Examining the three contigs, we manually joined two of these to produce a full-length virus genome.
We compared the virus genome length contigs produced by DNA and RNA sequencing as a mutual scaffold to confirm the presence of two full-length virus genomes. One full-length virus genome of 6,966 bp was tentatively named Canna yellow mottle associated virus 1 (CaYMAV-1) (Figure 3A). The second genome was 7,385 bp and a variant of CaYMV infecting Alpinia purpurata (CaYMV-Ap01) (Figure 3A).
Finally, PCR primers which were designed to clone ~1,000 bp fragment of each virus, were used to differentially detect both genomes in a population of 227 canna plants representing nine commercial varieties. In many instances individual plants were infected with both viruses. We provide an example of RT-PCR detection of CaYMAV-1 and CaYMV-Ap01 in the 12 plants. Three of these were positive only for CaYMV-Ap01 and nine were positive for both viruses (Figure 3B).

Figure 1: Virus nucleic acid preparations and NGS workflow. (A) Agarose (1.0%) gel electrophoresis of 565 bp PCR fragments of CaYMV genomes. Two PCR products were detected in samples prepared from the white pellet (lanes 1, 2) but not in the green pellet sample (lane 3). Positive control (+) represents a PCR product amplified from infected plant DNA that was isolated using an automated method involving standard paramagnetic cellulose particles. Lane L contains the DNA ladder used as a standard for measuring the size of linear DNA bands in sample lanes. (B) Example of virus particle viewed by transmission electron microscopy in the white pellet recovered by crude fractionation of infected canna leaves. (C) Agarose (0.8%) gel electrophoresis of DNA recovered from the green (lane 1) and white (lane 2) pellets that tested positive by PCR in panel A. The red and yellow dots next to lane 2 identify two high molecular weight DNA bands that occur in the white fraction. (D) Agarose (1%) gel electrophoresis of total RNA recovered by column-based RNA purification. Lane L contains the DNA ladder used as a standard for measuring the size of linear bands in sample lanes. Lane 1-6 contains RNA isolated from infected canna leaves which were pooled to a single sample for ribo-depletion and RNA-seq. (E) Schematic pipeline of nucleic acid isolations, library preparation, sequencing, contig assembly, and virus genome discovery. Please click here to view a larger version of this figure.

Figure 2: Krona charts visualizing the taxonomic categories of contigs. (A) The chart on the left shows the abundance and taxonomic distribution of contigs assembled from the crude virus preparation. The right chart depicts the proportions of virus contigs associated with the Caulimoviridae family, Badnavirus genus, and three closely related species. (B) The panel on the left shows the abundance of contigs derived from RNA-seq based on their taxonomic distribution. On the right is the graph depicting the abundance of contigs within the population of virus contigs associated with the Caulimoviridae family, Badnavirus genus, and three closely related species. Please click here to view a larger version of this figure.

Figure 3. Characterization of CaYMAV-1 and CaYMV-Ap01 genomes. (A) Diagrammatic representation of Canna yellow mottle associate virus 1 (CaYMAV) and Canna yellow mottle virus similar to the genome isolated from Alpinia purpurata (CaYMV-Ap01). Nucleotide positions 1-10 is identified as the start of the genome and contains a tRNAmet anticodon site typical of most badnavirus genomes. The stop and start positions for translation of open reading frame (ORF) 1 and 2 are adjacent. These proteins have unknown functions. ORF3 is a polyprotein containing zinc finger (ZnF), protease (Pro), reverse transcriptase (RT), and RNAse H domains. A 3' poly(A) signal sequence is conserved for both virus genomes. (B) RT-PCR analysis was carried out using RNA isolated from virus infected leaves and primers that detect CaYMAV and CaYMV-Ap01. In the same population of 12 plants, three were infected with CaYMV-Ap01 only, whereas the remaining were infected with both CaYMAV and CaYMV-Ap01. (+) indicates positive control and (-) indicates negative control. This figure is reproduced/modified from Wijayasekara et al.24 with permission. Please click here to view a larger version of this figure.