This protocol was developed and tested using samples collected as part of studies approved by the UCSF Institutional Review Board.

Figure 1: Cerebrospinal fluid (CSF) metagenomic next-generation sequencing (mNGS) sample processing and analytic workflow for pathogen detection. This diagram demonstrates the primary steps in the CSF mNGS wet-lab and bioinformatic pathways. Please click here to view a larger version of this figure.
1. CSF sample handling considerations
NOTE: mNGS results are optimal if CSF library preparation occurs at the time of collection, or if CSF samples are preserved via snap-freezing in liquid nitrogen, followed by long-term storage at -80 °C15. Alternatively, the use of preservative solutions such as DNA/RNA shield at the time of sample collection can allow for storage at room temperature for longer periods of time16,17. This may be preferable in situations when there are concerns about the reliability of the cold chain. In addition, avoidance of unnecessary freeze-thaw cycles can also help preserve nucleic acid integrity, and it may be advisable to pre-aliquot samples at the time of collection if multiple analytic workflows are anticipated for the CSF samples18,19.
2. Extraction of CSF nucleic acid
CAUTION: Perform in an appropriate biosafety level cabinet with gloves and a laboratory coat to reduce the risk of transmission of infectious organisms. Reagents are harmful if swallowed, inhaled, or come into contact with skin. Dispose of used materials as hazardous chemical waste in compliance with local regulations.
- The Quick-DNA/RNA Kit generally works well for samples with a low abundance of nucleic acids.
- Thaw the CSF samples on ice, then transfer 100 µL–1 mL of each sample to a 1.5 mL tube and centrifuge at 16,000 x g at 4 °C for 10 min.
- Carefully pipette off the supernatant, leaving behind the pellet (which is often invisible to the naked eye) and 100 µL of supernatant.
- Add 100 µL of DNA/RNA Shield to the remaining 100 µL of supernatant and pellet, and pipette mix thoroughly.
- Then, add 600 µL of DNA/RNA lysis buffer and pipette mix until the solution is homogeneous.
- Follow the remainder of the DNA/RNA extraction protocol as described by the manufacturer in Supplementary Appendices A, with the following important modifications specific to CSF:
- Perform all centrifugation steps for 1 min at max speed (~21,000 x g) except for the final DNA/RNA wash step, which should run for 3 min at max speed.
- Elute the sample by adding 22 µL (rather than 25 µL) of nuclease-free water directly onto the column matrix, then incubate for 3 min at room temperature. The higher volume generally results in a greater total quantity of nucleic acid eluted from the column.
- After elution, pipette the flow-through back onto the column matrix and centrifuge again for one minute at max speed. Transfer the sample to a new 1.5 mL storage tube, and either proceed immediately to library preparation or store at -80 °C.
3. CSF RNA library preparation
CAUTION: Perform in an appropriate biosafety cabinet, wearing gloves and a laboratory coat. Reagents are harmful if swallowed, inhaled, or come into contact with skin. Dispose of used materials as hazardous chemical waste in compliance with local regulations.
NOTE: Important considerations prior to beginning: rRNA depletion is preferred over polyadenylation enrichment for isolation of mRNA, given that the rRNA depletion does not require additional bead clean steps (which can reduce the yield of usable mRNA), and some neuroinvasive RNA virus transcripts lack polyadenylation. ERCC (External RNA controls consortium) RNA Spike-In can be added as a quality control measure, though it is not strictly necessary. If the RNA was stored at -80 °C following extraction, thaw on ice prior to beginning.
- RNA fragmentation and priming: Pipette mix the following in a polymerase chain reaction (PCR) tube: 3.5 µL RNA sample, 0.5 µL 1:2500 diluted ERCC spike-in, 4 µL first-strand reaction buffer (5x), 1 µL random primers, and 1 µL 1:100 diluted rRNA (Table of Materials). In a thermocycler, set the heated lid to 105 °C and incubate the samples at 75 °C for 2 min, 70 °C for 2 min, 65 °C for 2 min, 60 °C for 2 min, 55 °C for 2 min, 37 °C for 5 min, and 25 °C for 5 min.
- First-strand cDNA synthesis: Pipette mix the fragmented and primed RNA (10 µL) with 8 µL nuclease-free water and 2 µL first-strand synthesis enzyme mix. In a thermocycler, set the heated lid to 105 °C and incubate the samples at 25 °C for 10 min, 42 °C for 15 min, and 70 °C for 15 min.
- Second-strand cDNA synthesis: Pipette mix the first-strand synthesized DNA (20 µL) with 8 µL second-strand synthesis reaction buffer, 4 µL second-strand synthesis enzyme mix, and 48 µL nuclease-free water. In a thermocycler, incubate the samples for 1 h at 16 °C, with the heated lid turned off.
- Perform a magnetic bead purification (Supplementary Appendices B) using a 1.8x ratio of SPRI magnetic beads (144 µL). Elute into 53 µL nuclease-free water and transfer 50 µL of the final supernatant (containing purified double-strand cDNA) to a clean nuclease-free PCR tube. At this point, the samples can be safely frozen at -20 °C overnight if needed.
- cDNA library end preparation: If the sample was stored at -20 °C overnight, thaw on ice prior to restarting. Pipette mix the purified ds-cDNA (50 µL) with 7 µL of end-prep reaction buffer and 3 µL of end-prep enzyme mix.
- In a thermocycler, incubate the samples at 20 °C for 30 min, and 65 °C for 30 min, with the heated lid turned off.
- Adapter ligation (perform this step on ice): Create a 1:100 dilution of adapter in nuclease-free water. Pipette the end-preparation reaction mixture (60 µL) into a tube containing 30 µL ligation master mix, 1 µL ligation enhancer, and 2.5 µL 1:100 diluted adapter. Ensure the adapter is added separately (e.g., not premixed with the ligation master mix or ligation enhancer) to avoid adapter dimer formation.
- In a thermocycler, incubate the samples at 20 °C for 15 min, with the heated lid turned off.
- Proceed immediately to magnetic bead purification (Supplementary Appendix 2) using a 0.9x ratio of SPRI magnetic beads (87 µL).
- Elute into 17 µL nuclease-free water, and transfer 15 µL of the final supernatant to a clean nuclease-free PCR tube
- Barcoding PCR: Pipette the purified, adapter-ligated cDNA (15 µL) into a tube with 3 µL USER enzyme, 25 µL Q5 Master Mix, and 10 µL unique barcoded primers. In a thermocycler, set the heated lid to 105 °C, incubate the samples at 37 °C for 15 min, then 98 °C for 30 s, and then perform 19 cycles of 98 °C for 10 s and 65 °C for 75 s. Finish the PCR by incubating at 65 °C for 5 min.
- Perform a final magnetic bead purification using a 0.8x ratio of magnetic beads (43 µL). Elute into 23 µL nuclease-free water, and following the final bead separation step, transfer 20 µL to a clean nuclease-free PCR tube. At this point, the samples can be safely frozen at -20 °C overnight if needed.
4. CSF DNA library preparation
CAUTION: Perform in an appropriate biosafety level cabinet with gloves and a laboratory coat. Reagents are harmful if swallowed, inhaled, or come into contact with skin. Dispose of used materials as hazardous chemical waste in compliance with local regulations.
- If the extracted DNA was stored at -20 °C overnight, thaw on ice prior to restarting.
- DNA fragmentation and end-preparation: Mix the first-strand reaction buffer thoroughly by vortexing and pipette mixing the solution to resuspend any precipitate. Pipette mix 3.5 µL of the extracted DNA, 22.5 µL nuclease-free water, 7 µL first-strand reaction buffer, and 2 µL first-strand enzyme mix.
- In a thermocycler, set the heated lid to 105 °C and incubate the samples at 37 °C for 5 min, then 65 °C for 30 min.
- The adapter ligation and barcoding PCR process is identical for the DNA and RNA libraries. Repeat sections 7–10 of the “CSF RNA Library Preparation” with the fragmented and end-prepped DNA to complete the DNA library preparation.
5. Library quality control and pooling
- Quantify the concentration of the sample libraries via the DNA quantification kit and corresponding fluorometer. Water controls should have substantially less (or unquantifiable) concentrations compared to the rest of the samples.
- Assess the sample library sizes by running the samples on an automated capillary electrophoresis machine.
- Alternatively, if an automated capillary electrophoresis machine is not available, further amplify 1-2 µL of the sample by performing a PCR reaction with Illumina universal primers, and then perform electrophoresis of 10 µL of the final product on a 2% agarose gel. Dispose of used materials as hazardous chemical waste in compliance with local regulations.
NOTE: The optimal library length is approximately 400–500 base pairs, so future fragmentation incubations should be adjusted accordingly if the library sizes are larger or shorter than desired. In addition, it is critical to assess for the presence of adapter dimers, which are approximately 150 base pairs in length and occur when two adapter molecules inadvertently ligate together without an insert sequence. Even if present in low quantities, adapter dimers tend to cluster and sequence efficiently on the flow cell, reducing the proportion of usable reads during a sequencing run. As such, if adapter dimers are present, perform an additional magnetic bead purification with a 0.8x ratio of SPRI magnetic beads. (Of note, if >10% of the sample is adapter dimer, it may take 2-3 rounds of magnetic bead purification to remove most of the adapter dimers).
- To ensure a similar depth of sequencing across samples within the same sequencing run, it is necessary to pool equimolar quantities of each sample. If the library sizes are comparable across all samples, simply pool equivalent masses of each RNA library and repeat separately for each DNA library.
6. Sequencing considerations
- The pooled samples are now ready for sequencing on Illumina sequencing devices. If sequencing in-house (rather than via a specialized sequencing core facility), carefully denature and dilute the pooled libraries according to the sequencing device instructions prior to loading the pool onto the reagent cartridge (typical concentration is 4 nM for most Illumina sequencing platforms).
- Calculate the estimated sequencing depth by dividing the expected data yield for the Illumina flow cell by the number of multiplexed samples in the run.
NOTE: Determining the desired sequencing depth depends on several factors, including cost, the sequencing device used, and the objectives of the experiment. Two examples of different sequencing depths are presented below in the Representative Results section.
7. Data analysis for pathogen detection
- Utilize the open-source web-based platform Chan Zuckerberg ID (CZID, https://www.czid.org, Illumina mNGS Pipeline v8.3) to perform the metagenomic analysis for pathogen detection. CZID also includes options to upload samples via command line interface and direct transfer from an Illumina BaseSpace account to CZID. A detailed overview of CZID mNGS data analysis is available at https://chanzuckerberg.zendesk.com/hc/en-us/articles/13770737266196-Guide-to-mNGS-Data-Analysis
- Log into CZID, click “upload,” select “metagenomics” for analysis type, and select the input fastq files from the sequencing run. Input the necessary information for each sample, including sample name, sample type (CSF), and nucleotide (RNA or DNA).
NOTE: The uploaded files undergo automated processing via the CZID pipeline, which incorporates several computational tools such as STAR, Bowtie2, Trimmomatic, PriceSeq, and GSNAP, among others, to remove low-quality reads and human read sequences. The pipeline then aligns and assembles the filtered sequences utilizing the National Center for Biotechnology Information (NCBI) nucleotide (NT) and protein (NR) databases via the Minimap2 and DIAMOND software programs, respectively. The final CZID output includes NT/NR taxon counts and contig counts (contiguous, overlapping segments generated during assembly) for each sample.
- Once processed, click the project folder containing the samples and navigate to the summary dashboard.
- Scroll down to examine the number of reads per sample (and compare to the expected number of reads based on the Illumina flow cell used and the number of samples multiplexed in the sequencing run), the percentage of reads passing quality control filtering, and the duplicate compression ratio to assess for biased PCR overamplification.
- The presence of contamination by environmental nucleic acid fragments - either acquired during the CSF sample collection process, or during extraction/library preparation - is a common occurrence20. To determine whether bacterial, fungal, and parasitic reads are true positives, potential methods utilizing the normalized read counts (reads per million, or rPM) can be considered:
- Create a background model (e.g., using water samples that have been extracted, prepared, and sequenced) directly within CZID to filter out any taxons with rPMs that are similar to or lower in abundance relative to the background model (and therefore, likely to represent contamination). Within the project folder, check each of the water control samples and then click “background model” to create. While analyzing each sample, select the background model and click “threshold filters” to apply an “NTZ Score” > 1 (or higher) to filter out likely contaminants.
- Utilize pre-specified rPM thresholds, such as rpMsample > 10*(average rPMcontrols), or log10(rPMsample) at least 1 log greater than log 10(average rPMentire cohort). First, download the taxon-specific rPM for each sample by checking each sample on the project page, clicking download, and selecting “Combined Sample Taxon Results.” The resulting .csv file can be analyzed using statistical software such as R or Stata as desired.
NOTE: Given the lower relative abundance of viral pathogens, any virus with known neuroinvasive potential and at least one read aligning to the viral genome should be considered a positive result (and ideally, confirmed via repeat mNGS or pathogen-specific clinical testing). In addition, the analytic pipeline requires each read in the fastq file to be computationally aligned to the most likely reference genome in the NCBI database. As a result, some reads may be relatively non-specific to a particular pathogen and could theoretically have aligned to a broad range of organisms rather than the specific taxon assigned in CZID. Given this risk, manual confirmation of positive calls should be performed by confirming the alignment via the web-based NCBI BLAST tool (the specific reads for a taxon can be directly sent from CZID to NCBI BLAST by clicking the blast icon next to the relevant taxon).