$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
An accessible RABV, nanopore-based, whole genome sequencing workflow was developed by Brunker et al.61, using resources from the ARTIC network46. Here, we present an updated workflow, with complete sample-to-sequence-to-interpretation steps. The workflow details the preparation of brain tissue samples for whole genome sequencing, presents a bioinformatics pipeline to process reads and generate consensus sequences, and highlights two rabies-specific tools to automate lineage assignment and determine phylogenetic context. The updated workflow also provides comprehensive instructions for the setup of appropriate computational and laboratory workspaces, with considerations for implementation in different contexts (including low-resource settings). We have demonstrated the successful implementation of the workflow in both academic and research institute settings in four RABV endemic LMICs with no or limited genomic surveillance capacity. The workflow has proven resilient to application across diverse settings, and comprehensible by users with varying expertise.
This workflow for RABV sequencing is the most comprehensive publicly available protocol (covering sample-to-sequence-to-interpretation steps) and specifically adapted to reduce both startup and running costs. The time and cost required for library preparation and sequencing with nanopore technology is greatly reduced relative to other platforms, such as Illumina61, and continual technology developments are improving sequence quality and accuracy to be comparable with Illumina62.
This protocol is designed to be resilient in diverse low-resource contexts. By referring to the troubleshooting and modifications guidance provided alongside the core protocol, users are supported to adapt the workflow to their needs. The addition of user-friendly bioinformatic tools to the workflow constitutes a major development to the original protocol, providing rapid and standardized methods that can be applied by users with minimal prior bioinformatics experience to interpret sequence data in local contexts. The capacity to do this in situ is often limited by the need to have specific programming and phylogenetic skills, which require an intensive and long-term skills training investment. While this skillset is important to thoroughly interpret sequence data, basic and accessible interpretation tools are equally desirable in order to capacitate local "sequencing champions", whose core expertise may be wet lab based, enabling them to interpret and take ownership over their data.
As the protocol has been undertaken for a number of years in several countries, we now can provide guidance on how to optimize multiplex primer schemes to improve coverage and deal with accumulated diversity. Efforts have also been made to help users improve the cost-effectiveness or to allow for ease of procurement in a given region, which is typically a challenge for the sustainability of molecular approaches63. For example, in Africa (Tanzania, Kenya, and Nigeria), we opted for blunt/TA ligase master mix at the adapter ligation step, which was more readily available from local suppliers and a cheaper alternative to other ligation reagents.
From experience, there are several ways of reducing the cost per sample and per run. Reducing the number of samples per run (e.g., from 24 down to 12 samples) can extend the life of flow cells over multiple runs, whereas increasing the number of samples per run maximizes the time and reagents. In our hands, we were able to wash and reuse flow cells for one in every three sequencing runs, enabling an additional 55 samples to be sequenced. Washing the flow cell immediately after use, or if not possible, removing the waste fluid from the waste channel after every run, seemed to preserve the number of pores available for a second run. Taking into consideration the initial number of pores available in a flow cell, one run can also be optimized to plan how many samples to run in a particular flow cell.
Though the workflow aims to be as comprehensive as possible, with the addition of detailed guidance and signposted resources, the procedure is still complex and can be daunting for a new user. The user is encouraged to seek in-person training and support, ideally locally, or alternatively through external collaborators. In the Philippines for example, a project on capacity building within regional laboratories for SARS-CoV-2 genomic surveillance using ONT has developed core competencies among health care diagnosticians that are readily transferrable to RABV sequencing. Important steps, such as SPRI bead clean up, can be difficult to master without hands-on training, and ineffective clean up can damage the flow cell and compromise the run. Sample contamination is always a major concern when amplicons are being processed in the lab and can be difficult to eliminate. In particular, cross-contamination between samples is extremely difficult to detect during post-run bioinformatics. Good laboratory technique and practices, such as maintaining clean work surfaces, separating pre- and post-PCR areas, and incorporating negative controls, are imperative to ensure quality control. The fast pace of nanopore sequencing developments is both an advantage and disadvantage for routine RABV genomic surveillance. Continuing improvements to nanopore's accuracy, accessibility, and protocol repertoire widen and improve the scope for its application. However, the same developments make it challenging to maintain standard operating procedures and bioinformatic pipelines. In this protocol, we provide a document assisting the transition from older to current nanopore library preparation kits (Table of Materials).
A common roadblock to sequencing in LMICs is accessibility, including not only the cost but also the ability to procure consumables in a timely manner (in particular sequencing reagents, which are relatively new to procurement teams and suppliers) and computational resources, as well as simply having access to stable power and the internet. Using portable nanopore sequencing technology as the foundation of this workflow helps with many of these accessibility issues, and we have demonstrated the use of our protocol across a range of settings, conducting the full protocol and analysis in-country. Admittedly, procuring equipment and sequencing consumables in a timely manner remains a challenge and, in many instances, we were forced to carry or ship reagents from the UK. However, in some areas, we were able to rely entirely on local supply routes for reagents, benefiting from investment in SARS-CoV-2 sequencing (e.g., the Philippines) that has streamlined procurement processes and begun to normalize the application of pathogen genomics.
The need for a stable internet connection is minimized by one-time-only installs; for instance, GitHub repositories, software download, and nanopore sequencing itself only require internet access to start the run (not throughout) or can be performed completely offline with agreement from the company. If mobile data is available, a phone can be used as a hotspot to the laptop to begin the sequencing run, before disconnecting for the run duration. When routinely processing samples, data storage requirements can grow rapidly, and ideally data would be stored on a server. Otherwise, solid state drive (SSD) hard drives are relatively cheap to source.
While we recognize that there are still barriers to genomic surveillance in LMICs, increasing investment in building genomics accessibility and expertise (e.g., Africa Pathogen Genomics Initiative [Africa PGI])64 suggests that this situation will improve. Genomic surveillance is critical for pandemic preparedness6, and capacity can be established through routinizing the genomic surveillance of endemic pathogens such as RABV. Global disparities in sequencing capacities highlighted during the SARS-CoV-2 pandemic should be a driver of catalytic change to address these structural inequities.
This sample-to-sequence-to-interpretation workflow for RABV, including accessible bioinformatics tools, has the potential to be used to guide control measures targeting the goal of zero human deaths from dog-mediated rabies by 2030, and ultimately for the elimination of RABV variants. Combined with relevant metadata, genomic data generated from this protocol facilitates rapid RABV characterization during outbreak investigations and in the identification of circulating lineages in a country or region60,61,65. We illustrate our pipeline mostly using examples from dog-mediated rabies; however, the workflow is directly applicable to wildlife rabies. This transferability and low cost minimize the challenges in making routine sequencing easily available, not only for rabies but also for other pathogens46,66,67, to improve disease management and control.