Obtaining high quality DNA for human genetic studies is essential in the disease gene discovery process. Blood, though requiring an invasive procedure and also being more expensive than saliva collection, is favored for creating immortalized cell lines as an infinite source of DNA, or iPSCs for functional studies, and sometimes blood DNA is used when cell lines are not available. However, obtaining blood requires a trained phlebotomist and blood has a shorter half-life than saliva1. DNA from saliva is less expensive and easier to obtain, since it can be collected and sent through the mail without the need for a phlebotomist, thereby increasing potential subject pools well beyond the catchment area of hospitals and laboratories2. Study enrollment may be improved when subjects have the option of giving a saliva sample instead of blood3, 4. Concerns about the quantity and quality of DNA from saliva may have limited its widespread use despite numerous studies recent studies showing the suitability of whole saliva, with an average of 4.3 x 105 cells per milliliter, for DNA testing over the older buccal swabs methods that did not obtain significant amounts of saliva2, 3, 4, 5, 6. While a modest literature exists showing the suitability of whole saliva derived DNA for genotyping applications including microarray-based methods8, 9, 10, no studies have examined next generation sequencing (NGS). The goal for optimizing this whole saliva DNA extraction protocol was to maximize quantity and quality for genetics applications in a cost effective way that is easily implemented in laboratories with common reagents and consumables.
DNA extraction from saliva requires several procedures: 1) collection and storage, 2) cell lysis, 3) RNase treatment, 4) protein precipitation, 5) ethanol precipitation, 6) DNA rehydration. The DNA Stabilization Buffer solution, described previously2, functions adequately without alteration. No attempt to optimize the RNase treatment and DNA rehydration steps was made. For each remaining step, several variables that could affect yield were identified. Each variable was manipulated individually and improvement in yield and quality was assessed statistically. For variables that were shown to improve yield and/or DNA quality, the optimal values were included in the final protocol.