All samples used to determine the open reading frames (ORFs) selected in this study were collected from Starbuck Island, Site 7 (STAR7) and Caroline Atoll, Site 9 (CAR9) of the southern Line Islands. An estimated 100 L of seawater from these sites was collected below the coral boundary layer using bilge pumps, as described previously5. Contents of the pumps were subject to fractionalization through large pored filters to remove small eukaryotes and subsequently concentrated using 100 kDa tangential flow filters leaving only microbes and viral like particles (VLPs). To separate the VLPs, remaining seawater was passed through 0.45 µm filters resulting in the virome. Chloroform was introduced to this viral fraction to arrest growth of any remaining cells and stored at 4 °C.
VLPs were purified using the cesium chloride method, in which density gradients separate through centrifugation and allow for the recovery of virions at ~1.35 g/ml to 1.5 g/ml 3. Viral DNA was extracted using a CTAB/phenol:chloroform protocol and amplified via multiple displacement amplification using Phi29 reagents. Sequencing of the virome was accomplished with commercially available pyrosequencing technology.
Bioinformatics used in the processing and selection of viral ORFs for this study are as follows. Three pre-processing steps were used on the CAR9 and STAR7 viral metagenomes. First, public software was used to remove tag sequences that resulted from amplification of the viral DNA prior to sequencing27. Second, common sequencing artifacts such as sequence duplicates and low copy number were filtered out of the dataset via an additional bioinformatics program28. Lastly, removal of foreign sequence contamination29 was performed for those sequences that had ≥ 90% coverage and ≥ 94% identity with sequences in the following databases: RefSeq virus genomes; Human — Reference GRCh37; Human — Celera Genomics; Human — Craig Venter (HuRef); Human — Seong-Jin Kim (Korean); Human — Chromosome 7 version 2 (TCAG); and Human — James Watson, YanHuang (YH; Asian), Yoruba (NA18507; African) reference sequences21. Following these processes, sequences from the CAR9 samples totaled 591,600 and STAR7 sequences totaled 939,311. These sequences were uploaded onto MGRAST and assembled via assembler software using default settings. Contigs were translated into 6 reading frames and putative open reading frames (pORFs) were identified using scripts, as previously described21.
To identify unknown ORFs a number of similarity based searches were performed to remove ORFs of known function. Briefly, the following searches were performed with their corresponding searching criteria21:
- Significant similarity of ≥ 95% identity over ≥ 40 base pair (bp) by MGRAST BLAT in M5NR database.
- Significant similarity (e-value ≤ 0.001) by TBLASTN against all public metagenomes in My Metagenome Database Resource.
- Significant similarity (e-value ≤ 0.001) by BLASTP and TBLASTN against NR database.
- Significant similarity (e-value ≤ 0.001) by RPS-BLAST against the Conserved Domain Database.
- Protein translations from a select fraction of each dataset compared against solved protein structures in the Protein Data Bank.
- Calculation of dinucleotide frequencies using Dinucleotide Signatures package.
The resulting pORFs were designed for expression in E. coli using publically available gene design software. Back-translation of the amino-acid sequences employed a Universal Codon Usage Table designed to accommodate expression in E. coli with a minimum usage threshold of 2%. Restriction enzyme recognition sequences for BamHI and HindIII were excluded from the sequences to facilitate cloning. An outside company synthesized the engineered gene sequences30 and then ORFs were cloned into a medium-copy number pBAD promoter vector, pEMB11, via standard restriction enzyme cloning. All clones were transformed into E. coli K-12 strain BW 2778423.
Multi-phenotype Assay Plates (MAPs)
A high throughput and robust software pipeline was leveraged for analysis of the MAPs, the PMAnalyzer24. The pipeline was developed in a Linux server environment and performs various steps including: parsing of optical density files, formatting data into readable text files, pre-processing of growth curves for quality assurance (QA), and performing mathematical modeling techniques to analyze growth curves. The primary modeling scripts were developed in Python version 2.7.5 to make use of the PyLab module.
MAPs reproducibility was assessed using the standard error (SE) for replicate data (Figure 2A). Raw growth curves were compared to logistic growth curves to determine whether the program PMAnalyzer accurately parameterized and modeled clone growth during experimentation (data not shown). For further information on the accuracy and validity of the MAPs and PMAnalyzer see Cuevas et al.24
Following validation of the method, MAPs data was analyzed using the multiple parameters, such as maximum growth rate (µmax) and growth level (GL), provided by the processing pipeline. Comparative visualization of growth curves is often used for growth data interpretation; however, the number of curves that can be visualized at a time for the comparison has limitations. To analyze numerous growth curves simultaneously, heat map derived plots were implored to compare tens of clones grown on a single substrate against that conditions’ average response (Figure 2B). The influence placed by overexpression of a novel phage protein is observed through changes in curve parameters, specifically: lag phase, exponential phase and the maximum biomass yield (asymptote). As an example, the steep climb from lag phase into exponential phase modeled in the growth curve for the Capsid protein (Figure 2A) is reproduced by a quick change in color intensity from black to white for the same clone in the dynamic plot of Figure 2B.
To obtain a global picture of clone distribution across substrates, phenotypic classifications derived from the GL were used (Figure 3). Here, the four phenotypes are separate into four charts where the height of each bar represents the number of clones displaying that phenotype for a specific substrate. Outliers in the data are recognized as clones falling into the “gain of function” or “loss of function” category. Outliers can then be sought out individually and investigated more closely experimentally. In addition, global analysis recognizes substrate biases in the assay. Substrates such as phenylalanine, malic acid, and glycine resulted in a “no growth” classification. Substrates that consistently fall into the no growth classification, across all clones, are not heavily weighted in downstream functional characterization.
Metabolomics
The catabolic products from clones expressing unknown phage genes were identified using metabolomics. Briefly, clones were grown under either a continuous culture or serial passage in batch culture environment prior to being sent out for GC-TOFMS analysis at a metabolomics core facility. For details on sample processing, analysis, and normalization for GC-TOFMS implemented by the chosen core facility see Fiehn et al.31 Briefly, 1 ml of cold extraction solvent is added to each sample, after which samples are vortexed and sonicated in a cold bath for 5 min. Samples are finally centrifuged and half of the sample is decanted and dried down for analysis. Extracts are purified and spiked with internal retention index markers before being loaded onto the gas chromatograph and then subsequently transferred to the mass spectrometer. Data from each sample is analyzed such that signal intensities for all detected signals in the chromatogram are reported. For normalization, the abundance of peaks for each sample is summed and the total peak abundances are averaged between all samples in the set. Metabolite abundances per sample are divided by the sample’s peak abundance and then multiplied by the average peak abundance of the sample set. The resulting data is used for the analysis of metabolomics in the research discussed.
Validation of metabolomics reproducibility was required to determine the appropriate sample size for each culturing method. To detect the precision seen within samples and the variation seen across sample sizes the standard error of the mean (SM) for both n = 3 and n = 6 datasets were reviewed (Figure 5B). Regardless of the continuous culture (CC) sample size, less than 1% of the data had a SM ≤ 1.5. Median SMs were 221 and 300, and values ranged from 0 to 7.55 x 105 and 3.74 x 105, for n = 3 and n = 6 respectively. SMs were also calculated for each set of sample replicates in the serial culture (SC) method. Again, less than 1% of the data had a SM ≤ 1.5, a median SM of 137, and a range from 0 to 3.51 x 105. To compare the SM distributions between each sample set (CC n = 3 vs. CC n = 6, CC n = 3 vs. SC n = 3, and CC n = 6 vs. SC n = 3) a permutation test was performed. The distribution of either continuous culture SM values dataset was not significantly different from the distribution of the serial culture SM values (p-value = 0.0). However, the distribution of SM values for continuous culture n = 3 data was significantly different from that of the SM values for the continuous culture n = 6 data (p-value = 1.908804 x 10-49). Lastly, the coefficient of variation per metabolite was compared before and after implementation of a quality assurance (QA) step (Figure 5C). In total, 210 metabolites were removed after implementation of the QA pipeline (40% of the data). Less than 1% of the removed data had a metabolite abundance of zero, ~2% was internal standard data, ~5% was data from metabolites never before observed in E. coli, and the remaining metabolites (> 30%) had a coefficient of variation greater than 1.
As with the MAPs analysis, global observations provided an initial understanding of the depth of information metabolomics offers. To obtain a global picture, clones were hierarchically clustered based on their relative metabolite abundances providing information on clone-metabolite profiles, potential clones with related functions, and clone-metabolite outliers (Figure 6). To highlight protein functions, metabolites are separated and grouped based on common metabolic pathways. Using this analysis with preliminary results, it was evident that metabolomics is capable of separating genes from different classes (Figure 6, highlighted clones). In addition, outlier identification with the metabolomics data was determined by calculating standard scores (z scores) for each clone-metabolite pair. To ensure statistical significance, outliers were defined as a clone-metabolite pair with a Z score value of 2, accounting for only 5 percent of the data (data not shown).

Figure 1. Definitions of phenotype classifications. (A) Relationship between the growth level (GL) and maximum growth rate. Data points circled in red represent growth curves showing little to no substrate utilization. (B) Boxplot representation defining growth threshold based on the distribution of growth curves with a minimal growth rate (< 0.15 OD/hr). (C) The variance and standard deviation of GL is calculated for substrate D-galactose. Short dashed lines represent two standard deviations away from the mean. Please click here to view a larger version of this figure.

Figure 2. MAPs validation via precision and differentiation. (A) Growth curves for annotated structural (Capsid) and metabolic clones (Thioredoxin), two novel metabolic clones (EDT2440, EDT2441), and the average response of clones grown on sucrose, D-galactose and D-mannose in the MAPs. Blue lines indicate the standard error seen between replicate data (n = 3). (B) The growth curves for 47 different clones are represented as heat maps for sucrose, D-galactose and D-mannose. The annotated structural (green circle) and metabolic clones (orange circle), two novel metabolic clones (dark and light blue circles), and the average response (red circle) are highlighted. Please click here to view a larger version of this figure.

Figure 3. Clone distribution for each phenotype across multiple substrates. The phenotype-clone count for 47 clones across 72 carbon-specific growth conditions. Table provides direct counts for each phenotype. Please click here to view a larger version of this figure.

Figure 4. Diagram detailing the construction of the continuous culture apparatus. (A) Steps used to build ports α-γ of the continuous culture reactor, (B) steps used to build the out flow port of the continuous culture reactor, and (C) the steps to build ports δ and ε of the continuous culture feeding bottle. Please click here to view a larger version of this figure.

Figure 5. Comparison of the phenomic methods presented. (A) Workflow for the preparation of the Multi-phenotype Assay Plates (MAPs), continuous cultures and serial cultures. (B) The percentage of Standard error of mean (SM) counts for both n = 3 and n = 6 sample sizes for the continuous culture (CC) and serial culture (SC) preparation methods for metabolomics. The y-axis is on a log scale. (C) The distributions of coefficients of variation (CV) per metabolite, before and after implementation of the QA pipeline for the continuous culture (CC) method. Please click here to view a larger version of this figure.

Figure 6. Metabolomic profiles of clones grown in continuous culture. The median metabolite abundances for a set of metabolites are plotted for 84 clones grown in continuous cultures. Metabolite profiles for annotated structural (Capsid) and metabolic clones (Thioredoxin), two novel metabolic clones (EDT2440, EDT2441), and the average metabolic response are highlighted in red. Please click here to view a larger version of this figure.
| Compound | Carbon | Nitrogen | Sulfur | Phosphorus |
| Glycerol | − | 0.40% | 0.40% | 0.40% |
| Ammonium chloride | 9.5 mM | − | 9.5 mM | 9.5 mM |
| Sodium sulfate | 0.250 mM | 0.250 mM | − | 0.250 mM |
| Magnesium sulfate | 1.0 mM | 1.0 mM | − | 1.0 mM |
| Potassium phosphate | 1.32 mM | 1.32 mM | 1.32 mM | − |
| Magnesium chloride | − | − | * | − |
| Potassium chloride | 10 mM | 10 mM | 10 mM | 10 mM |
| Calcium chloride | 0.5 µM | 0.5 µM | 0.5 µM | 0.5 µM |
| Sodium chloride | 5 mM | 5 mM | 5 mM | 5 mM |
| Ferric chloride | 6 µM | 6 µM | 6 µM | 6 µM |
| L- arabinose | 0.10% | 0.10% | 0.10% | 0.10% |
| MOPS pH 7.4 | 1x | 1x | 1x | 1x |
Table 1. The compounds and concentrations of the different basal media used in the MAPs. *1.0 mM of magnesium chloride is substituted. 1x MOPS = 40 mM MOPS, 4 mM Tricine.
| Carbon substrates | Nitrogen substrates | Sulfur substrates | Phosphorus substrates |
| 2 deoxy-D-ribose | 2-deoxy-D-ribose | 1-butane-sulfonic acid | adenosine-5-monophoshate |
| 4 hydroxy-phenyl acetic acid | acetamide | acetyl cysteine | beta-glycerophosphate |
| acetic acid | adenine | D-cysteine | creatinephosphate |
| adenosine-5-monophosphate | adenosine | D-methionine | D-glucose-6-phosphate |
| adonitol | allantoin | diethyl-dithiophosphate | diethyl-dithiophosphate |
| alpha-D-glucose | beta-phenylethylamine | DL-ethionine | DL-alpha-glycerophosphate |
| alpha-D-lactose | biuret | glutathione | potassium phosphate |
| alpha-D-melebiose | cytidine | isethionic acid | sodium pyrophosphate |
| citric acid | cytosine | L-cysteic acid | sodium thiophosphate |
| D-alanine | D-alanine | L-Cysteine | |
| D-arabinose | D-asparagine | L-djenkolic acid |
| D-arabitol | D-aspartate | L-methionine |
| D-asparagine | D-cysteine | magnesium sulfate |
| D-aspartate | D-glucosamine | methane sulfonic acid |
| D-cellubiose | D-glutamic acid | N-acetyl-DL-methionine |
| D-cysteine | DL-alpha-amino-N-butyric acid | N-acetyl-L-cysteine |
| D-fructose | D-methionine | potassium-tetra-thionate |
| D-galactose | D-serine | sodium thiosulfate |
| D-glucosamine | D-valine | sulfanic acid |
| D-glucose | gamma-amino-N-butyric acid | taurine |
| D-glucose-6-phosphate | glycine | taurocholic acid |
| D-glutamate | guanidine | thiourea |
| D-mannose | histamine | |
| D-raffinose | inosine |
| D-ribose | L-alanine |
| D-salicin | L-arginine |
| D-serine | L-asparagine |
| D-trehalose | L-citrulline |
| D-xylose | L-cysteine |
| dulcitol | L-glutamic acid |
| glycerol | L-glutamine |
| glycine | L-glutathione |
| i-erythritol | L-histidine |
| inosine | L-isoleucine |
| L-alanine | L-leucine |
| L-arabinose | L-lysine |
| L-arabitol | L-methionine |
| L-asparagine | L-ornithine |
| L-aspartate | L-phenyl-alanine |
| L-cysteic acid | L-proline |
| L-cysteine | L-pyro-glutamic acid |
| L-fucose | L-serine; L-threonine |
| L-glutamic Acid | L-tryptophan |
| L-glutamine | L-valine |
| L-isoleucine | N-acetyl-D-glucosamine |
| L-leucine | putrescine |
| L-lysine | thiourea |
| L-methionine | thymidine |
| L-phenylalanine | thymine |
| L-pyro-glutamic acid | tyramine |
| L-rhamnose | tyrosine |
| L-serine | uridine |
| L-sorbose | |
| L-threonine |
| L-tryptophan |
| L-valine |
| L-xylose |
| lactate |
| lactulose |
| malate |
| myo-inositol |
| oxalic acid |
| potassium sorbate |
| propionic acid |
| putrescine |
| quinic acid |
| sodium pyruvate |
| sodium succinate |
| sucrose |
| thymidine |
| xylitol |
Table 2. List of substrates used in the MAP experiments.