$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
In the field of structural biology, X-ray crystallography is regarded as the gold standard technique to determine the atomic-resolution structures of macromolecules. It has been utilized extensively to understand the molecular basis of diseases, guide rational drug design projects and elucidate the catalytic mechanism of enzymes1,2. Although structural data provides a wealth of knowledge, the process of protein expression and purification, crystallization and structure determination can be extremely laborious. Several bottlenecks are commonly encountered that hinder the progress of these projects and this must be addressed to efficiently streamline the crystal structure determination pipeline.
Following recombinant expression and purification, preliminary conditions that are conducive to crystallization must be identified which is often an arduous and time-consuming aspect of X-ray crystallography. Commercial sparse matrix screens that consolidate known and published conditions have been developed to ease this bottleneck3,4. However, it is common to generate few hits from these initial screens despite using highly pure and concentrated protein samples. Observing clear drops indicates that the protein may not be reaching the levels of supersaturation required to nucleate a crystal. To encourage crystal nucleation and growth, seeds produced from pre-existing crystals can be added to the conditions and this allows for increased sampling of the crystallization space. Ireton and Stoddard first introduced the microseed matrix screening method5. Poor quality crystals were crushed to make a seed stock and then added systematically to crystallization conditions containing different salts to generate new diffraction-quality crystals that would not have otherwise formed. This technique was further improved by D'Arcy et al. who developed random microseed matrix screening (rMMS) in which seeds were introduced into a spare matrix crystallization screen6,7. This improved the quality of crystals and increased the number of crystallization hits on average by a factor of 7.
After crystals are successfully produced and an X-ray diffraction pattern is obtained, another bottleneck in the form of solving the 'phase problem' is encountered. During the data acquisition process, the intensity of diffraction (proportional to the square of the amplitude) is recorded but the phase information is lost, giving rise to the phase problem that halts immediate structure determination8. If the target protein shares high sequence identity to a protein with a previously determined structure, molecular replacement can be used to estimate the phase information9,10,11,12. Although this method is fast and inexpensive, model structures may not be available or suitable. The success of the homology model-based molecular replacement method drops significantly as sequence identity falls below 35%13. In the absence of a suitable homology model, ab initio methods, such as ARCIMBOLDO14,15 and AMPLE16, can be tested. These methods use computationally predicted models or fragments as starting points for molecular replacement. AMPLE, which uses predicted decoy models as starting points, struggles to solve structures of large (>100 residues) proteins and proteins containing predominately β-sheets. ARCIMBOLDO, which attempts to fit small fragments to extend into a larger structure, is limited to high resolution data (≤2 Å) and by the ability of algorithms to expand the fragments into a full structure.
If molecular replacement methods fails, direct methods such as isomorphous replacement17,18 and anomalous scattering at a single wavelength (SAD19) or multiple wavelengths (MAD20) must be used. This is often the case for truly novel structures, where the crystal must be formed or derivatized with a heavy atom. This can be achieved by soaking or co-crystallizing with a heavy atom compound, chemical modification (such as 5-bromouracil incorporation in RNA) or labelled protein expression (such as incorporating selenomethionine or selenocysteine amino acids into the primary structure)21,22. This further complicates the crystallization process and requires additional screening and optimization.
A new class of phasing compounds, including I3C (5-amino-2,4,6-triiodoisophthalic acid) and B3C (5-amino-2,4,6-tribromoisophthalic acid), offer exciting advantages over pre-existing phasing compounds23,24,25. Both I3C and B3C feature an aromatic ring scaffold with an alternating arrangement of anomalous scatters required for direct phasing methods and amino or carboxylate functional groups that interact specifically with the protein and provide binding site specificity. The subsequent equilateral triangular arrangement of heavy metal groups allows for simplified validation of the phasing substructure. At the time of writing, there are 26 I3C-bound structures in the Protein Data Bank (PDB), of which 20 were solved using SAD phasing26.
This protocol improves the efficacy of the structure determination pipeline by combining the methods of heavy metal derivatization and rMMS screening to simultaneously increase the number of crystallization hits and simplify the crystal derivatization process. We demonstrated this method was extremely effective with hen egg white lysozyme and a domain of a novel lysin protein from bacteriophage P6827. Structure solution using the highly automated Auto-Rickshaw structure determination pipeline is described, specifically tailored for the I3C phasing compound. There exists other automated pipelines that can be used such as AutoSol28, ELVES29 and CRANK230. Non-fully automated packages such as SHELXC/D/E can also be used31,32,33. This method is particularly beneficial to researchers who are studying proteins lacking homologous models in the PDB, by significantly reducing the number of screening and optimization steps. A prerequisite for this method is protein crystals or a crystalline precipitate of the target protein, obtained from previous crystallization trials.