$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Determining the three-dimensional structure of specific proteins is crucial in biology. The information that is derived from doing so sheds light on the biological function and on the shape and specificity of active and/or binding sites contained in the molecule under study. In many cases, this allows mechanisms of action to be determined or, where appropriate, potential therapeutic molecules to be developed. MX is the technique most commonly used to obtain structural information, but a bottleneck is the determination of the optimal conditions to obtain well-diffracting crystals. Therefore, crystallization trials are carried out in numerous different conditions and are then screened, to find the best crystals to be used for diffraction data collection. The automation of the setup of crystallization trials1 has clearly helped in this regard. However, the subsequent steps (i.e., crystal mounting, diffraction screening, and diffraction data collection) are usually carried out manually, taking up a lot of time, effort, and resources. The automation of diffraction screening and data collection would, therefore, mean an enormous gain in time and efficiency.
Diffraction screening and data collection in MX is most often carried out at synchrotron MX beamlines at which automation has largely facilitated this process. However, in most cases, it is necessary for the scientist to be present at the beamline during an experiment or to operate it remotely. Recently, a new generation of completely automated MX beamlines has been developed2. Here, users do not need to be present, either physically or remotely, during an experimental session. This allows scientists to spend more time on less routine tasks, rather than spending entire days, and often nights, screening crystals and collecting diffraction data. The world's first fully automated beamline is the Massively Automated Sample Selection Integrated Facility (MASSIF-1, ID30A-1)2,3 at the European Synchrotron Radiation Facility (ESRF). It has a unique sample environment in which a high-capacity sample-containing dewar operates in tandem with a robotic sample changer that also acts as the beamline's goniometer4,5. MASSIF-1 is an undulator beamline equipped with a single-photon-counting hybrid pixel detector6, that operates at a fixed wavelength of 0.969 Å (12.84 keV) with an intense X-ray beam (2 x 1012 photons/s). The beam size at the sample position can be adjusted between a minimum of 10 µm (round beam) to a maximum of 100 µm x 65 µm (horizontal by vertical beam size). On average, the beamline can process, in a completely automatic fashion (see below), 120 crystals in 24 h. The operation of the beamline is based on a series of workflows7, each of which takes intelligent decisions based on the outcome of previous steps in the workflow, to ensure the measurement of the best possible data from the sample under study. In particular, the evaluation of the diffraction characteristics of an individual sample takes into account crystal volume and flux and ensures, where the crystal is larger than the X-ray beam, that only the best region of the crystal is used for subsequent data collection. Diffraction data sets are, thus, optimized for maximum resolution with minimized radiation damage2,3. Demanding data collection protocols, such as pseudo-helical (multi-position) data collection strategies for both native and single-wavelength anomalous diffraction (SAD) data collection, are also available8.
Completely automatic experiments at MASSIF-1 involve cryocooling and mounting the crystals on a magnetic sample mount suitable for the desired beamline equipment standard pins SPINE9, entering the desired experimental parameters in the 'diffraction plan' table in the Integrated System for Protein Crystallography beamlines (ISPyB)10, a web-based information management system for MX experiments, and sending the samples to the beamline. At the ESRF, all costs of the transport of the samples to/from the beamline are supported by the ESRF User Office (see the website of the ESRF11 for details). At MASSIF-1, no restrictions are placed on the loop size or crystal quality. When choosing a diffraction plan for a given crystal, the user can either use default settings or choose from specific workflows, which can be customized for each sample. Several preprogrammed workflows are available. In the MXPressE3 workflow, the sample-containing loop is first aligned to the sample position using optical centering. Then, X-ray-based centering ensures that the best region of the crystal is centered to the X-ray beam. Data collection strategies are then calculated using eEDNA, a framework for developing plugin-based applications especially for online data analysis in the X-ray experiments field, taking into account crystal volume and the real-time flux at the beamline. Following the collection of a full diffraction data set, this is then processed using a series of automatic data processing pipelines12 and the results are made available for inspection and download in ISPyB. The MXPressE SAD3 workflow is aimed at selenomethionine-containing crystals of the target protein and exploits the fact that the operating energy of MASSIF-1 is just above the Se K edge. Here, the MXPressE eEDNA data collection strategy is optimized for SAD data collection (i.e., high redundancy, and with the resolution set to where the Rmerge between Bijvoet pairs is below 5%). To screen the diffraction properties of a series of crystals without subsequent data collection, the MXScore3 workflow can be used to produce a full quality assessment of the crystals analyzed. In the MXPressI3 workflow, 180° of rotation data are collected using 0.2° oscillations and using the starting phi angle and the resolution determined by an eEDNA strategy. MXPressO3 includes a preobserved resolution into the workflow (default: dmin = 2 Å). To make an initial assessment of the crystals resulting from a crystallization trial, the MXPressM3 workflow is offered. This performs a high-dose mesh scan over the widest orientation of sample support with no data collection or centering. Recently, two new experiment workflows, MXPressP and MXPressP_SAD, which perform pseudohelical data collections, have been implemented8. The execution of all steps in all workflows can be followed online and in real-time by the user, via ISPyB.
Here we show how to prepare a fully automated MX experiment at MASSIF-1 and how to retrieve and analyze the data resulting from the experiment. As an example, we use human mitochondrial glycine cleavage system protein H (GCSH). This lipoic acid-containing protein is part of the glycine cleavage system responsible for the degradation of glycine. This system further includes the P protein, a pyridoxal phosphate-dependent glycine decarboxylase, the T protein, a tetrahydrofolate-requiring enzyme, and the L protein, a lipoamide dehydrogenase. GCSH transfers the methylamine group of glycine from the P protein to the T protein. Defects in the H protein are the cause of nonketotic hyperglycinemia (NKH) in humans13.