Figure 1 shows the workflow for using the AI software to create a model for the MN assay. The user loads the desired .daf files into the AI software, then assigns objects to the ground truth model classes using the AI-assisted cluster (Figure 2) and predict (Figure 3) tagging algorithms. Once all ground truth model classes have been populated with sufficient objects, the model can be trained using the RF or CNN algorithms. Following training, the performance of the model can be assessed using tools including class distribution histograms, accuracy statistics, and an interactive confusion matrix (Figure 4). From the results screen in the AI software, the user can either return to the training portion of the workflow to enhance the ground truth data or, if sufficient accuracy has been achieved, the user can use the model to classify additional data.
Using both the cluster and predict algorithms, 190 segments with a total of 285,000 objects were assigned to the proper ground truth classes until all classes were populated with between 1,500 and 10,000 images. In total, 31,500 objects (only 10.5% of the initial objects loaded) were used in the training of this model. Precision (percentage of false positives), recall (percentage of false negatives), and F1 score (balance between precision and recall) are available in the deep learning software package to quantify model accuracy. Here, these statistics ranged from 86.0% to 99.4%, indicating high model accuracy (Figure 4).
Using Cyt-B, background MN frequencies for all control samples were between 0.43% and 1.69%, comparing well to literature17. Statistically significant increases in MN frequency, ranging from 2.09% to 9.50% for (Mitomycin C) MMC and from 2.99% to 7.98% for Etoposide, were observed when compared to solvent controls and compared well to manual microscopy scoring. When examining the negative control Mannitol, no significant increases in MN frequency were observed. Additionally, increasing cytotoxicity with the dose was observed for both Etoposide and MMC, with both microscopy and AI showing similar trends across the dose range. For Mannitol, no observable increase in cytotoxicity was seen (Figure 5).
When not using Cyt-B, background MN frequencies for all control samples were between 0.38% and 1.0%, consistent with results published in the literature17. Statistically significant increases in MN frequency, ranging from 2.55% to 7.89% for MMC and from 2.37% to 5.13% for Etoposide, were observed when compared to solvent controls and compared well to manual microscopy scoring. When examining the negative control Mannitol, no significant increases in MN frequency were observed. Further, increasing cytotoxicity with the dose was observed for both Etoposide and MMC, with both microscopy and AI showing similar trends across the dose range. For Mannitol, no observable increase in cytotoxicity was seen (Figure 5).
When scoring by microscopy, from each culture, 1,000 binucleated cells were scored to assess MN frequency and another 500 mononucleated, binucleated, or polynucleated cells were scored to determine cytotoxicity in the Cyt-B version of the assay. In the non-Cyt-B version of the assay, 1,000 mononucleated cells were scored to assess MN frequency. By IFC, an average of 7,733 binucleated cells, 6,493 mononucleated cells, and 2,649 polynucleated cells were scored per culture to determine cytotoxicity. MN frequency was determined from within the binucleated cell population for the Cyt-B version of the assay. For the non-Cyt-B version of the assay, an average of 27,866 mononucleated cells were assessed for the presence of MN (Figure 5).

Figure 1: AI software workflow. The user begins by selecting the .daf files to be loaded into the AI software. Once the data has been loaded, the user begins to assign objects to the ground truth model classes through the user interface. To aid in ground truth population, the cluster and predict algorithms can be used to identify imagery with similar morphology. Once sufficient objects have been added to each model class, the model can be trained. Following training, the user can assess the performance of the model using the tools provided, including an interactive confusion matrix. Finally, the user can either return to the training portion of the workflow to enhance the ground truth data or, if sufficient accuracy has been achieved, the user can step out of the training/tagging workflow loop and use the model to classify additional data. Please click here to view a larger version of this figure.

Figure 2: Cluster algorithm. The cluster algorithm can be run at any time on a segment of 1,500 objects randomly selected from the input data. This algorithm groups similar objects within a segment together according to the morphology of both unclassified objects and objects that have been assigned to the ground truth model classes. Example imagery shows binucleated, mononucleated, and multinucleated cells, and cells with irregular morphology. Clusters containing mononucleated cells fall on one side of the object map, while clusters with multinucleated cells are on the opposite side of the object map. Binucleated cell clusters fall somewhere between mono- and multinucleated cell clusters. Finally, clusters with irregular morphology fall in a different area of the object map altogether. The user interface permits adding entire clusters, or select objects within clusters, to the ground truth model classes. Please click here to view a larger version of this figure.

Figure 3: Predict algorithm. The predict algorithm requires a minimum of 25 objects in each ground truth model class and attempts to predict the most appropriate model class to assign unclassified objects within a segment. The predict algorithm is more robust in comparison to the cluster algorithm with respect to the identification of subtle morphologies in images (i.e., mononucleated cells with MN [yellow] versus mononucleated cells without MN [red]). Objects with these similarities are placed in close proximity on the object map; however, the user is easily able to inspect the images in each predicted class and assign objects to the appropriate model class. Objects that the algorithm is unable to predict a class for will remain as 'unknown'. The predict algorithm permits users to rapidly populate the ground truth model classes, particularly in the case of events that are considered rare and challenging to find within the input data, such as micronucleated cells. Please click here to view a larger version of this figure.

Figure 4: Confusion matrix with model results. The results screen of the AI software presents the user with three different tools to assess model accuracy. (A) The class distribution histograms permit the user to click on the bins of the histogram to assess the relationship between objects in the truth populations and objects that were predicted to belong to that model class. In general, the closer the percentage values between the truth and predicted populations are to one another for a given model class, the more accurate the model. (B) The accuracy statistics table allows the user to assess three common machine learning metrics to assess model accuracy: precision, recall, and F1. In general, the closer these metrics are to 100%, the more accurate the model is at identifying events in the model classes. Finally, (C) the interactive confusion matrix provides an indication of where the model is misclassifying events. The on-axis entries (green) indicate objects from the ground truth data that were classified correctly during training. Off-axis entries (shaded orange) indicate objects from the ground truth data that were incorrectly classified. Various examples of misclassified objects are shown, including (i) a mononucleated cell classified as a mononucleated cell with MN, (ii) a binucleated cell classified as a binucleated cell with MN, (iii) a mononucleated cell classified as a cell having irregular morphology, (iv) a binucleated cell with MN classified as a cell having irregular morphology, and (v) a binucleated cell with a MN classified as a binucleated cell. Please click here to view a larger version of this figure.

Figure 5: Genotoxicity and cytotoxicity results. Genotoxicity measured by the percentage of MN by microscopy (clear bars) and AI (dotted bars) following a 3 h exposure and 24 h recovery for Mannitol, Etoposide, and MMC using both the (A-C) Cyt-B and (D-F) non-Cyt-B methods. Statistically significant increases in MN frequency compared to controls are indicated by asterisks (*p < 0.001, Fisher's Exact Test). Error bars represent the standard deviation of the mean from three replicate cultures at each dose point except for MMC by microscopy, where only duplicate cultures were scored. This figure has been modified from Rodrigues et al.15. Please click here to view a larger version of this figure.