Individual identification
The ability of the footprint identification technique to classify individual cheetah is contingent on two factors, the use of a standardized footprint collection protocol and a new statistical model based on a cross-validated pairwise discriminant analysis with a Ward's clustering analysis. These are facilitated by an integrated graphic user interface for data visualization (Fig. 1). Minimal equipment is needed, making this technique cost-effective (Materials List). Data collected with the footprint included the number of cheetahs, number of footprint images collected, range of footprints per cheetah, number of trails, range of trails per cheetah and age-range of cheetahs (Table 1).
781 footprints (M:F 395:386) belonging to 110 trails, from 38 individuals, were collected for the training dataset. Table 1 gives a summary of data collected. Using the feature extraction window (Fig. 2) a set of 25 landmark points were able to generate 15 derived points on each footprint image. From these landmark and derived points 136 variables were generated for each footprint, comprising distances, angles and areas. Each row in the database therefore represented the 136 variables generated by a single footprint. Footprints were processed by trail. A varying number of rows represented each trail, and were marked as such.
These data were duplicated into the data table as an entity then referred to as the Reference Centroid Value (RCV) that acts to stabilize the pairwise comparison of trails necessary for individual classification. The pair-wise analysis window (Fig. 3) was designed to help validate the data and/or test for data from unknown populations. Figure 4 shows the outcome of a pair-wise comparison of trails from the same individual (A) and two different individuals (B) based on the footprint identification technique customized model. The classifier incorporated into the model is based on the presence or absence of overlap between the ellipses. Note that the analysis is performed for each pairwise comparison in the presence of a third entity, i.e., the reference centroid value (RCV).
Using a robust pairwise cross-validated discriminant analysis with a Ward's clustering analysis, an algorithm was generated to provide effective classification of individuals. The footprint identification technique algorithm is based on three adjustable entities; the number of measurements used, the ellipse size (confidence interval used), and the threshold value that determines the cut-off value for the clusters. Each of these entities is adjusted in the software until the highest accuracy for classification is achieved for the training set of animals of known identity. This same algorithm can then be used to identify unknown cheetahs. For example, Figures 5a, b & c show a dendrogram of a sample of trails from seven cheetahs showing the correct prediction when the algorithm is optimized (a) and when the algorithm is suboptimal (b & c).
Holdback trials were conducted to validate the algorithm derived from the training set of 'known' individuals. These were carried out sequentially by varying the proportion of cheetahs in the test and training sets. Rather than apportioning cheetahs to training and test sets arbitrarily, analyses were performed sequentially increasing the test set size. For each test set, 10 iterations were performed with cheetahs being selected randomly for each iteration. For each test set, this allowed a mean value to be calculated. Figure 6. shows the varying test size plotted against itself (red), and on the y-axis the predicted value for each test size iteration (green) and the mean predicted value for each test size (blue). The plot demonstrates that even when the test set size is increased considerably (n=28) compared with the training set size (n=10), the mean predicted value is similar to the expected value.
Using several holdback trials, the accuracy of individual identification was consistently >90% for both the predicted number of individuals and, just as importantly, the classification of trails, i.e., whether the trails from the same individual (self-trails) and those from different individuals (non-self-trails) classified correctly. A cluster dendrogram representing all 38 individual cheetahs is shown (Fig. 7). There were 110 trails, generating a total of 5,886 pairwise comparisons. Of these, there were 46 misclassifications giving an accuracy of 99% (Table 2).
| # of cheetahs | # of footprint images | Range of footprints per cheetah | # of trails | Range of trails per cheetah | Age range (yrs) |
| Females | 16 | 386 | 12 - 36 | 55 | 2 - 5 | 2.5 - 8.5 |
| Males | 22 | 395 | 7 - 32 | 54 | 1 - 4 | 1 - 11 |
| Total | 38 | 781 | 7 - 36 | 109 | 1 - 5 | 1 - 11 |
Table 1. Summary of collected data. The number of cheetahs, the number of footprint images collected, the range of footprints per cheetah, the number of trails, the range of trails per cheetah and the age-range of cheetahs.
| Self | Non-self | Total | Misclassifications |
| Self (N) | 117 | 9 | 126 | 9 |
| Self (%) | 93 | 7 | 100 | 7 |
| Non-self (N) | 37 | 5,723 | 5,760 | 37 |
| Non-self (%) | 1 | 99 | 100 | 1 |
| Total (N) | - | - | 5,886 | 46 |
| Total (%) | - | - | 100 | 1 |
Table 2. The output in the footprint identification technique software shows the classification of trails based on pairwise comparison. 'Self ' refers to trails from the same individual and 'non-self', trails from different individuals. Each trail was matched against every other trail using a customized robust cross-validated discriminant analysis model. 110 trails resulted in 5,886 pairwise comparisons and the overall classification accuracy was 99%.

Figure 1. The opening main menu window in the footprint identification technique. This is an image identification add-in to the data visualization software, designed to classify footprints by individual, sex and age-class from morphometric measurements. A graphic user interface allows the seamless navigation between different options. Please click here to view a larger version of this figure.

Figure 2. The feature extraction window. Capabilities include drag and drop images, automatic resizing to the window, rotation of images for standardization, substrate depth factoring, etc. Pre-assigned landmark points are manually positioned and generate a series of scripted derived points to enable the extraction of metrics in the form of distances, angles and areas. The output is in the form of a row of data providing the x.y co-ordinates and the metrics. Please click here to view a larger version of this figure.

Figure 3. Pairwise data analysis window in the footprint identification technique. Once a database of measurements has been created, the pair-wise analysis window is designed to help validate the data and/or test for data from unknown populations. The analysis is based on a customized model incorporating a constant, reference centroid value (RCV), which compares pairs of trails sequentially16,17. The final output is in the form of a cluster dendrogram that provides a prediction for the number of individuals and the relationship between the trails. Please click here to view a larger version of this figure.

Figure 4. Pairwise comparisons. The figure shows the outcome of a pair-wise comparison of trails from the same individual (A) and two different individuals (B) based on a customized model in the data visualization software. The classifier incorporated into the model is based on the presence or absence of overlap between the ellipses. Note that the analysis is performed for each pairwise comparison in the presence of a third entity, i.e., the reference centroid value (RCV). Please click here to view a larger version of this figure.

Figure 5. A dendrogram of a sample of trails from seven cheetahs showing the correct prediction when the algorithm is optimized (a) and when the algorithm is suboptimal (b & c). d shows the outcome with 18 variables, with the sliding scale moved in one direction to show that the chance of ten cheetahs is less than 50%. The algorithm is based on three adjustable entities; the number of measurements used, the ellipse size (confidence interval used) and finally, the threshold value that determines the cut-off value for the clusters. Please click here to view a larger version of this figure.

Figure 6. A holdback trial carried out sequentially by varying the proportion of cheetahs in the test and training sets. Rather than apportioning cheetahs to training and test sets arbitrarily, an analysis was performed sequentially increasing the test set size. For each test set, ten iterations were performed with cheetahs being selected randomly for each iteration. For each test set, this allowed a mean value to be calculated. The figure shows the varying test size plotted against itself (red), and on the y-axis the predicted value for each test size iteration (green) and the mean predicted value for each test size (blue). The plot demonstrates that even when the test set size is increased considerably (n=28) compared with the training set size (n=10), the mean predicted value is similar to the expected value. Please click here to view a larger version of this figure.

Figure 7. Dendrogram showing the predicted outcome when all 110 trails from 38 cheetahs are included in the analysis. Note the fidelity of the trails forming the clusters. Interestingly, many of the misclassifications were between littermates, e.g., cheetah Letotse /Duma and Vincent/Bonsai. Please click here to view a larger version of this figure.