A model may perform well on biological examples used during development yet make poorer predictions on an independent dataset. This discrepancy signals overfitting, meaning the algorithm has adapted too closely to the training examples rather than capturing patterns that generalize. Comparing performance across the two phases therefore provides evidence about whether predictions extend to new biological samples.
Training changes the model’s parameters as the algorithm identifies patterns in biological data. These parameters encode the model’s learned response to the examples supplied during development, whether the data are labeled or unlabeled. The resulting parameter settings determine how the model later produces predictions, making training the phase in which the model’s behavior is established.
Keeping the testing dataset separate prevents information from the evaluation examples from influencing model development. If those examples affect parameter adjustment or other development decisions, measured performance may no longer represent prediction on genuinely new samples. An unseen dataset consequently gives a more reliable indication of generalization beyond the biological data used for training.
The training phase can identify patterns in either labeled or unlabeled biological data, but the available information differs between these formats. Labeled data include designated categories or outcomes that can guide the model, whereas unlabeled data provide examples without those designations. This distinction affects what patterns the algorithm can learn and how later predictions are interpreted.
First, biological data are separated into development and independent evaluation sets. The algorithm then adjusts its parameters using the training data and identifies relevant patterns. After development is complete, the model generates predictions for the unseen testing data, and those predictions are measured to estimate performance and assess behavior on new biological samples.
This workflow supports several types of biological analysis. A model can classify cell types, predict molecular properties, or analyze genomic sequences, with testing used to examine how well its learned patterns transfer beyond the development examples. Separating the datasets helps researchers judge whether results reflect broader biological structure rather than only the supplied training samples.
Testing provides an estimate of how the model performs when it encounters biological data that were not used during development. Its predictions can be examined to determine whether the learned patterns generalize beyond familiar examples. This outcome is especially important when the intended task involves new samples, such as cells, molecules, or genomic sequences.