Accuracy may appear high when a model correctly predicts many majority-class examples, even if it fails to identify uncommon cases. Because the majority class contributes most examples to the learning objective, the algorithm can favor that category during training. In medical prediction, this creates an overly reassuring result when the missed minority cases represent rare diseases or adverse clinical events.
Class weighting changes the learning objective so that errors involving the underrepresented category receive greater influence. Resampling instead changes the training data by altering the representation of existing examples, while synthetic minority examples add generated cases for that category. These strategies pursue the same broad goal, but they modify either the contribution of errors or the composition of the training data.
Decision-threshold adjustment changes the criterion used to assign a prediction to one class or another after the model has produced its outputs. This can make the system more responsive to minority cases that would otherwise be overlooked. Researchers should then examine precision, recall, F1 score, and the area under the precision-recall curve to understand the resulting performance.
Precision indicates how often positive predictions are correct, whereas recall indicates how effectively the model identifies the relevant minority cases. F1 score combines these two measures, and the area under the precision-recall curve summarizes performance across decision thresholds. Together, they provide a more informative view than accuracy when uncommon medical outcomes are especially important.
A practical approach is to select a strategy from the available options: class weighting, resampling, synthetic minority examples, or decision-threshold adjustment. The model should then be evaluated with measures suited to uneven class representation, including precision, recall, F1 score, and area under the precision-recall curve. This process focuses assessment on detection of uncommon outcomes rather than majority-class frequency alone.
The issue is especially important when the target outcome is uncommon but consequential, such as a rare disease or an adverse clinical event. In these settings, overlooking minority cases can undermine the usefulness of a prediction model even when its overall accuracy is high. Improving minority-case detection supports more meaningful evaluation of clinical prediction performance.