15.18
During a data collection project, a student is interested in gathering the heights of adult men in their city.
Upon returning to the classroom, the pupils graph the frequency of heights in the sample population. The resulting curve is bell-shaped, with a single peak at the center of which lies the mean.
While a single data point, such as the mean, is crucial to the analysis of these results, so too is the variation. Defined as the dispersion of measurements within a data set, this quantity describes the spread of results, giving a sense of the distance between points.
Additionally, the graph is symmetric, with half of the individuals demonstrating a stature taller than and half shorter than the average. This is referred to as a normal distribution or curve.
To evaluate the variation, they first calculate the range of the results, which is the difference between the highest and lowest heights.
Although the range describes the spread of data, it can be dramatically affected by outliers—like the school’s tallest basketball player—and doesn’t elucidate how measurements are positioned around the mean.
To address this, the student uses an equation to compute a second gauge of variation, termed the standard deviation—the average amount that measurements differ from the mean.
Here, the standard deviation is two-and-a-half inches, so males are—on average—two-and-a-half inches shorter or taller than the mean. Based on the properties of a normal distribution, within this one negative and positive standard deviation, 68% of individuals will fall.
This number will increase to 95% for two standard deviations—here five inches above or below the average height—and 99.7% for three standard deviations. Importantly, the lower the standard deviation, the more tightly results cluster around the mean, which produces a tall and narrow normal curve.
So, if a data set has a small standard deviation, it will have low variation. Thus, means for measurements with low variability are more likely to be a reliable representation of the sample population than those derived from results with high variation, which may be disproportionally affected by outliers.
In the field of psychology, there are several ways to organize measurements of a trait, feature, or characteristic (i.e., variables). Qualitative data…
Copyright © 2026 MyJoVE Corporation. All rights reserved.