15.18
In the field of psychology, there are several ways to organize measurements of a trait, feature, or characteristic (i.e., variables). Qualitative data…
During a data collection project, a student is interested in gathering the heights of adult men in their city.
Upon returning to the classroom, the pupils graph the frequency of heights in the sample population. The resulting curve is bell-shaped, with a single peak at the center of which lies the mean.
While a single data point, such as the mean, is crucial to the analysis of these results, so too is the variation. Defined as the dispersion of measurements within a data set, this quantity describes the spread of results, giving a sense of the distance between points.
Additionally, the graph is symmetric, with half of the individuals demonstrating a stature taller than and half shorter than the average. This is referred to as a normal distribution or curve.
To evaluate the variation, they first calculate the range of the results, which is the difference between the highest and lowest heights.
Although the range describes the spread of data, it can be dramatically affected by outliers—like the school’s tallest basketball player—and doesn’t elucidate how measurements are positioned around the mean.
To address this, the student uses an equation to compute a second gauge of variation, termed the standard deviation—the average amount that measurements differ from the mean.
Here, the standard deviation is two-and-a-half inches, so males are—on average—two-and-a-half inches shorter or taller than the mean. Based on the properties of a normal distribution, within this one negative and positive standard deviation, 68% of individuals will fall.
This number will increase to 95% for two standard deviations—here five inches above or below the average height—and 99.7% for three standard deviations. Importantly, the lower the standard deviation, the more tightly results cluster around the mean, which produces a tall and narrow normal curve.
So, if a data set has a small standard deviation, it will have low variation. Thus, means for measurements with low variability are more likely to be a reliable representation of the sample population than those derived from results with high variation, which may be disproportionally affected by outliers.
View the full transcript and gain access to JoVE Core videos
Q1: What is a normal distribution and why does it matter in data analysis?
A normal distribution is a bell-shaped curve where data clusters symmetrically around the mean, with half the values above and half below the average. This pattern is important because it allows researchers to predict how measurements spread across a population and assess the reliability of the mean as a representative value for the sample.
Q2: How does standard deviation differ from range when measuring data spread?
Range measures only the difference between the highest and lowest values, making it vulnerable to outliers. Standard deviation calculates the average distance all measurements differ from the mean, providing a more comprehensive view of how tightly data clusters around the center and better reflecting true variation in the dataset.
Q3: What does it mean when data has a low standard deviation?
Low standard deviation indicates measurements cluster tightly around the mean, producing a tall and narrow normal curve. This means results have low variation and the mean is a more reliable representation of the sample population, whereas high variation may be disproportionally affected by outliers.
Q4: How do you calculate standard deviation from deviation scores?
First, subtract the mean from each raw score to create deviation scores. Square these deviations to convert negatives to positives, then sum them to get the sum of squares. Divide by the number of data points or degrees of freedom, then take the square root of that result to obtain the standard deviation.
Q5: What percentage of data falls within one standard deviation in a normal distribution?
Within one standard deviation above and below the mean, approximately 68% of individuals fall in a normal distribution. This percentage increases to 95% for two standard deviations and 99.7% for three standard deviations, allowing researchers to predict where most measurements will cluster.
Q6: Why is variance an important step before calculating standard deviation?
Variance estimates the average distance of all scores around the mean by squaring deviations, which converts negative values to positive and prevents them from canceling out. Taking the square root of variance yields standard deviation, which returns the measurement to its original units and makes it more interpretable for describing data spread.
Q7: How can outliers affect the reliability of the range as a measure of variation?
Outliers like an exceptionally tall basketball player can dramatically inflate the range, making it appear that data spreads more widely than it actually does. This misrepresentation makes range a less precise method for measuring variation compared to standard deviation, which accounts for all data points rather than just extremes.