15.18
在心理学领域,存在 测量某一性状、特征或属性的多种方法即,变量)。定性数据(如种族)可整理为频数表,以提供样本或总体中各组比例及多样性的信息。另一方面,研究人员可对定量数据进行更广泛的计算。例如,均值、众数和中位数是用于识别给定数值数据集中变量典型值的集中趋势指标。同样,也有若干方法用于估算数据点之…
在一项数据收集项目中,一名学生希望收集其所在城市成年男性的身高数据。
回到教室后,学生们绘制出样本群体中身高的频率分布图。所得曲线呈钟形,中心处有一个峰值,该峰值即为平均值。
尽管单个数据点(例如均值)对于这些结果的分析至关重要,但变异度同样重要。变异度定义为数据集中测量值的离散程度,该指标描述了结果的分布范围,能够反映数据点之间距离的大小。
此外,该图形呈对称分布,其中一半个体的身高高于平均值,另一半则低于平均值。这种分布被称为正态分布或曲线。
为了评估变异程度,他们首先计算结果的极差,即最高值与最低值之间的差值。
尽管极差可以描述数据的离散程度,但它可能受到异常值(例如该校最高的篮球运动员)的显著影响,且无法说明测量值在均值周围的分布情况。
为解决这一问题,学生使用一个方程来计算另一种变异程度的度量,称为标准差——即测量值与平均值之间的平均差异量。
此处的标准差为2.5英寸,因此男性平均比平均身高矮或高2.5英寸。根据正态分布的特性,在正负一个标准差范围内,将有68%的个体落入该区间。
这个数值在两个标准差范围内将增加到95%——此处为平均身高上下五英寸的范围——在三个标准差范围内则达到99.7%。重要的是,标准差越小,结果越集中在均值附近,从而形成一个高而窄的正态曲线。
因此,如果一个数据集的标准差较小,则其变异程度较低。因此,与来自高变异结果的均值相比,低变异性的测量均值更有可能可靠地代表样本总体,因为高变异结果可能受到离群值的过度影响。
View the full transcript and gain access to JoVE Core videos
Q1: What is a normal distribution and why does it matter in data analysis?
A normal distribution is a bell-shaped curve where data clusters symmetrically around the mean, with half the values above and half below the average. This pattern is important because it allows researchers to predict how measurements spread across a population and assess the reliability of the mean as a representative value for the sample.
Q2: How does standard deviation differ from range when measuring data spread?
Range measures only the difference between the highest and lowest values, making it vulnerable to outliers. Standard deviation calculates the average distance all measurements differ from the mean, providing a more comprehensive view of how tightly data clusters around the center and better reflecting true variation in the dataset.
Q3: What does it mean when data has a low standard deviation?
Low standard deviation indicates measurements cluster tightly around the mean, producing a tall and narrow normal curve. This means results have low variation and the mean is a more reliable representation of the sample population, whereas high variation may be disproportionally affected by outliers.
Q4: How do you calculate standard deviation from deviation scores?
First, subtract the mean from each raw score to create deviation scores. Square these deviations to convert negatives to positives, then sum them to get the sum of squares. Divide by the number of data points or degrees of freedom, then take the square root of that result to obtain the standard deviation.
Q5: What percentage of data falls within one standard deviation in a normal distribution?
Within one standard deviation above and below the mean, approximately 68% of individuals fall in a normal distribution. This percentage increases to 95% for two standard deviations and 99.7% for three standard deviations, allowing researchers to predict where most measurements will cluster.
Q6: Why is variance an important step before calculating standard deviation?
Variance estimates the average distance of all scores around the mean by squaring deviations, which converts negative values to positive and prevents them from canceling out. Taking the square root of variance yields standard deviation, which returns the measurement to its original units and makes it more interpretable for describing data spread.
Q7: How can outliers affect the reliability of the range as a measure of variation?
Outliers like an exceptionally tall basketball player can dramatically inflate the range, making it appear that data spreads more widely than it actually does. This misrepresentation makes range a less precise method for measuring variation compared to standard deviation, which accounts for all data points rather than just extremes.