11.3
Współczynnik korelacji r, opracowany przez Karla Pearsona na początku XX wieku, ma charakter liczbowy i stanowi miarę siły i kierunku liniowego powiąz…
Weź pod uwagę zestaw danych dotyczących poziomu dwutlenku węgla w funkcji rocznej temperatury w określonym okresie. Wykres punktowy punktów danych pokazuje prawdopodobny wzorzec liniowy między dwiema zmiennymi.
Aby potwierdzić wzorzec liniowy, obliczany jest współczynnik korelacji liniowej r.
Najpierw określa się x kwadrat, y kwadrat i iloczyn x i y, a następnie dodaje. Liczba punktów danych wynosi 7.
Na podstawie tych wartości obliczany jest współczynnik korelacji.
Znaczenie wartości współczynnika korelacji można zinterpretować za pomocą tabeli wartości krytycznych.
Przy poziomie istotności 0,05, a n jest równe 7, wartość krytyczna wynosi 0,754.
Ponieważ moduł r jest większy niż wartość krytyczna, istnieją wystarczające dowody na poparcie wniosku, że istnieje korelacja liniowa między zmiennymi.
Wartość r kwadrat wskazuje, że 76,2% zmienności rocznej temperatury można wytłumaczyć liniową zależnością między poziomem dwutlenku węgla a roczną temperaturą.
View the full transcript and gain access to JoVE Core videos
Q1: What does the linear correlation coefficient measure?
The linear correlation coefficient, r, is a numerical measure developed by Karl Pearson that quantifies the strength and direction of the linear association between two variables. It indicates whether variables move together in a predictable straight-line pattern. Values range from -1 to +1, where values closer to -1 or +1 indicate stronger linear relationships, while values near 0 suggest weak or no linear association.
Q2: How do you determine if a correlation coefficient is statistically significant?
Compare the absolute value of the calculated correlation coefficient to the critical value from a significance table at your chosen significance level. If the absolute value of r exceeds the critical value, the correlation is statistically significant. For example, at a 0.05 significance level with 7 data points, a critical value of 0.754 means any |r| greater than 0.754 indicates significant linear correlation between variables.
Q3: What is the coefficient of determination and what does it tell you?
The coefficient of determination, r², is the square of the correlation coefficient expressed as a percentage. It represents the percent of variation in the dependent variable that can be explained by the independent variable using the regression line. For instance, an r² of 76.2% means 76.2% of temperature variation is explained by carbon dioxide levels, while 23.8% remains unexplained.
Q4: What calculations are needed to compute the linear correlation coefficient?
To calculate r, you must first determine x², y², and the product of x and y for each data point, then sum these values. The correlation coefficient formula uses these sums along with the number of data points, n. These intermediate calculations provide the necessary components to quantify the linear relationship strength between your two variables.
Q5: Why is the linear correlation coefficient also called the Pearson product-moment correlation coefficient?
The coefficient is named after Karl Pearson, who developed it in the early 1900s. The term 'product-moment' refers to the calculation method, which involves computing products of paired values and their deviations from the mean. This naming convention distinguishes it from other correlation measures and honors its mathematical foundation.
Q6: What does it mean when the correlation coefficient falls between the critical values?
When r falls between the positive and negative critical values from the significance table, the correlation coefficient is not statistically significant. This means there is insufficient evidence to conclude a true linear relationship exists between the variables. In such cases, using the regression line for prediction is not recommended.
Q7: How does the unexplained variation relate to data scatter around the regression line?
The unexplained variation, calculated as 1 – r² and expressed as a percentage, represents the portion of dependent variable variation not explained by the regression line. This unexplained variation appears as the scattering of observed data points about the best-fit line, which can be examined using residual plots. Greater scatter indicates more unexplained variation and a weaker linear relationship.