The theoretical probability determines the proportion of observations assigned to a category under the chosen model. Because the calculation multiplies that probability by the total number of observations, changing the model’s probability changes the expected count even when the dataset stays the same. This makes the result a model-based benchmark rather than a description of the observed data.
In a contingency table, row totals and column totals determine how the grand total is distributed across categories. Multiplying the relevant row total by the relevant column total, then dividing by the grand total, produces the table cell’s expected count. The result represents the count implied by those marginal totals under the comparison model.
A difference shows that the data do not match the specified probability model perfectly, but the difference alone does not establish a meaningful departure. Chi-square analysis evaluates whether the overall pattern of discrepancies is consistent with chance or instead suggests an association between categorical variables. Expected frequencies therefore provide the reference point for interpreting observed counts.
The same comparison principle serves two related statistical purposes. In goodness-of-fit analysis, observed category counts are compared with counts predicted by a theoretical probability model. In a contingency table, cell counts are compared with values derived from row and column totals to assess association between categorical variables. The analysis changes according to the structure of the data and the question.
First identify the categories or table cells being analyzed, the total number of observations, and the theoretical probability for each category when using a single-category model. For a contingency table, collect each row total, column total, and the grand total. These inputs establish the appropriate baseline before observed and expected counts are compared.
A basic workflow calculates the expected count for every category or contingency-table cell, places those values alongside the observed counts, and then evaluates the discrepancies through a chi-square test. The resulting comparison addresses whether the departures from the model can reasonably be attributed to chance or whether they indicate a categorical association.