A chi-square test checks whether observed categorical counts differ from what you would expect by chance. The test of independence asks whether two categorical variables are related (e.g. is treatment associated with outcome?); the goodness-of-fit test asks whether one variable matches an expected distribution. It works on raw counts, not percentages, and expects at least about 5 observations per cell. A significant result (p < α) means the variables are associated, or the distribution departs from expectation. Results here are validated against SciPy to full precision.

Test whether observed frequencies match expected frequencies.

Test whether two categorical variables are independent. Enter the contingency table — one row per line, values separated by commas.

When to use a chi-square test

Use the test of independence when you have two categorical variables and want to know if they are related — laid out as a contingency table (e.g. smoker/non-smoker × disease/no-disease). Use the goodness-of-fit test when you have one categorical variable and a theoretical distribution to compare it against (e.g. are die rolls uniform?).

Chi-square is for frequencies of categories. If your outcome is numeric, you want a t-test or ANOVA instead.

Assumptions, and how to check them

Worked example

A 2×2 table of treatment (A/B) by outcome (recovered/not): [[30, 10], [20, 40]]. The recovery rate is 75% under A versus 33% under B, and chi-square returns a large statistic with a small p-value — evidence that treatment and outcome are associated. Enter the table above to get the statistic, df = (rows−1)(cols−1), and p-value.

Common mistakes

If assumptions fail

For small samples or sparse 2×2 tables, use Fisher's exact test. For larger sparse tables, the G-test (likelihood-ratio) is an alternative. Details and validation are on the methodology page.

Guide: Chi-Square Test

The chi-square (χ²) test is a statistical test for categorical data. It compares observed counts to expected counts to determine whether a significant difference exists.

There are two main variants: the goodness-of-fit test (one variable, comparing to a theoretical distribution) and the test of independence (two variables, testing if they are associated).

Goodness of fit: You have one categorical variable and want to test whether the observed frequencies match an expected distribution. Example: "Are the outcomes of a die roll uniformly distributed?"

Test of independence: You have a contingency table with two categorical variables and want to know if they are associated. Example: "Is gender associated with product preference?" Enter the raw cell counts as a matrix.

Cramér's V is an effect size measure for chi-square tests, ranging from 0 (no association) to 1 (perfect association). It corrects for the fact that chi-square grows with sample size, making it a pure measure of effect regardless of n.

Interpretation: small (~0.1), medium (~0.3), large (~0.5). Use Cramér's V alongside the p-value to distinguish statistical significance from practical importance.

  • Categorical data: Variables must be counts/frequencies, not means
  • Independence: Each observation must belong to only one cell
  • Expected frequencies: The rule of thumb is that at least 80% of cells should have an expected count ≥ 5. If this is violated, consider Fisher's Exact Test or combining categories
  • Random sample: Data should come from a random sample