How to Compare Groups With Chi-Square and Nonparametric Tests
When your outcome is a category, use a chi-square test (or Fisher’s exact test for small counts); when it is ordinal or clearly non-normal with a small sample, use a rank-based test such as Mann–Whitney, Wilcoxon or Kruskal–Wallis. Which one depends on two questions: what kind of outcome you have, and whether the groups contain different people or the same people measured more than once.
Choosing the test
| Outcome | 2 independent groups | 3+ independent groups | Same people, 2 conditions |
|---|---|---|---|
| Categories (yes/no, type) | Chi-square or Fisher’s exact | Chi-square | McNemar |
| Ordinal or non-normal numbers | Mann–Whitney U | Kruskal–Wallis | Wilcoxon signed-rank |
| Roughly normal numbers | t-test (Welch) | ANOVA | Paired t-test |
Unsure? The which-test wizard asks the questions for you.
Categories: the chi-square test
Suppose 40 patients receive treatment A and 60 receive B. Of A, 30 recover; of B, 20 do. Laid out as a 2×2 table (recovered / not by treatment), the chi-square test compares the observed counts with those expected if treatment made no difference — here 20, 20, 30 and 30. The result is χ²(1, N = 100) = 16.67, p < 0.001, with Cramér’s V = 0.41 as the effect size: 75% recovered under A against 33% under B. (With Yates’ continuity correction, common for 2×2 tables, χ² = 15.04; the conclusion is the same.)
Try your own table:
Small counts: the chi-square approximation needs expected counts of about 5 or more in most cells. For the table [7, 1 / 2, 6] the expected counts are only 4.5 and 3.5, so use Fisher’s exact test instead: p = 0.041. And always enter counts, never percentages.
Ordinal or non-normal data: rank-based tests
Rank-based tests replace the values with their ranks, so they work for Likert scores and are robust to outliers and skew.
Mann–Whitney U (two independent groups). Satisfaction ratings of 3, 4, 4, 5, 5 in group A against 1, 2, 2, 3, 4 in group B give U = 22.5 and p ≈ 0.043 using the normal approximation with a tie correction (what our calculator reports). An exact calculation that ignores the ties gives 0.056 — with five values per group and several ties, the evidence is modest either way, which is worth saying in a report. The exact critical values are in the Mann–Whitney U table.
Kruskal–Wallis (three or more groups). On three groups of four scores (12–15, 17–20 and 22–26), H(2) = 9.85, p = 0.007. Like ANOVA, it says at least one group differs; follow up with pairwise comparisons and a multiple-comparisons correction.
Wilcoxon signed-rank and McNemar (same people twice). For before/after scores use Wilcoxon on the paired differences; for before/after yes/no outcomes use McNemar’s test, which looks only at the people whose answer changed.
Report an effect size too
- Chi-square: Cramér’s V (0.41 above), or the risk difference and odds ratio for 2×2 tables.
- Rank tests: a rank-based r; our Mann–Whitney calculator reports r = Z/√N (0.66 in the example).
Common mistakes
- Running chi-square on percentages instead of counts.
- Using Mann–Whitney on paired data — use Wilcoxon signed-rank instead.
- Calling Mann–Whitney a test of medians without checking that the two distributions have a similar shape.
- Reaching for nonparametric tests by default. With reasonably large, roughly symmetric samples, the t-test and ANOVA are more powerful.