How to Compare Groups With Chi-Square and Nonparametric Tests

By Dackohn · 2026-10-06

When your outcome is a category, use a chi-square test (or Fisher’s exact test for small counts); when it is ordinal or clearly non-normal with a small sample, use a rank-based test such as Mann–Whitney, Wilcoxon or Kruskal–Wallis. Which one depends on two questions: what kind of outcome you have, and whether the groups contain different people or the same people measured more than once.

Choosing the test

Outcome2 independent groups3+ independent groupsSame people, 2 conditions
Categories (yes/no, type)Chi-square or Fisher’s exactChi-squareMcNemar
Ordinal or non-normal numbersMann–Whitney UKruskal–WallisWilcoxon signed-rank
Roughly normal numberst-test (Welch)ANOVAPaired t-test

Unsure? The which-test wizard asks the questions for you.

Categories: the chi-square test

Suppose 40 patients receive treatment A and 60 receive B. Of A, 30 recover; of B, 20 do. Laid out as a 2×2 table (recovered / not by treatment), the chi-square test compares the observed counts with those expected if treatment made no difference — here 20, 20, 30 and 30. The result is χ²(1, N = 100) = 16.67, p < 0.001, with Cramér’s V = 0.41 as the effect size: 75% recovered under A against 33% under B. (With Yates’ continuity correction, common for 2×2 tables, χ² = 15.04; the conclusion is the same.)

Try your own table:

Small counts: the chi-square approximation needs expected counts of about 5 or more in most cells. For the table [7, 1 / 2, 6] the expected counts are only 4.5 and 3.5, so use Fisher’s exact test instead: p = 0.041. And always enter counts, never percentages.

Ordinal or non-normal data: rank-based tests

Rank-based tests replace the values with their ranks, so they work for Likert scores and are robust to outliers and skew.

Mann–Whitney U (two independent groups). Satisfaction ratings of 3, 4, 4, 5, 5 in group A against 1, 2, 2, 3, 4 in group B give U = 22.5 and p ≈ 0.043 using the normal approximation with a tie correction (what our calculator reports). An exact calculation that ignores the ties gives 0.056 — with five values per group and several ties, the evidence is modest either way, which is worth saying in a report. The exact critical values are in the Mann–Whitney U table.

Kruskal–Wallis (three or more groups). On three groups of four scores (12–15, 17–20 and 22–26), H(2) = 9.85, p = 0.007. Like ANOVA, it says at least one group differs; follow up with pairwise comparisons and a multiple-comparisons correction.

Wilcoxon signed-rank and McNemar (same people twice). For before/after scores use Wilcoxon on the paired differences; for before/after yes/no outcomes use McNemar’s test, which looks only at the people whose answer changed.

Report an effect size too

Common mistakes