How to Document Statistical Analyses in a Report or Thesis
Document an analysis so that a reader could repeat it and reach the same numbers: say what data you used, what you excluded and why, which test you ran and why it fits, how you checked its assumptions, and report each result with its test statistic, degrees of freedom, exact p-value, effect size and confidence interval. This post gives a checklist for the methods section and ready-to-adapt reporting examples.
The methods section: what to include
- Research question and hypotheses, stated before the analysis, and which tests are confirmatory and which exploratory.
- Data and sample: where the data came from, the final sample size, and how many cases were excluded at each step, with reasons.
- Variables: how each was measured or coded, including any transformations.
- Tests and justification: which test answers which question and why it fits the data (see choosing a statistical test).
- Assumption checks: how you checked normality, equal variances or independence, and what you did when an assumption failed.
- Significance level and multiple comparisons: the α you used and how you corrected for running many tests.
- Sample-size justification: a power analysis, or an honest statement that the sample size was fixed by circumstance (sample size calculator).
- Software: the program or library and its version, and where readers can find code and data.
Reporting results: examples
These examples follow common APA style: statistics in italics, exact p-values to two or three decimals, p < .001 for very small values, and no leading zero for quantities that cannot exceed 1 (p, r, η²). Every number below comes from the worked examples in our calculators.
- Welch’s t-test: Group A (M = 86.8, SD = 5.07) scored higher than group B (M = 80.0, SD = 3.81), t(7.42) = 2.40, p = .046, d = 1.52, mean difference 6.8, 95% CI [0.17, 13.43].
- One-way ANOVA: Scores differed across the three methods, F(2, 9) = 55.56, p < .001, η² = .93.
- Chi-square test: Recovery was associated with treatment, χ²(1, N = 100) = 16.67, p < .001, V = .41.
- Correlation: Hours studied correlated with exam score, r(3) = .997, p < .001 — with only five people, a result to replicate before trusting.
- Mann–Whitney U: Satisfaction was higher in group A (Mdn = 4) than group B (Mdn = 2), U = 22.5, p = .043, r = .66.
Note that the t-test reports Welch’s fractional degrees of freedom (7.42); say in the methods that you used Welch’s test.
Reproducibility
- Share the data (or a de-identified version) and the analysis code or exact tool settings.
- For simulation or bootstrap methods, report the number of resamples and the random seed.
- Keep a log of every analysis you ran, not just the ones in the final report.
- If you used an online calculator, name it and cite it. Our methodology page has a ready-made citation in APA and BibTeX format.
Common mistakes
- “p = .000”. A p-value is never exactly zero; write p < .001.
- Significance without size. Always report an effect size and its interval; see interpreting effect size.
- Confusing SD and SE. Label which one you report, especially in tables and error bars (SD vs SE).
- “A trend towards significance”. Report the exact p-value and the interval instead.
- Reporting only the tests that worked. Readers need to know how many you ran to interpret any of them.
- Unexplained exclusions. Every dropped case needs a reason stated in advance.
For the formulas behind these tests, keep the statistics formula sheet (also available as a PDF) at hand.