Start
- Choose a statistical test
For students and analysts: answer a few questions about your goal and your data, and get a named test, why it fits and what to check.
These are general guidelines from university statistics guides, not hard rules. The same data can often be analysed more than one way.
First name your outcome (what you measured) and what you think explains it. Choose the test before you look at results.
- Are the observations independent, apart from planned pairs or repeats?
Independent means one observation tells you nothing about another. Pupils in the same class, patients in the same hospital or repeated readings from one machine are clustered.
- No, clustered: go to step 3, Talk to a statistician about a mixed or multilevel model
- Yes: go to step 4, What is your main goal?
- Talk to a statistician about a mixed or multilevel model
The tests below assume independent observations or a simple paired or repeated design. Ignoring clustering usually makes p-values too small.
- What is your main goal?
- Fit a distribution: go to step 5, Use a chi-square goodness-of-fit test
- Compare groups: go to step 6, What type of outcome are you comparing?
- Relationship: go to step 33, How are the two variables measured?
- Predict an outcome: go to step 37, What type of outcome do you want to predict?
- Use a chi-square goodness-of-fit test
Fits when you have counts in categories or bins and want to know if they match expected proportions, such as a fair die or a Poisson model.
Check: you need a large enough sample for the chi-square approximation, and the result depends on how you bin the data.
To test normality of a continuous variable, the Shapiro-Wilk test is designed for that. Anderson-Darling and Kolmogorov-Smirnov suit other continuous distributions.
Compare groups
- What type of outcome are you comparing?
Continuous means measured on a scale, like weight or time. Ordinal means ranked categories, like a 1 to 5 rating. Categorical means labels, like yes or no.
- Continuous: go to step 7, How many groups are you comparing?
- Ordinal: go to step 25, How are the ranked groups arranged?
- Categorical or binary: go to step 26, Are the groups independent or paired?
- Time to an event: go to step 31, Do you need to adjust for other variables?
- Count: go to step 40, Is the variance much larger than the mean?
- How many groups are you comparing?
- One vs a known value: go to step 8, Use a one-sample t-test
- Two: go to step 9, Are the two sets of values paired?
- Three or more: go to step 17, Are the same subjects measured in every group?
- Use a one-sample t-test
Tests whether the mean of one sample differs from a hypothesised value, such as a target or a published norm.
Check: the values should be roughly normal, or the sample large. If not, use a one-sample median (sign or signed-rank) test.
- Are the two sets of values paired?
Paired means each value in one set matches one value in the other: the same people before and after, twins, or matched cases.
- Yes, paired: go to step 10, Are the paired differences roughly normal?
- No, independent: go to step 13, Are both groups roughly normal, with similar spread?
- Are the paired differences roughly normal?
Look at a histogram or Q-Q plot of the differences. With a large sample, the mean difference is close to normal even if the data are not.
- Yes: go to step 11, Use a paired t-test
- No: go to step 12, Use the Wilcoxon signed-rank test
- Use a paired t-test
Compares the mean of the within-pair differences to zero.
Check: run it on the differences, not on the two columns separately. Outliers in the differences can drive the result.
- Use the Wilcoxon signed-rank test
The non-parametric version of the paired t-test. It ranks the differences, so it suits skewed or ordinal differences.
Check: it tests whether differences tend to be above or below zero. Report medians, not means.
- Are both groups roughly normal, with similar spread?
Normal or large samples. For spread, compare standard deviations. If they look clearly different, treat them as unequal.
- Normal, similar spread: go to step 14, Use Student's independent t-test
- Normal, unequal or unsure: go to step 15, Use Welch's t-test
- Not normal: go to step 16, Use the Mann-Whitney U test
- Use Student's independent t-test
The pooled two-sample t-test. It assumes both groups have the same variance.
Check: if group sizes differ and spreads differ, this test gives wrong p-values. Welch's t-test is the safer default.
- Use Welch's t-test
The two-sample t-test without the equal-variance assumption. NIST notes its Welch-Satterthwaite degrees of freedom are robust to unequal sizes and variances.
Check: it still compares means, so it needs roughly normal data or reasonably large groups.
- Use the Mann-Whitney U test
Also called the Wilcoxon rank-sum or Wilcoxon-Mann-Whitney test. The non-parametric analogue of the independent t-test.
Check: if the two groups have very different shapes or spreads, a significant result does not simply mean the medians differ.
- Are the same subjects measured in every group?
- Yes, repeated: go to step 18, Are the outcomes roughly normal?
- No, independent: go to step 21, Are the groups roughly normal, with similar spread?
- Are the outcomes roughly normal?
- Yes: go to step 19, Use repeated-measures ANOVA
- No: go to step 20, Use the Friedman test
- Use repeated-measures ANOVA
One categorical within-subjects factor and a normally distributed outcome measured at least twice per subject.
Check: sphericity (equal variances of the differences between conditions). Apply a correction such as Greenhouse-Geisser if it fails.
- Use the Friedman test
For one within-subjects factor with two or more levels when the outcome is ordinal or not normal.
Check: a significant result says the conditions differ somewhere. Follow up with pairwise signed-rank tests and a multiple-comparison correction.
- Are the groups roughly normal, with similar spread?
- Normal, similar spread: go to step 22, Use one-way ANOVA
- Normal, unequal or unsure: go to step 23, Use Welch's ANOVA
- Not normal: go to step 24, Use the Kruskal-Wallis test
- Use one-way ANOVA
Tests whether the means of three or more independent groups are equal, assuming normal populations.
Check: a significant F only says some means differ. Use Tukey's HSD for all pairwise comparisons.
- Use Welch's ANOVA
A one-way ANOVA that does not assume equal variances.
Check: follow up with the Games-Howell test, which is built for unequal variances and unequal group sizes.
- Use the Kruskal-Wallis test
The non-parametric version of one-way ANOVA, for an ordinal or non-normal outcome across independent groups.
Check: follow up with pairwise rank tests (for example Dunn's test) and correct for multiple comparisons.
- How are the ranked groups arranged?
- Two independent: go to step 16, Use the Mann-Whitney U test
- Two paired: go to step 12, Use the Wilcoxon signed-rank test
- 3+ independent: go to step 24, Use the Kruskal-Wallis test
- 3+ repeated: go to step 20, Use the Friedman test
- Are the groups independent or paired?
- Independent: go to step 27, Are any expected cell counts below 5?
- Paired yes or no: go to step 30, Use McNemar's test
- Are any expected cell counts below 5?
Expected count for a cell = row total x column total / grand total. Your software reports it.
- No: go to step 28, Use a chi-square test of independence
- Yes: go to step 29, Use Fisher's exact test
- Use a chi-square test of independence
Tests whether two categorical variables are related, using a contingency table of counts.
Check: it assumes every cell has an expected value of five or more. Use counts, not percentages.
- Use Fisher's exact test
Used instead of chi-square when one or more cells has a small expected frequency. It has no minimum-count assumption.
Check: it is most common for 2 x 2 tables. Larger tables can be slow to compute.
- Use McNemar's test
For two binary outcomes from the same subjects or matched pairs, such as pass or fail before and after training.
Check: only the discordant pairs (changed from yes to no or no to yes) carry the information, so you need enough of them.
- Do you need to adjust for other variables?
Time-to-event data usually has censoring: some subjects leave the study or the study ends before their event happens.
- No: go to step 32, Use the log-rank test
- Yes: go to step 43, Use Cox proportional hazards regression
- Use the log-rank test
Compares survival curves (Kaplan-Meier) between groups and handles censored observations.
Check: it tests one grouping variable at a time, can't adjust for confounders and gives no size of effect. Censoring should be unrelated to risk.
Relationships
- How are the two variables measured?
- Both continuous: go to step 34, Roughly normal with a straight-line pattern?
- At least one ranked: go to step 36, Use Spearman rank correlation
- Both categorical: go to step 27, Are any expected cell counts below 5?
- One binary, one continuous: go to step 7, How many groups are you comparing?
- Roughly normal with a straight-line pattern?
Plot a scatter plot first. Correlation only measures straight-line (or for Spearman, steadily rising or falling) patterns.
- Yes: go to step 35, Use Pearson correlation
- No: go to step 36, Use Spearman rank correlation
- Use Pearson correlation
Measures the strength of a linear relationship between two normally distributed interval variables.
Check: one or two outliers can create or hide a correlation. Correlation is not causation.
- Use Spearman rank correlation
Used when one or both variables are not normal or are ordinal. It correlates ranks, so it measures any steadily increasing or decreasing pattern.
Check: ties are common with ordinal data. Most software corrects for them, but say how you handled them.
Prediction
- What type of outcome do you want to predict?
- Continuous: go to step 38, Use linear regression
- Binary: go to step 39, Use logistic regression
- Count: go to step 40, Is the variance much larger than the mean?
- Time to an event: go to step 43, Use Cox proportional hazards regression
- Use linear regression
Models a normally distributed interval outcome from one or more predictors, which can be continuous or categorical.
Check: the residuals, not the raw outcome, should be roughly normal with constant spread. Plot residuals against fitted values.
- Use logistic regression
For a yes or no outcome coded 0 and 1. Reports odds ratios.
Check: you need enough events. A common rule of thumb is at least 10 events per predictor.
- Is the variance much larger than the mean?
A count is a whole number of events, such as visits or defects. Compare the variance of the outcome with its mean within groups.
- No: go to step 41, Use Poisson regression
- Yes: go to step 42, Use negative binomial regression
- Use Poisson regression
Models count outcomes. It assumes the conditional variance equals the conditional mean.
Check: test for over-dispersion after fitting. If the variance is larger, standard errors are too small.
- Use negative binomial regression
Generalises Poisson regression with an extra parameter for over-dispersion, when the variance exceeds the mean.
Check: if over-dispersion comes from many extra zeros (two kinds of zero, such as people who never fish and people who fished but caught nothing), consider a zero-inflated model.
- Use Cox proportional hazards regression
Adjusts for several risk factors at once, allows continuous predictors and reports hazard ratios. It handles right-censored data.
Check: the proportional hazards assumption, that a predictor's effect is the same early and late in follow-up. Observations must be independent.