Which statistical test should I use? Flowchart from goal to test

Pick a statistical test from your goal, outcome type, groups, pairing and normality: t-tests, ANOVA, rank tests, chi-square, correlation, regression.

Which statistical test should I use? Flowchart from goal to testSTARTCOMPARE GROUPSRELATIONSHIPSPREDICTIONNo, clusteredFit a distributionContinuousOne vs a known valueTwoYes, pairedYesNoNo, independentNormal, similar spreadNormal, unequal or unsureNot normalThree or moreYes, repeatedYesNoNo, independentNormal, similar spreadNormal, unequal or unsureNot normalOrdinalCategorical or binaryIndependentNoYesPaired yes or noTime to an eventNoBoth continuousYesNoContinuousBinaryCountNoYesTime to an eventYesCompare groupsRelationshipPredict an outcomeTwo independentTwo paired3+ independent3+ repeatedCountYesAt least one rankedBoth categoricalOne binary, one continuousChoose a statistical testFor students and analysts: answer afew questions about your goal andyour data, and get a named test, whyit fits and what to check.These are general guidelines fromuniversity statistics guides, not hardrules. The same data can often beanalysed more than one way.First name your outcome (what youmeasured) and what you thinkexplains it. Choose the test beforeyou look at results.Are the observationsindependent, apart fromplanned pairs or repeats?Independent means one observationtells you nothing about another.Pupils in the same class, patients inthe same hospital or repeatedreadings from one machine areclustered.Talk to a statistician about amixed or multilevel modelThe tests below assume independentobservations or a simple paired orrepeated design. Ignoring clusteringusually makes p-values too small.What is your maingoal?Use a chi-squaregoodness-of-fit testFits when you have counts incategories or bins and want to knowif they match expected proportions,such as a fair die or a Poisson model.Check: you need a large enoughsample for the chi-squareapproximation, and the resultdepends on how you bin the data.To test normality of a continuousvariable, the Shapiro-Wilk test isdesigned for that. Anderson-Darlingand Kolmogorov-Smirnov suit othercontinuous distributions.What type of outcome areyou comparing?Continuous means measured on ascale, like weight or time. Ordinalmeans ranked categories, like a 1 to 5rating. Categorical means labels, likeyes or no.How many groupsare you comparing?Use a one-sample t-testTests whether the mean of onesample differs from a hypothesisedvalue, such as a target or a publishednorm.Check: the values should be roughlynormal, or the sample large. If not,use a one-sample median (sign orsigned-rank) test.Are the two sets of valuespaired?Paired means each value in one setmatches one value in the other: thesame people before and after, twins,or matched cases.Are the paired differencesroughly normal?Look at a histogram or Q-Q plot ofthe differences. With a large sample,the mean difference is close tonormal even if the data are not.Use a paired t-testCompares the mean of thewithin-pair differences to zero.Check: run it on the differences, noton the two columns separately.Outliers in the differences can drivethe result.Use the Wilcoxonsigned-rank testThe non-parametric version of thepaired t-test. It ranks the differences,so it suits skewed or ordinaldifferences.Check: it tests whether differencestend to be above or below zero.Report medians, not means.Are both groups roughlynormal, with similar spread?Normal or large samples. For spread,compare standard deviations. If theylook clearly different, treat them asunequal.Use Student's independentt-testThe pooled two-sample t-test. Itassumes both groups have the samevariance.Check: if group sizes differ andspreads differ, this test gives wrongp-values. Welch's t-test is the saferdefault.Use Welch's t-testThe two-sample t-test without theequal-variance assumption. NISTnotes its Welch-Satterthwaitedegrees of freedom are robust tounequal sizes and variances.Check: it still compares means, so itneeds roughly normal data orreasonably large groups.Use the Mann-Whitney UtestAlso called the Wilcoxon rank-sum orWilcoxon-Mann-Whitney test. Thenon-parametric analogue of theindependent t-test.Check: if the two groups have verydifferent shapes or spreads, asignificant result does not simplymean the medians differ.Are the samesubjects measuredin every group?Are the outcomesroughly normal?Use repeated-measuresANOVAOne categorical within-subjectsfactor and a normally distributedoutcome measured at least twice persubject.Check: sphericity (equal variances ofthe differences between conditions).Apply a correction such asGreenhouse-Geisser if it fails.Use the Friedman testFor one within-subjects factor withtwo or more levels when theoutcome is ordinal or not normal.Check: a significant result says theconditions differ somewhere. Followup with pairwise signed-rank testsand a multiple-comparisoncorrection.Are the groupsroughly normal, withsimilar spread?Use one-way ANOVATests whether the means of three ormore independent groups are equal,assuming normal populations.Check: a significant F only says somemeans differ. Use Tukey's HSD for allpairwise comparisons.Use Welch's ANOVAA one-way ANOVA that does notassume equal variances.Check: follow up with theGames-Howell test, which is built forunequal variances and unequalgroup sizes.Use the Kruskal-Wallis testThe non-parametric version ofone-way ANOVA, for an ordinal ornon-normal outcome acrossindependent groups.Check: follow up with pairwise ranktests (for example Dunn's test) andcorrect for multiple comparisons.How are the rankedgroups arranged?Are the groupsindependent orpaired?Are any expected cell countsbelow 5?Expected count for a cell = row totalx column total / grand total. Yoursoftware reports it.Use a chi-square test ofindependenceTests whether two categoricalvariables are related, using acontingency table of counts.Check: it assumes every cell has anexpected value of five or more. Usecounts, not percentages.Use Fisher's exact testUsed instead of chi-square when oneor more cells has a small expectedfrequency. It has no minimum-countassumption.Check: it is most common for 2 x 2tables. Larger tables can be slow tocompute.Use McNemar's testFor two binary outcomes from thesame subjects or matched pairs,such as pass or fail before and aftertraining.Check: only the discordant pairs(changed from yes to no or no toyes) carry the information, so youneed enough of them.Do you need to adjust forother variables?Time-to-event data usually hascensoring: some subjects leave thestudy or the study ends before theirevent happens.Use the log-rank testCompares survival curves(Kaplan-Meier) between groups andhandles censored observations.Check: it tests one grouping variableat a time, can't adjust forconfounders and gives no size ofeffect. Censoring should beunrelated to risk.How are the twovariables measured?Roughly normal with astraight-line pattern?Plot a scatter plot first. Correlationonly measures straight-line (or forSpearman, steadily rising or falling)patterns.Use Pearson correlationMeasures the strength of a linearrelationship between two normallydistributed interval variables.Check: one or two outliers can createor hide a correlation. Correlation isnot causation.Use Spearman rankcorrelationUsed when one or both variables arenot normal or are ordinal. Itcorrelates ranks, so it measures anysteadily increasing or decreasingpattern.Check: ties are common with ordinaldata. Most software corrects forthem, but say how you handled them.What type ofoutcome do youwant to predict?Use linear regressionModels a normally distributedinterval outcome from one or morepredictors, which can be continuousor categorical.Check: the residuals, not the rawoutcome, should be roughly normalwith constant spread. Plot residualsagainst fitted values.Use logistic regressionFor a yes or no outcome coded 0 and1. Reports odds ratios.Check: you need enough events. Acommon rule of thumb is at least 10events per predictor.Is the variance much largerthan the mean?A count is a whole number of events,such as visits or defects. Comparethe variance of the outcome with itsmean within groups.Use Poisson regressionModels count outcomes. It assumesthe conditional variance equals theconditional mean.Check: test for over-dispersion afterfitting. If the variance is larger,standard errors are too small.Use negative binomialregressionGeneralises Poisson regression withan extra parameter forover-dispersion, when the varianceexceeds the mean.Check: if over-dispersion comes frommany extra zeros (two kinds of zero,such as people who never fish andpeople who fished but caughtnothing), consider a zero-inflatedmodel.Use Cox proportionalhazards regressionAdjusts for several risk factors atonce, allows continuous predictorsand reports hazard ratios. It handlesright-censored data.Check: the proportional hazardsassumption, that a predictor's effectis the same early and late infollow-up. Observations must beindependent.

Start

  1. Choose a statistical test

    For students and analysts: answer a few questions about your goal and your data, and get a named test, why it fits and what to check.

    These are general guidelines from university statistics guides, not hard rules. The same data can often be analysed more than one way.

    First name your outcome (what you measured) and what you think explains it. Choose the test before you look at results.

  2. Are the observations independent, apart from planned pairs or repeats?

    Independent means one observation tells you nothing about another. Pupils in the same class, patients in the same hospital or repeated readings from one machine are clustered.

  3. Talk to a statistician about a mixed or multilevel model

    The tests below assume independent observations or a simple paired or repeated design. Ignoring clustering usually makes p-values too small.

  4. What is your main goal?
  5. Use a chi-square goodness-of-fit test

    Fits when you have counts in categories or bins and want to know if they match expected proportions, such as a fair die or a Poisson model.

    Check: you need a large enough sample for the chi-square approximation, and the result depends on how you bin the data.

    To test normality of a continuous variable, the Shapiro-Wilk test is designed for that. Anderson-Darling and Kolmogorov-Smirnov suit other continuous distributions.

Compare groups

  1. What type of outcome are you comparing?

    Continuous means measured on a scale, like weight or time. Ordinal means ranked categories, like a 1 to 5 rating. Categorical means labels, like yes or no.

  2. How many groups are you comparing?
  3. Use a one-sample t-test

    Tests whether the mean of one sample differs from a hypothesised value, such as a target or a published norm.

    Check: the values should be roughly normal, or the sample large. If not, use a one-sample median (sign or signed-rank) test.

  4. Are the two sets of values paired?

    Paired means each value in one set matches one value in the other: the same people before and after, twins, or matched cases.

  5. Are the paired differences roughly normal?

    Look at a histogram or Q-Q plot of the differences. With a large sample, the mean difference is close to normal even if the data are not.

  6. Use a paired t-test

    Compares the mean of the within-pair differences to zero.

    Check: run it on the differences, not on the two columns separately. Outliers in the differences can drive the result.

  7. Use the Wilcoxon signed-rank test

    The non-parametric version of the paired t-test. It ranks the differences, so it suits skewed or ordinal differences.

    Check: it tests whether differences tend to be above or below zero. Report medians, not means.

  8. Are both groups roughly normal, with similar spread?

    Normal or large samples. For spread, compare standard deviations. If they look clearly different, treat them as unequal.

  9. Use Student's independent t-test

    The pooled two-sample t-test. It assumes both groups have the same variance.

    Check: if group sizes differ and spreads differ, this test gives wrong p-values. Welch's t-test is the safer default.

  10. Use Welch's t-test

    The two-sample t-test without the equal-variance assumption. NIST notes its Welch-Satterthwaite degrees of freedom are robust to unequal sizes and variances.

    Check: it still compares means, so it needs roughly normal data or reasonably large groups.

  11. Use the Mann-Whitney U test

    Also called the Wilcoxon rank-sum or Wilcoxon-Mann-Whitney test. The non-parametric analogue of the independent t-test.

    Check: if the two groups have very different shapes or spreads, a significant result does not simply mean the medians differ.

  12. Are the same subjects measured in every group?
  13. Are the outcomes roughly normal?
  14. Use repeated-measures ANOVA

    One categorical within-subjects factor and a normally distributed outcome measured at least twice per subject.

    Check: sphericity (equal variances of the differences between conditions). Apply a correction such as Greenhouse-Geisser if it fails.

  15. Use the Friedman test

    For one within-subjects factor with two or more levels when the outcome is ordinal or not normal.

    Check: a significant result says the conditions differ somewhere. Follow up with pairwise signed-rank tests and a multiple-comparison correction.

  16. Are the groups roughly normal, with similar spread?
  17. Use one-way ANOVA

    Tests whether the means of three or more independent groups are equal, assuming normal populations.

    Check: a significant F only says some means differ. Use Tukey's HSD for all pairwise comparisons.

  18. Use Welch's ANOVA

    A one-way ANOVA that does not assume equal variances.

    Check: follow up with the Games-Howell test, which is built for unequal variances and unequal group sizes.

  19. Use the Kruskal-Wallis test

    The non-parametric version of one-way ANOVA, for an ordinal or non-normal outcome across independent groups.

    Check: follow up with pairwise rank tests (for example Dunn's test) and correct for multiple comparisons.

  20. How are the ranked groups arranged?
  21. Are the groups independent or paired?
  22. Are any expected cell counts below 5?

    Expected count for a cell = row total x column total / grand total. Your software reports it.

  23. Use a chi-square test of independence

    Tests whether two categorical variables are related, using a contingency table of counts.

    Check: it assumes every cell has an expected value of five or more. Use counts, not percentages.

  24. Use Fisher's exact test

    Used instead of chi-square when one or more cells has a small expected frequency. It has no minimum-count assumption.

    Check: it is most common for 2 x 2 tables. Larger tables can be slow to compute.

  25. Use McNemar's test

    For two binary outcomes from the same subjects or matched pairs, such as pass or fail before and after training.

    Check: only the discordant pairs (changed from yes to no or no to yes) carry the information, so you need enough of them.

  26. Do you need to adjust for other variables?

    Time-to-event data usually has censoring: some subjects leave the study or the study ends before their event happens.

  27. Use the log-rank test

    Compares survival curves (Kaplan-Meier) between groups and handles censored observations.

    Check: it tests one grouping variable at a time, can't adjust for confounders and gives no size of effect. Censoring should be unrelated to risk.

Relationships

  1. How are the two variables measured?
  2. Roughly normal with a straight-line pattern?

    Plot a scatter plot first. Correlation only measures straight-line (or for Spearman, steadily rising or falling) patterns.

  3. Use Pearson correlation

    Measures the strength of a linear relationship between two normally distributed interval variables.

    Check: one or two outliers can create or hide a correlation. Correlation is not causation.

  4. Use Spearman rank correlation

    Used when one or both variables are not normal or are ordinal. It correlates ranks, so it measures any steadily increasing or decreasing pattern.

    Check: ties are common with ordinal data. Most software corrects for them, but say how you handled them.

Prediction

  1. What type of outcome do you want to predict?
  2. Use linear regression

    Models a normally distributed interval outcome from one or more predictors, which can be continuous or categorical.

    Check: the residuals, not the raw outcome, should be roughly normal with constant spread. Plot residuals against fitted values.

  3. Use logistic regression

    For a yes or no outcome coded 0 and 1. Reports odds ratios.

    Check: you need enough events. A common rule of thumb is at least 10 events per predictor.

  4. Is the variance much larger than the mean?

    A count is a whole number of events, such as visits or defects. Compare the variance of the outcome with its mean within groups.

  5. Use Poisson regression

    Models count outcomes. It assumes the conditional variance equals the conditional mean.

    Check: test for over-dispersion after fitting. If the variance is larger, standard errors are too small.

  6. Use negative binomial regression

    Generalises Poisson regression with an extra parameter for over-dispersion, when the variance exceeds the mean.

    Check: if over-dispersion comes from many extra zeros (two kinds of zero, such as people who never fish and people who fished but caught nothing), consider a zero-inflated model.

  7. Use Cox proportional hazards regression

    Adjusts for several risk factors at once, allows continuous predictors and reports hazard ratios. It handles right-censored data.

    Check: the proportional hazards assumption, that a predictor's effect is the same early and late in follow-up. Observations must be independent.

Outcomes

Talk to a statistician about a mixed or multilevel model

The tests below assume independent observations or a simple paired or repeated design. Ignoring clustering usually makes p-values too small.

You get here from step 2, Are the observations independent, apart from planned pairs or repeats? (No, clustered).

Use a chi-square goodness-of-fit test

Fits when you have counts in categories or bins and want to know if they match expected proportions, such as a fair die or a Poisson model.

Check: you need a large enough sample for the chi-square approximation, and the result depends on how you bin the data.

To test normality of a continuous variable, the Shapiro-Wilk test is designed for that. Anderson-Darling and Kolmogorov-Smirnov suit other continuous distributions.

You get here from step 4, What is your main goal? (Fit a distribution).

Use a one-sample t-test

Tests whether the mean of one sample differs from a hypothesised value, such as a target or a published norm.

Check: the values should be roughly normal, or the sample large. If not, use a one-sample median (sign or signed-rank) test.

You get here from step 7, How many groups are you comparing? (One vs a known value).

Use a paired t-test

Compares the mean of the within-pair differences to zero.

Check: run it on the differences, not on the two columns separately. Outliers in the differences can drive the result.

You get here from step 10, Are the paired differences roughly normal? (Yes).

Use the Wilcoxon signed-rank test

The non-parametric version of the paired t-test. It ranks the differences, so it suits skewed or ordinal differences.

Check: it tests whether differences tend to be above or below zero. Report medians, not means.

You get here from step 10, Are the paired differences roughly normal? (No), step 25, How are the ranked groups arranged? (Two paired).

Use Student's independent t-test

The pooled two-sample t-test. It assumes both groups have the same variance.

Check: if group sizes differ and spreads differ, this test gives wrong p-values. Welch's t-test is the safer default.

You get here from step 13, Are both groups roughly normal, with similar spread? (Normal, similar spread).

Use Welch's t-test

The two-sample t-test without the equal-variance assumption. NIST notes its Welch-Satterthwaite degrees of freedom are robust to unequal sizes and variances.

Check: it still compares means, so it needs roughly normal data or reasonably large groups.

You get here from step 13, Are both groups roughly normal, with similar spread? (Normal, unequal or unsure).

Use the Mann-Whitney U test

Also called the Wilcoxon rank-sum or Wilcoxon-Mann-Whitney test. The non-parametric analogue of the independent t-test.

Check: if the two groups have very different shapes or spreads, a significant result does not simply mean the medians differ.

You get here from step 13, Are both groups roughly normal, with similar spread? (Not normal), step 25, How are the ranked groups arranged? (Two independent).

Use repeated-measures ANOVA

One categorical within-subjects factor and a normally distributed outcome measured at least twice per subject.

Check: sphericity (equal variances of the differences between conditions). Apply a correction such as Greenhouse-Geisser if it fails.

You get here from step 18, Are the outcomes roughly normal? (Yes).

Use the Friedman test

For one within-subjects factor with two or more levels when the outcome is ordinal or not normal.

Check: a significant result says the conditions differ somewhere. Follow up with pairwise signed-rank tests and a multiple-comparison correction.

You get here from step 18, Are the outcomes roughly normal? (No), step 25, How are the ranked groups arranged? (3+ repeated).

Use one-way ANOVA

Tests whether the means of three or more independent groups are equal, assuming normal populations.

Check: a significant F only says some means differ. Use Tukey's HSD for all pairwise comparisons.

You get here from step 21, Are the groups roughly normal, with similar spread? (Normal, similar spread).

Use Welch's ANOVA

A one-way ANOVA that does not assume equal variances.

Check: follow up with the Games-Howell test, which is built for unequal variances and unequal group sizes.

You get here from step 21, Are the groups roughly normal, with similar spread? (Normal, unequal or unsure).

Use the Kruskal-Wallis test

The non-parametric version of one-way ANOVA, for an ordinal or non-normal outcome across independent groups.

Check: follow up with pairwise rank tests (for example Dunn's test) and correct for multiple comparisons.

You get here from step 21, Are the groups roughly normal, with similar spread? (Not normal), step 25, How are the ranked groups arranged? (3+ independent).

Use a chi-square test of independence

Tests whether two categorical variables are related, using a contingency table of counts.

Check: it assumes every cell has an expected value of five or more. Use counts, not percentages.

You get here from step 27, Are any expected cell counts below 5? (No).

Use Fisher's exact test

Used instead of chi-square when one or more cells has a small expected frequency. It has no minimum-count assumption.

Check: it is most common for 2 x 2 tables. Larger tables can be slow to compute.

You get here from step 27, Are any expected cell counts below 5? (Yes).

Use McNemar's test

For two binary outcomes from the same subjects or matched pairs, such as pass or fail before and after training.

Check: only the discordant pairs (changed from yes to no or no to yes) carry the information, so you need enough of them.

You get here from step 26, Are the groups independent or paired? (Paired yes or no).

Use the log-rank test

Compares survival curves (Kaplan-Meier) between groups and handles censored observations.

Check: it tests one grouping variable at a time, can't adjust for confounders and gives no size of effect. Censoring should be unrelated to risk.

You get here from step 31, Do you need to adjust for other variables? (No).

Use Pearson correlation

Measures the strength of a linear relationship between two normally distributed interval variables.

Check: one or two outliers can create or hide a correlation. Correlation is not causation.

You get here from step 34, Roughly normal with a straight-line pattern? (Yes).

Use Spearman rank correlation

Used when one or both variables are not normal or are ordinal. It correlates ranks, so it measures any steadily increasing or decreasing pattern.

Check: ties are common with ordinal data. Most software corrects for them, but say how you handled them.

You get here from step 34, Roughly normal with a straight-line pattern? (No), step 33, How are the two variables measured? (At least one ranked).

Use linear regression

Models a normally distributed interval outcome from one or more predictors, which can be continuous or categorical.

Check: the residuals, not the raw outcome, should be roughly normal with constant spread. Plot residuals against fitted values.

You get here from step 37, What type of outcome do you want to predict? (Continuous).

Use logistic regression

For a yes or no outcome coded 0 and 1. Reports odds ratios.

Check: you need enough events. A common rule of thumb is at least 10 events per predictor.

You get here from step 37, What type of outcome do you want to predict? (Binary).

Use Poisson regression

Models count outcomes. It assumes the conditional variance equals the conditional mean.

Check: test for over-dispersion after fitting. If the variance is larger, standard errors are too small.

You get here from step 40, Is the variance much larger than the mean? (No).

Use negative binomial regression

Generalises Poisson regression with an extra parameter for over-dispersion, when the variance exceeds the mean.

Check: if over-dispersion comes from many extra zeros (two kinds of zero, such as people who never fish and people who fished but caught nothing), consider a zero-inflated model.

You get here from step 40, Is the variance much larger than the mean? (Yes).

Use Cox proportional hazards regression

Adjusts for several risk factors at once, allows continuous predictors and reports hazard ratios. It handles right-censored data.

Check: the proportional hazards assumption, that a predictor's effect is the same early and late in follow-up. Observations must be independent.

You get here from step 37, What type of outcome do you want to predict? (Time to an event), step 31, Do you need to adjust for other variables? (Yes).