Short answer

A p-value calculator turns a test statistic into a p-value: the probability of a result at least as extreme as yours if the null hypothesis were true. Convertica's free p-value calculator takes a z, t or chi-square statistic, a tail and a significance level, and returns the p-value, the critical value and the decision.

  • Reject the null hypothesis when the p-value is below the significance level you chose before looking at the data. Otherwise you fail to reject it, which is not the same as showing it is true.
  • A p-value is not the probability that the null hypothesis is true, and it says nothing about how large an effect is.
  • z and t tests can be left-tailed, right-tailed or two-tailed. Chi-square tests of fit and independence are right-tailed.
  • The p-value and the critical value are two routes to the same decision.
  • For an A/B test you start from visitors and conversions, not from a z score, and the p-value only holds if the test ran to a sample size planned in advance.

Find the p-value from a test statistic

Free to use, no sign-up. The result updates as you type.

The statistic your test produced.

Keep the minus sign on a negative z or t.

For t and chi-square. A z score needs none.

Match your alternative hypothesis. Chi-square is always right-tailed.

Usually 0.05. Fix it before you look at the data.

Two-tailed p-value

0.0171

Reject the null hypothesis at a significance level of 0.05: the p-value is below 0.05. The result is statistically significant at that level.

Critical values: z = ±1.9600. Your z of 2.384 is outside them, in the rejection region.

If the null hypothesis were true, a z statistic at least this far from 0 in either direction would occur about 1.71% of the time.

All three tails for this statistic: left-tailed 0.9914, right-tailed 0.0086, two-tailed 0.0171.

Formula

left-tailed    p = F(x)
right-tailed   p = 1 - F(x)
two-tailed     p = 2 * min(F(x), 1 - F(x))
chi-square     p = 1 - F(x)     right tail only

decision       reject H0 if p < alpha, else fail to reject H0

x is your test statistic. F is its cumulative distribution when the null hypothesis (H0) is true: the standard normal for z, Student's t with your degrees of freedom for t, and chi-square with your degrees of freedom for chi-square.

How to use the p-value calculator

  1. Pick the test statistic. z for a test based on the normal distribution, t for a t-test, chi-square for a test of fit or independence. Your software output or textbook formula names it.
  2. Enter the statistic. Keep the sign on a z or t score: it shows which side of the null value your result fell on.
  3. Enter the degrees of freedom for t and chi-square. A one-sample t-test has n - 1. A chi-square test on a table has (rows - 1) × (columns - 1).
  4. Choose the tail that matches your alternative hypothesis: two-tailed for "different from", left-tailed for "less than", right-tailed for "greater than".
  5. Enter the significance level (alpha) you fixed before collecting data. 0.05 is the usual choice.
  6. Read the p-value first, then the decision and the critical value under it.

What is a p-value?

A p-value is the probability of getting a test statistic at least as extreme as the one you observed, if the null hypothesis is true and the assumptions of the test hold. A small p-value means your data would be unusual under the null hypothesis, which counts as evidence against it. It describes the data, not the hypothesis.

The null hypothesis (H0) is the claim of no effect or no difference. The alternative hypothesis (Ha) is what you suspect instead. Before you look at the data you choose a significance level, alpha: the share of true null hypotheses you are prepared to reject by mistake. The decision rule is then mechanical:

  • If p is below alpha, reject the null hypothesis. The result is statistically significant at that level.
  • If p is not below alpha, fail to reject the null hypothesis.

The wording is deliberate. "Fail to reject" is not "accept": a large p-value means the data did not supply enough evidence against H0, which also happens when a real effect exists and the sample is too small to show it. And "reject" is not "prove": at alpha 0.05, about 1 in 20 tests of a true null hypothesis is rejected by chance alone.

Four things a p-value does not tell you
A p-value is notWhy
The probability that the null hypothesis is trueIt is calculated on the assumption that the null hypothesis is true, so it cannot also measure how likely that assumption is.
The probability that the result happened by chanceIt is the probability of data this extreme if chance alone were at work. That is not the probability that chance alone was at work.
The size or importance of an effectThe same effect gives a smaller p-value in a larger sample. Read an effect size or a confidence interval for size.
Evidence of no effect when it is largeA p-value above alpha means the test did not find enough evidence. It does not show that nothing is there.

The American Statistical Association says the same in its 2016 statement on p-values: they do not measure the probability that the studied hypothesis is true, and they do not measure the size of an effect or the importance of a result. Its six principles are in the ASA statement on statistical significance and p-values.

How to find the p-value from a test statistic

To find a p-value, calculate the test statistic from your data, identify the distribution it follows when the null hypothesis is true, and measure the area of that distribution at least as extreme as your statistic. The tail you measure comes from the alternative hypothesis. The p-value formula is the same for every test: a tail area.

  1. State the hypotheses. H0 and Ha, and whether Ha is "less than", "greater than" or "not equal".
  2. Fix alpha before you see the result.
  3. Calculate the test statistic with the formula for your test: z, t or chi-square.
  4. Find the tail area beyond the statistic in the matching distribution. That area is the p-value.
  5. Compare p with alpha and state the decision in terms of H0.

The p-value formula for each tail

F is the cumulative distribution of the statistic under H0, and x is the statistic you observed
Alternative hypothesisTestP-value formulaReject H0 when
Less than the null valueLeft-tailedp = F(x)x is below the critical value
Greater than the null valueRight-tailedp = 1 - F(x)x is above the critical value
Not equal to the null valueTwo-tailedp = 2 × min(F(x), 1 - F(x))x is beyond either critical value
Counts do not fit, or two variables are relatedChi-square, right-tailedp = 1 - F(x)x is above the critical value

For z and t, which are symmetric around 0, the two-tailed formula is the same as doubling the area beyond the absolute value of the statistic: p = 2 × (1 - F(|x|)).

How do you find the p-value from a z score?

To find the p-value from a z score, take the area of the standard normal distribution beyond it. For a right-tailed test that is 1 - Phi(z), for a left-tailed test Phi(z), and for a two-tailed test 2 × (1 - Phi(|z|)), where Phi is the standard normal cumulative distribution. No degrees of freedom are needed.

A z score fits when the statistic is a standardized distance and the sample is large enough for the normal distribution to apply: a test of one or two proportions, or a test of a mean when the population standard deviation is known.

Worked example: a z score from an A/B test

An A/B test shows 200 conversions from 10,000 visitors on the control (2.00%) and 250 from 10,000 on the variant (2.50%). The two-proportion z-test gives z = 2.384, the default value in the calculator above. The numbers are made up to show the arithmetic.

  1. H0: the two versions convert at the same rate. Ha: the rates differ, so the test is two-tailed. Alpha is 0.05.
  2. Area beyond the statistic: 1 - Phi(2.384) = 0.008563.
  3. Two tails: p = 2 × 0.008563 = 0.017126, which rounds to 0.0171.
  4. 0.0171 is below 0.05, so reject H0: the difference is statistically significant at the 0.05 level. The critical values are ±1.9600, and 2.384 lies outside them, which is the same decision.

In words: if the two versions really converted at the same rate, a gap this large in either direction would turn up in about 1.71% of tests of this size.

P-values for common z scores. One-tailed is the tail beyond the score; a negative z of the same size gives the same values.
z scoreOne-tailed p-valueTwo-tailed p-value
1.0000.15870.3173
1.2820.09990.1998
1.6450.05000.1000
1.9600.02500.0500
2.0000.02280.0455
2.3260.01000.0200
2.5760.00500.0100
3.0000.00130.0027
3.2910.0004990.000998

How do you find the p-value from a t score?

To find the p-value from a t score, take the area of Student's t distribution beyond it, using the degrees of freedom of your test. The t distribution has heavier tails than the normal, so the same score gives a larger p-value, most noticeably in small samples. As the degrees of freedom grow, t approaches z.

A t score fits when you test a mean, or a difference in means, with the standard deviation estimated from the sample. Degrees of freedom are n - 1 for a one-sample or paired test and n1 + n2 - 2 for a two-sample test with equal variances. Welch's test for unequal variances gives a decimal number of degrees of freedom, which the calculator accepts.

Worked example: a one-sample t-test

A sample of 25 orders has a mean order value of 53.36 and a standard deviation of 8.00. The question is whether the true mean differs from 50.00. The numbers are made up to show the arithmetic.

  1. H0: the mean is 50.00. Ha: it is not, so the test is two-tailed.
  2. Standard error: 8.00 / √25 = 8.00 / 5 = 1.60.
  3. t = (53.36 - 50.00) / 1.60 = 3.36 / 1.60 = 2.10, with 25 - 1 = 24 degrees of freedom.
  4. Area beyond 2.10 in a t distribution with 24 degrees of freedom: 0.023211. Two tails: p = 2 × 0.023211 = 0.046422, which rounds to 0.0464.
  5. At alpha 0.05 the critical values are ±2.0639: reject H0. At alpha 0.01 they are ±2.7969: fail to reject H0.

The same result is significant at one common level and not at the other, which is why alpha has to be chosen before the data are in. Read as a z score, 2.10 would give p = 0.0357: using z where t belongs overstates the evidence. The table shows how much the degrees of freedom matter for this one score.

Two-tailed p-value of t = 2.10 by degrees of freedom, with the critical value and the decision at alpha 0.05
Degrees of freedomTwo-tailed p-valueCritical value (±)Decision at 0.05
20.17054.303Fail to reject
50.08982.571Fail to reject
100.06212.228Fail to reject
240.04642.064Reject
600.03992.000Reject
1200.03781.980Reject
1,0000.03601.962Reject
z (normal)0.03571.960Reject

How do you find the p-value from a chi-square statistic?

To find the p-value from a chi-square statistic, take the area of the chi-square distribution to the right of it, using the degrees of freedom of your test. The statistic sums squared gaps between observed and expected counts, so it is never negative, and a larger value always means a worse fit to the null hypothesis.

That is why chi-square tests of goodness of fit, independence and homogeneity are right-tailed. Gaps in either direction are squared before they are added, so both push the statistic up, and only the right tail holds the evidence against H0. A chi-square test for a single variance is the exception: it can be left-tailed or two-tailed. For that case the calculator also prints the area to the left of your statistic.

Worked example: a chi-square test on three variants

An A/B/C test sends 10,000 visitors to each of three versions. The question is whether conversion depends on the version at all. The numbers are made up to show the arithmetic.

Observed counts, and the counts expected if all three versions converted at the pooled rate of 2.27% (680 of 30,000)
VersionConvertedDid not convertExpected convertedExpected notContribution to chi-square
A (2.00%)2009,800226.679,773.333.2100
B (2.50%)2509,750226.679,773.332.4577
C (2.30%)2309,770226.679,773.330.0502
  1. H0: the conversion rate is the same for all three versions. Ha: at least one differs.
  2. Each cell adds (observed - expected)² / expected. Summed over the six cells, chi-square = 5.718.
  3. Degrees of freedom: (2 - 1) × (3 - 1) = 2.
  4. Area to the right of 5.718 with 2 degrees of freedom: p = 0.0573. With 2 degrees of freedom this tail has a simple form you can check by hand, e raised to minus half the statistic: exp(-2.8590) = 0.057326.
  5. 0.0573 is not below 0.05, so fail to reject H0. The critical value is 5.991, and 5.718 is not above it.

Versions A and B here are the two groups from the z score example, where the two of them alone gave p = 0.0171. Once a third version is in the test, the evidence that the versions differ at all is weaker. Picking out the best-looking pair after the fact, and testing only that, is how false winners are made.

Chi-square critical values, right tail: reject H0 when your statistic is above the value
Degrees of freedomAlpha 0.10Alpha 0.05Alpha 0.01
12.7063.8416.635
24.6055.9919.210
36.2517.81511.345
47.7799.48813.277
59.23611.07015.086
610.64512.59216.812
1015.98718.30723.209

One-tailed or two-tailed: which p-value do you need?

Use a two-tailed p-value when the alternative hypothesis is "different from", and a one-tailed p-value only when it named one direction before the data were collected. When the statistic falls on the predicted side, the one-tailed p-value is half the two-tailed one. When it falls on the other side, the one-tailed p-value is above 0.5.

That halving is the reason to decide first. Switching to one tail after seeing which way the result went turns a 0.05 test into a 0.10 test without saying so. A two-tailed test also catches a result in the direction you did not expect, which in practice is often the more useful thing to know.

P-value or critical value: do they give the same answer?

Yes. The p-value approach and the critical value approach always give the same decision, because the critical value is the test statistic whose p-value equals alpha exactly. A statistic beyond the critical value has a p-value below alpha, and the reverse. The p-value adds one thing: how far past, or short of, the line the result is.

Critical values are what printed tables give you. If you have to work from a table, NIST publishes critical values of Student's t and critical values of chi-square. A table brackets the p-value between two columns. The calculator gives it exactly.

Get your free CRO audit

Enter your website and email and the audit starts right away. Watch it check your page live: eight checks, each scored out of 100, and three fixes you can make now.

What does a p-value tell you about an A/B test?

In an A/B test, the p-value tells you how often a gap as large as yours would appear if the two versions truly converted at the same rate. A small p-value is a reason to doubt "no difference". It is not the chance that the variant is better, and it is only as sound as the test behind it.

You rarely start from a z score. A testing tool gives you visitors and conversions for each version, and the statistical significance calculator turns those into the z score, the p-value, the lift and a confidence interval for the difference. Use this page when you already have the statistic, or to check a number in a report.

Three things a p-value from an A/B test cannot tell you:

  • How large the lift is. A tiny lift on a very large sample can have a smaller p-value than a large lift on a small one. Read the interval, not only the verdict.
  • Whether the lift is worth it. That depends on what the change costs to build and keep, and on what it does to revenue per visitor, not on the p-value.
  • Whether the test was sound. A broken traffic split, a tracking fault or a test stopped early produces a p-value that looks exactly like a good one.

The peeking problem: why the first p below 0.05 is not the answer

A p-value from a fixed-sample test is valid for one look at the data, at a sample size chosen in advance. While a test is running, its p-value drifts up and down as visitors arrive. If you check it every day and stop the first time it dips below 0.05, you give chance many opportunities to cross the line, and the share of false winners rises well above 5%.

  1. Plan the sample first. Work out how many visitors each version needs with the A/B test sample size calculator.
  2. Run to that number, over whole weeks. The guide on how long to run an A/B test covers the timing and when stopping early is justified.
  3. Read the p-value once, at the end. If you need to monitor results as they come in, use a sequential method built for that and read its own output.

Convertica's A/B testing guide walks through the whole cycle, from hypothesis to reading the result. A test also only answers the question you put to it. The free CRO audit is where the questions worth testing come from: it checks a page in eight areas and ranks what to fix first.

What assumptions does a p-value depend on?

A p-value is only as good as the model behind it. It assumes the null hypothesis, and it also assumes the conditions that make the statistic follow its distribution. If those conditions fail, the number is still printed, but it no longer means what it claims. The calculator converts a statistic. It cannot check how the statistic was produced.

  • Every test: observations are independent, and the sample was drawn or assigned at random.
  • z tests: the sampling distribution is close to normal. For proportions that needs enough successes and failures in each group; for a mean it needs a known population standard deviation, or a large sample.
  • t tests: the data are roughly normal, or the sample is large enough for the mean to behave that way. A few extreme values in a small sample can mislead the test.
  • Chi-square tests: the cells hold counts, not percentages, each observation falls in one cell, and expected counts are not too small. A common rule of thumb is at least 5 in every cell.
  • One planned test: the hypothesis, the tail, alpha and the sample size were set in advance. Running many tests and reporting the best one breaks the meaning of every p-value involved.

How to find a p-value in Excel or Google Sheets

Excel and Google Sheets have a function for each tail area, so you can find a p-value from a test statistic without a table. Put the statistic in A1 and, for t and chi-square, the degrees of freedom in B1. The formulas below work in both.

Spreadsheet formulas: statistic in A1, degrees of freedom in B1
StatisticTailFormula
z scoreTwo-tailed=2*(1-NORM.DIST(ABS(A1),0,1,TRUE))
z scoreRight-tailed=1-NORM.DIST(A1,0,1,TRUE)
z scoreLeft-tailed=NORM.DIST(A1,0,1,TRUE)
t scoreTwo-tailed=T.DIST.2T(ABS(A1),B1)
t scoreRight-tailed=T.DIST.RT(A1,B1)
t scoreLeft-tailed=T.DIST(A1,B1,TRUE)
Chi-squareRight-tailed=CHISQ.DIST.RT(A1,B1)

T.DIST.2T needs a positive number, which is what ABS is for. For critical values, use =T.INV.2T(0.05,B1) for a two-tailed t test and =CHISQ.INV.RT(0.05,B1) for chi-square, with your own alpha in place of 0.05. Microsoft's documentation says Excel cuts a decimal number of degrees of freedom down to a whole number in these functions, so a Welch test can come out slightly different there than in this calculator, which uses the decimal value.

A p-value answers one question about a result. These tools answer the ones on either side of it.

See all free tools and how they fit together.

P-value calculator FAQ

What does a p-value calculator do?

A p-value calculator converts a test statistic into a p-value. You give it a z score, a t score with its degrees of freedom or a chi-square statistic with its degrees of freedom, and say which tail the test uses. It returns the area of the distribution at least as extreme as your statistic, which is the p-value.

How do you calculate a p-value?

Work out the test statistic from your data, pick the distribution it follows if the null hypothesis is true (normal, t or chi-square), and find the area in the tail or tails beyond the statistic. For a two-tailed z test that is 2 × (1 - Phi(|z|)). By hand you read the area from a table; a calculator or a spreadsheet gives it exactly.

What does a p-value of 0.04 mean?

A p-value of 0.04 means that if the null hypothesis were true, a result at least as extreme as yours would occur about 4% of the time. It is below 0.05, so you reject the null hypothesis at the 0.05 level, but not at the 0.01 level. It does not mean a 4% chance that the null hypothesis is true or a 96% chance that the effect is real.

Is a p-value of 0.005 significant?

Yes, at the two most common levels: 0.005 is below both 0.05 and 0.01, so the result is statistically significant at either. It is not below 0.001. Whether it matters is a separate question: a p-value measures how surprising the data would be under the null hypothesis, not how large or useful the effect is.

How do I find the p-value from a z score?

Look up the standard normal area beyond the z score. Right-tailed: p = 1 - Phi(z). Left-tailed: p = Phi(z). Two-tailed: p = 2 × (1 - Phi(|z|)). For example, z = 1.96 gives a two-tailed p-value of 0.0500 and a one-tailed p-value of 0.0250.

How do I find the p-value from a t score?

You need the t score and its degrees of freedom, because the t distribution changes shape with the sample size. The p-value is the area of that t distribution beyond the score, doubled for a two-tailed test. For example, t = 2.10 with 24 degrees of freedom gives a two-tailed p-value of 0.0464; the same score read as a z score would give 0.0357.

Why is the chi-square p-value right-tailed?

A chi-square statistic adds up squared gaps between observed and expected counts, so it can only grow as the data move away from the null hypothesis, whatever the direction of the gaps. Only a large value counts as evidence against the null hypothesis, so the p-value is the area to the right of the statistic.

Should I use a one-tailed or a two-tailed p-value?

Use a two-tailed p-value unless your alternative hypothesis named one direction before you saw the data. A one-tailed p-value is half the two-tailed one when the statistic falls on the side you predicted, which is why choosing the tail after seeing the result makes false positives more likely than your significance level says.

Is the p-value the same as the z score?

No. The z score is the test statistic: how many standard errors your result is from the value the null hypothesis expects. The p-value is the probability that goes with it. A bigger z score in either direction means a smaller p-value: a two-tailed test at the 0.05 level rejects when z is beyond ±1.96.

Can a p-value be 0, negative or greater than 1?

No. A p-value is a probability, so it lies between 0 and 1. A very large test statistic gives a p-value too small to print with four decimals, which is why this calculator shows "below 0.0001" and then writes the value out in scientific notation. If a formula gives a negative p-value or one above 1, the tail has been mixed up.

How do I calculate a p-value in Excel or Google Sheets?

For a z score in A1, use =2*(1-NORM.DIST(ABS(A1),0,1,TRUE)) for a two-tailed p-value. For a t score in A1 with degrees of freedom in B1, use =T.DIST.2T(ABS(A1),B1), or =T.DIST.RT(A1,B1) for the right tail. For a chi-square statistic, use =CHISQ.DIST.RT(A1,B1). The same formulas work in Google Sheets.

Can I use this p-value calculator for an A/B test?

Yes, if you already have the z score. Most people do not: an A/B test gives you visitors and conversions for each version. Convertica's statistical significance calculator starts from those counts, works out the z score for you and reports the p-value, the lift and a confidence interval for the difference.

Find out what is costing you conversions

Enter your website and email and the audit starts right away. Watch it check your page live: eight checks, each scored out of 100, and three fixes you can make now.