Significance Testing
You measured something and got a number. Was it a real effect or just random noise? Hypothesis testing is a fixed, mechanical procedure for answering that question. Once you know the six steps, every question in this unit is the same six steps with different numbers.
We always assume nothing is happening (the null hypothesis H₀). We then ask: "if H₀ were true, how unlikely would it be to see data this extreme?" If it is very unlikely, we reject H₀ and call it a real effect. If it is not unlikely, we have found nothing — and that is a perfectly valid scientific answer, not a failure.
1. Null and Alternate Hypothesis
H₀ — the Null Hypothesis
The default position: nothing unusual is happening.
- Contains an = sign (usually)
- It is the easy one to calculate — the maths works out cleanly
- For a mean: μ = μ₀ (the population mean equals the stated value)
- For a proportion: p = p₀
H₁ — the Alternate Hypothesis
What the researcher hopes to prove.
- Uses >, < or ≠
- ≠ means two-tailed. > or < means one-tailed
- Never called "accepting H₀"
Failing to find evidence against H₀ does not prove H₀ is true. It only means the data was not strong enough. Always write "fail to reject H₀". This is a standard marking point and is frequently the difference between a full-mark answer and a half-mark one.
Read the question's verb.
- "has it increased?", "is it more than?", "did it rise?" → H₁: μ > μ₀ (one-tailed, right)
- "has it decreased?", "is it less than?", "did it fall?" → H₁: μ < μ₀ (one-tailed, left)
- "has it changed?", "is it different?", "does it agree?" → H₁: μ ≠ μ₀ (two-tailed)
If the question does not say which direction, it is two-tailed. When in doubt, choose two-tailed — it is the safer default.
2. Level of Significance, Critical Region and Critical Values
Common values: α = 0.10 (10%), 0.05 (5%), 0.01 (1%)
| Term | Definition |
|---|---|
| Critical region | The set of values that make you reject H₀. Also called the rejection region. |
| Critical value | The boundary number between the critical region and the acceptance region. |
| Acceptance region | Where you fail to reject H₀. |
| Test statistic | The single number you calculate (z or t) and compare with the critical value. |
Students swap these constantly. Region = a range of numbers (an area on the curve, or x > 1.645). Value = one number (1.645). If the question says "find the critical values", it wants numbers. If it says "state the critical region", it wants a range.
3. One-Tailed vs Two-Tailed
| One-tailed | Two-tailed | |
|---|---|---|
| Used when | you predicted a direction in advance | any change in either direction counts |
| H₁ looks like | μ > μ₀ or μ < μ₀ | μ ≠ μ₀ |
| Shaded area | all of α in one tail | α split into two halves |
| Critical value at α = 0.05 | 1.645 | 1.96 |
| Harder to reject H₀? | Yes — all evidence is in one direction | No — twice the area helps you |
| Key words | "greater than", "less than", "increase", "decrease" | "different", "changed", "differs" |
At α = 0.05, the critical value is 1.645 for one tail and 1.96 for two tails.
Shortcut: the two-tailed value is always the larger number. If you write 1.96 for a one-tailed test your answer is wrong, and vice versa. Writing the wrong one is the single most common error in this unit.
Rule of thumb: 1.645 is about 1.6, 1.96 is about 2. If you cannot remember which is which, remember two tails ⇒ 2.
4. Standard Critical Values
| Level α | One-tailed | Two-tailed | ||
|---|---|---|---|---|
| Critical value | Area in the tail | Critical value | Area in each tail | |
| 0.10 | 1.282 | 0.10 | 1.645 | 0.05 |
| 0.05 | 1.645 | 0.05 | 1.96 | 0.025 |
| 0.025 | 1.960 | 0.025 | 2.241 | 0.0125 |
| 0.01 | 2.326 | 0.01 | 2.576 | 0.005 |
| 0.005 | 2.576 | 0.005 | 2.807 | 0.0025 |
| 0.001 | 3.090 | 0.001 | 3.291 | 0.0005 |
1. Two-tailed at α is the same number as one-tailed at α/2. Two-tailed α = 0.05 gives 1.96, and one-tailed α = 0.025 also gives 1.96. Check the two-tailed row: it is exactly the one-tailed row shifted down one line.
2. Bigger α ⇒ smaller critical value. α = 0.01 (strict) needs a bigger z to reject than α = 0.10 (lenient). If your computed z is larger than the critical value, that is fine — it just means very strong evidence.
t-distribution critical values (small samples)
When the population standard deviation is unknown, or the sample is small (n < 30), you use the t-distribution instead of z. The shape is the same bell curve but the tails are fatter, so the critical values are larger.
| df (n−1) | Two-tailed α = 0.05 | One-tailed α = 0.05 | ||
|---|---|---|---|---|
| t | z (for comparison) | t | z | |
| 1 | 12.71 | 1.96 | 6.314 | 1.645 |
| 2 | 4.303 | 2.920 | ||
| 5 | 2.571 | 2.015 | ||
| 10 | 2.228 | 1.812 | ||
| 20 | 2.086 | 1.725 | ||
| 30 | 2.042 | 1.697 | ||
| ∞ (large n) | 1.960 | 1.645 | ||
df = n − 1 for a single sample.
Why n − 1? Because the sample mean already "uses up" one degree of freedom. It is a standard result — just apply it. As n grows, df grows and t → z, so with a big sample the two tests give the same answer anyway.
5. Test for a Single Mean
Step-by-step: the z-test for a mean
Read the fraction as "x̄ is this many standard errors away from μ₀". The quantity σ/√n is the standard error — it is the standard deviation of the sample mean, not of the data.
Question: A machine is supposed to fill bottles with 500 ml. A sample of 36 bottles has mean 495 ml. The population SD is 6 ml. Test at the 5% level whether the machine is filling incorrectly.
Step 1 · Hypotheses. "Incorrectly" means a change in either direction.
H₁: μ ≠ 500 (machine is filling incorrectly)
Step 2 · Significance level and type. α = 0.05, two-tailed (because ≠).
Step 3 · Critical value. Two-tailed, α = 0.05 → ±1.96.
Step 4 · Standard error and test statistic.
z = (495 − 500)/1 = −5.0
Step 5 · Decision. |−5.0| > 1.96, so z lies in a critical region.
Step 6 · Conclusion in context.
Question: A new battery claims to last more than 100 hours. A sample of 16 batteries has mean 108 hours with SD 12 hours. Test at the 5% level.
Step 1 · Hypotheses. "more than" → right-tailed, one-tailed.
H₁: μ > 100
Step 2 · Test type. σ unknown and n = 16 < 30 → t-test, one-tailed, α = 0.05.
Step 3 · Critical value. df = n − 1 = 15. One-tailed α = 0.05, df = 15 → t = 1.753.
Step 4 · Test statistic.
t = (108 − 100)/3 = 8/3 = 2.667
Step 5 · Decision. 2.667 > 1.753 → in the critical region.
Question: A sample of 25 students has mean score 72 with SD 10. The exam board claims the mean is 75. Test at 5%.
σ unknown, n = 25 < 30 → t-test, df = 24
Two-tailed α = 0.05, df = 24 → t = 2.064
SE = 10/√25 = 10/5 = 2
t = (72 − 75)/2 = −1.5
6. Test for a Sample Proportion
Sometimes you are testing a percentage, not a mean — "is the defect rate really under 5%?" The structure is identical; only the formula changes.
p₀ = the value in H₀ (the claimed proportion)
p̂ ("p-hat") is the experimental value — what you actually observed. It has a hat because it is an estimate of the truth.
p₀ ("p-nought") is the null value — what H₀ claims. It has a zero because it is the value you are nillying-out.
Structure of the whole formula, both times: (observed − claimed) ÷ standard error. Write it that way in your notes and you will never mix them up.
Question: A company claims fewer than 5% of its products are defective. A sample of 400 products contains 14 defective items. Test the claim at the 5% level.
Step 1 · Hypotheses. "fewer than" → left-tailed, one-tailed.
H₁: p < 0.05 (more than 5% are defective — claim is false)
Step 2 · Test type and critical value. One-tailed, α = 0.05 → critical value z = −1.645. Reject if z < −1.645.
Step 3 · Sample proportion.
Step 4 · Standard error (using p₀ = 0.05, not 0.035).
Step 5 · Test statistic.
Step 6 · Decision. −1.3766 is not less than −1.645, so it is not in the critical region.
The proportion test needs np₀ ≥ 5 and n(1−p₀) ≥ 5.
Here: 400(0.05) = 20 ≥ 5 ✓ and 400(0.95) = 380 ≥ 5 ✓. Both fine.
Always use p₀ in this check too, not p̂ — the condition is about H₀, not the data.
7. The Six-Step Recipe
If you memorise nothing else from this unit, memorise these six steps. They are the marking scheme, written in order.
- State H₀ and H₁. H₀ has the = sign. H₁ tells you the tail count. [2 marks]
- Name the level α and the type of test. Write "one-tailed, left/right" or "two-tailed". [1 mark]
- State the critical value(s) from the table, and the critical region. [1 mark]
- Calculate the test statistic — show the standard error separately. [2 marks]
- Decide. Compare |statistic| with the critical value. Write REJECT H₀ or FAIL TO REJECT H₀. [1 mark]
- Conclusion in context. Say what it means about the real world, quoting α. [1 mark]
- Name the significance level ("at the 5% level of significance")
- Use "significant evidence", not "proof"
- Reference the original context ("the machine", "the defect rate"), not just the variable
- Match the direction: if you rejected a right-tailed H₀, say the value has increased
- Did I halve α? Two-tailed means α/2 in each tail, so the critical value must be bigger than the one-tailed value. Did I use 1.96 and not 1.645?
- Did I use σ or p₀, not p̂? In the denominator of a proportion test, always the null value. In a mean test, the population σ (or sample s) not the mean.
- Did I write "fail to reject" and never "accept"?
These three checks catch the overwhelming majority of lost marks in this unit.
8. Type I and Type II Errors
A test is never certainly right. It can fail in two opposite ways, and they have formal names.
| H₀ is actually TRUE | H₀ is actually FALSE | |
|---|---|---|
| Reject H₀ | TYPE I ERROR false alarm (probability = α) |
Correct found a real effect |
| Fail to reject H₀ | Correct no real effect |
TYPE II ERROR false negative (probability = β) |
- Lower α (e.g. 0.05 instead of 0.10) → fewer Type I errors, but more Type II errors (β rises). The test becomes more cautious.
- Higher n (bigger sample) → standard error shrinks, z gets larger, both errors fall. Bigger samples are simply better.
You cannot make both errors small at once. This trade-off is a very common 2-mark question.
Never write "there is no effect" or "the null hypothesis is true". Write "there is insufficient evidence to reject the null hypothesis". Failing to reject is a statement about the evidence, not about the world.
Self-check
Open to see the answers
1. What are the critical values for a two-tailed test at α = 0.01, and for a one-tailed test at α = 0.01?
One-tailed: z = 1.645… no — for α = 0.01 one-tailed the value is 2.326
Note the two-tailed value is the LARGER one. That check alone catches the most common error.
2. A sample of 49 items has mean 20.4. The population SD is 3.0. Test at 5% whether the mean is 20. (Assume two-tailed.)
SE = 3.0/√49 = 3/7 = 0.4286
z = (20.4 − 20)/0.4286 = 0.4/0.4286 = 0.933
Critical ±1.96. 0.933 < 1.96 → fail to reject H₀
3. In a sample of 500, 220 say "yes". Test at the 1% level whether the true proportion differs from 0.5.
p̂ = 220/500 = 0.44
SE = √[0.5(0.5)/500] = √0.0005 = 0.02236
z = (0.44 − 0.5)/0.02236 = −0.06/0.02236 = −2.683
Two-tailed α = 0.01 → critical ±2.576. |−2.683| > 2.576 → REJECT H₀
Validity: 500(0.5) = 250 ≥ 5 ✓
4. A quality inspector rejects a batch when more than 5 out of 100 items are faulty. A batch has 8 faulty. What type of error is possible here?