Resources For Teachers For Tutors For Students & Parents Pricing
Year 12 Methods (Unit 3 & 4) Sampling and estimation

Approximating the distribution of the sample proportion

20 practice questions 0 video lessons Theory + worked examples

In Year 12 Mathematical Methods (Queensland, QCAA), for a large sample the sample proportion \(\hat p=\dfrac{X}{n}\) is approximately normal: \(\hat p\approx N\!\left(p,\ \dfrac{p(1-p)}{n}\right)\), with mean \(p\) and standard deviation (the standard error) \(\sqrt{\dfrac{p(1-p)}{n}}\). The approximation is reliable when \(np\ge 5\) and \(n(1-p)\ge 5\); probabilities are found by standardising and using technology, and a larger \(n\) gives a smaller standard error.

Take a random sample of size \(n\) from a large population in which the proportion of successes is \(p\). The number of successes is \(X\sim B(n,p)\), and the sample proportion is \(\hat p=\dfrac{X}{n}\). Its value varies from sample to sample, so \(\hat p\) is a random variable with mean \(p\) and standard deviation \(\sqrt{\dfrac{p(1-p)}{n}}\).

For a large sample the distribution of \(\hat p\) is close to a normal curve: \(\hat p\approx N\!\left(p,\ \dfrac{p(1-p)}{n}\right)\). The centre is the population proportion \(p\), and the spread is the standard error \(\text{SD}(\hat p)=\sqrt{\dfrac{p(1-p)}{n}}\). This lets us find probabilities such as \(P(\hat p>a)\) or \(P(a<\hat p

The approximation is only good when the sample is large enough, using the rule of thumb \(np\ge 5\) and \(n(1-p)\ge 5\). Its accuracy depends on both \(n\) and \(p\): it is poor for a small sample or for a \(p\) close to \(0\) or \(1\). Increasing \(n\) both improves the approximation and shrinks the standard error, so the sample proportion clusters more tightly around \(p\).

Key idea. For large \(n\), \(\hat p\approx N\!\left(p,\ \dfrac{p(1-p)}{n}\right)\): mean \(p\), standard error \(\sqrt{\dfrac{p(1-p)}{n}}\); valid when \(np\ge 5\) and \(n(1-p)\ge 5\); larger \(n\) means a smaller standard error.
The approximating normal curve for the sample proportionA bell curve centred at the population proportion p with the upper tail beyond a value a shaded, representing the probability that the sample proportion exceeds a. P(p̂ > a) p a
\(\hat p\approx N\!\left(p,\ \dfrac{p(1-p)}{n}\right)\); \(P(\hat p>a)\) is the shaded area
The effect of sample size on the standard errorTwo normal curves for the sample proportion with the same mean p; the tall narrow curve is a large sample with a small standard error, the short wide curve is a small sample with a large standard error. p large n small n
Same \(p\), larger \(n\): a smaller standard error, so a tighter distribution

The approximating normal for a large sample:

\[\hat p\approx N\!\left(p,\ \dfrac{p(1-p)}{n}\right) \qquad \text{mean } p,\quad \text{SD}(\hat p)=\sqrt{\dfrac{p(1-p)}{n}}\]
p^N(p,p(1-p)n)

When the approximation is appropriate (rule of thumb):

\[np\ge 5 \qquad \text{and} \qquad n(1-p)\ge 5\]

Finding a probability - standardise, then use technology:

\[z=\dfrac{a-p}{\sqrt{\dfrac{p(1-p)}{n}}} \qquad P(\hat p>a),\quad P(a<\hat p
The effect of \(n\). Because \(n\) is under the square root, multiplying the sample size by \(4\) divides the standard error \(\sqrt{p(1-p)/n}\) by \(2\), tightening the distribution of \(\hat p\) about \(p\).

How to use the normal approximation for \(\hat p\)

  1. Check the rule of thumb. Confirm \(np\ge 5\) and \(n(1-p)\ge 5\); if either fails, the normal approximation is unreliable.
  2. State the distribution. Write \(\hat p\approx N\!\left(p,\ \dfrac{p(1-p)}{n}\right)\) using the true population proportion \(p\).
  3. Find the standard error. Compute \(\text{SD}(\hat p)=\sqrt{\dfrac{p(1-p)}{n}}\) - remember the square root.
  4. Standardise. For a boundary \(a\), use \(z=\dfrac{a-p}{\text{SD}(\hat p)}\).
  5. Read the probability. Use technology (a normal cdf) for \(P(\hat p>a)\), \(P(a<\hat p
Reverse problems. The mean of \(\hat p\) is \(p\), so a given mean gives \(p\) directly; substitute into \(\sqrt{p(1-p)/n}\), square both sides, and solve for the sample size \(n\).
Example 1 — State the distribution
A random sample of \(n=600\) voters is taken from a population where \(p=0.6\) support a proposal. State the approximate distribution of \(\hat p\).
Solution

Mean \(p\); standard error \(\sqrt{p(1-p)/n}\).

\(\text{SD}(\hat p)\)\(=\)\(\sqrt{\dfrac{0.6\times 0.4}{600}}=\sqrt{0.0004}\)
\(\text{SD}(\hat p)\)\(=\)\(0.02\)
\(\hat p\)\(\approx\)\(N(0.6,\ 0.02^{2})\)
0.02
Example 2 — A probability
A fair coin is tossed \(100\) times (\(p=0.5\)). Find \(P(\hat p\le 0.45)\).
Solution

\(\text{SD}=\sqrt{0.25/100}=0.05\); standardise.

\(z\)\(=\)\(\dfrac{0.45-0.5}{0.05}=-1\)
\(P(\hat p\le 0.45)\)\(=\)\(P(Z\le -1)\)
\(=\)\(0.1587\)
Lower tail P(p-hat <= 0.45) for a fair coin, n=100A bell curve centred at 0.5 with the lower tail below 0.45, one standard error to the left, shaded. 0.5 0.45
0.1587
Example 3 — Is it appropriate?
A rare side effect has \(p=0.02\). A researcher plans the normal approximation for \(\hat p\) in a sample of \(n=100\). Is it appropriate?
Solution

Check the rule of thumb.

\(np\)\(=\)\(100\times 0.02=2\)
\(np=2\)\(<\)\(5\)

No — since \(np<5\) the approximation is poor. A larger sample would fix it.

np=2
Example 4 — The effect of \(n\)
For \(p=0.5\), what sample size \(n\) makes \(\text{SD}(\hat p)=0.01\)?
Solution

Set the standard error equal and square.

\(\sqrt{\dfrac{0.25}{n}}\)\(=\)\(0.01\)
\(\dfrac{0.25}{n}\)\(=\)\(0.0001\)
\(n\)\(=\)\(2500\)
n=2500

Common pitfalls

Take the square root. The standard error is \(\sqrt{\dfrac{p(1-p)}{n}}\), not \(\dfrac{p(1-p)}{n}\) — the latter is the variance. For \(p=0.5,\ n=100\), \(\text{SD}=\sqrt{0.0025}=0.05\), not \(0.0025\).
Use the true \(p\). When the population proportion \(p\) is given, put it in \(\sqrt{p(1-p)/n}\); do not use an observed \(\hat p\) such as \(\sqrt{\hat p(1-\hat p)/n}\).
\(\hat p\) is a proportion, not a count. \(X\sim B(n,p)\) is the number of successes; \(\hat p=X/n\) lies between \(0\) and \(1\). Standardise with the mean \(p\) and standard error \(\sqrt{p(1-p)/n}\), not \(np\) and \(\sqrt{np(1-p)}\).
Check both conditions. The approximation needs \(np\ge 5\) and \(n(1-p)\ge 5\); an extreme \(p\) can fail one even when \(n\) looks large.

Frequently asked questions

How is the distribution of p-hat approximated?

For a large sample, \(\hat p\approx N\!\left(p,\ \dfrac{p(1-p)}{n}\right)\): mean \(p\), standard error \(\sqrt{p(1-p)/n}\).

When is the normal approximation appropriate?

When \(np\ge 5\) and \(n(1-p)\ge 5\). For a small \(n\) or an extreme \(p\) the approximation is poor.

What is the standard error of p-hat?

\(\text{SD}(\hat p)=\sqrt{\dfrac{p(1-p)}{n}}\) — the square root of \(p(1-p)/n\), which itself is the variance.

How do you find a probability for p-hat?

Standardise with \(z=\dfrac{a-p}{\sqrt{p(1-p)/n}}\), then read the probability with technology.

What effect does n have?

A larger \(n\) gives a smaller standard error, so \(\hat p\) is more tightly concentrated about \(p\). Multiplying \(n\) by \(4\) halves the standard error.

Is p-hat the same as X?

No. \(X\sim B(n,p)\) is the count; \(\hat p=X/n\) is a proportion between \(0\) and \(1\), approximately \(N(p, p(1-p)/n)\).

Create a free accountTrack your progress and save your work as you go.
Create free account