Resources For Teachers For Tutors For Students & Parents Pricing
Year 12 Methods (Unit 3 & 4) Sampling and estimation

Populations and samples

20 practice questions 0 video lessons Theory + worked examples

In Year 12 Mathematical Methods (Queensland, QCAA), a population parameter is a fixed number describing the whole population — for a proportion it is \(p\), the true (usually unknown) proportion. A sample statistic is computed from a sample: the sample proportion \(\hat p=\dfrac{X}{n}\) estimates \(p\) but varies from sample to sample. A random sample gives every member an equal chance, which avoids bias so that \(\hat p\) is likely to be close to \(p\).

A population is the whole group we want to know about; a sample is the part of it we actually examine. A number that describes the population is a parameter; a number worked out from the sample is a statistic. For a proportion, the parameter is \(p\), the true proportion of the population with a feature, and the statistic is the sample proportion \(\hat p=\dfrac{X}{n}\), where \(X\) is the number of successes in a sample of size \(n\).

The key distinction is that \(p\) is fixed (one value for the whole population, usually unknown) while \(\hat p\) varies: take a different sample and you generally get a different \(\hat p\). We use \(\hat p\) as an estimate of \(p\). This spread of \(\hat p\) between samples is called sampling variability.

How the sample is chosen decides whether \(\hat p\) is trustworthy. In a simple random sample every member of the population is equally likely to be chosen, which keeps the sample representative. Methods that are not random introduce bias: a convenience sample (only easy-to-reach members), a voluntary response sample (only people who choose to reply) or any other non-random selection systematically pushes \(\hat p\) above or below \(p\).

Key idea. \(p\) is a fixed population parameter; \(\hat p=\dfrac{X}{n}\) is a sample statistic that varies between samples. Random sampling avoids bias so \(\hat p\) estimates \(p\) well.
Sampling variability of the sample proportionA number line for p-hat from 0 to 1 with a dashed line at the true proportion p = 0.4. Seven dots from repeated random samples are scattered close to 0.4, showing that p-hat varies from sample to sample about the fixed p. p = 0.4 0 0.2 0.4 0.6 0.8 1 ĥ
Sampling variability: repeated random samples give different \(\hat p\), scattered about the fixed \(p\)
Unbiased versus biased samplingA number line for p-hat with a dashed line at the true proportion p = 0.4. Blue dots from an unbiased random method cluster on 0.4, while red dots from a biased method cluster near 0.65, away from p. p = 0.4 0 0.2 0.4 0.6 0.8 1 unbiased biased ĥ
An unbiased (random) method centres \(\hat p\) on \(p\); a biased method centres it away from \(p\)

The sample proportion (successes over sample size):

\[\hat p=\dfrac{X}{n} \qquad X=\text{number of successes},\quad n=\text{sample size}\]
p^=Xn

Parameter versus statistic:

\[p=\text{population proportion (fixed parameter)} \qquad \hat p=\text{sample proportion (varies)}\]

The sample proportion estimates the population proportion, and its value depends on the sample:

\[\hat p \approx p \qquad \text{but } \hat p \text{ generally differs from } p \text{ and from other samples}\]
Random sampling. In a simple random sample every member of the population is equally likely to be chosen. Random selection avoids bias, so \(\hat p\) is likely to be close to \(p\); a convenience or voluntary-response sample is not random and biases \(\hat p\).

How to work with populations and samples

  1. Identify the parameter and the statistic. The true proportion for the whole population is \(p\) (fixed); the proportion from a sample is \(\hat p\) (varies).
  2. Compute the sample proportion. \(\hat p=\dfrac{X}{n}\) — number of successes over sample size.
  3. Check how the sample was chosen. A simple random sample (every member equally likely) is unbiased; a convenience or voluntary-response sample is biased.
  4. Judge the estimate. If the method is random, \(\hat p\) should be close to \(p\); if it is biased, \(\hat p\) will systematically over- or under-state \(p\).
  5. Remember variability. A different random sample gives a different \(\hat p\); the parameter \(p\) does not change, only the statistic does.
Which value can change? Only \(\hat p\). The population parameter \(p\) is a single fixed number; taking another sample changes the statistic \(\hat p\), not \(p\).
Example 1 — Compute a sample proportion
Of a random sample of \(200\) homes, \(80\) have air conditioning. Find the sample proportion \(\hat p\).
Solution

The sample proportion is successes over sample size.

\(\hat p\)\(=\)\(\dfrac{X}{n}=\dfrac{80}{200}\)
\(\hat p\)\(=\)\(0.4\)

So \(\hat p=0.4\).

0.4
Example 2 — Parameter vs statistic
A factory's globes are \(8\%\) faulty. In a random sample of \(50\), \(3\) are faulty. State \(p\) and \(\hat p\), and say which can change.
Solution

The \(8\%\) is the population value; \(\hat p\) comes from the sample.

\(p\)\(=\)\(0.08\) (population parameter)
\(\hat p\)\(=\)\(\dfrac{3}{50}=0.06\)

Only \(\hat p\) changes with the sample; \(p\) is fixed.

0.06
Example 3 — Identify bias
A radio station asks listeners to phone in; \(350\) of \(500\) callers vote ``yes''. Find \(\hat p\) and say whether it is reliable.
Solution

Only listeners who chose to call are included.

\(\hat p\)\(=\)\(\dfrac{350}{500}=0.7\)
sample\(=\)voluntary response

It is self-selected and biased, so \(0.7\) is unreliable.

0.7
Example 4 — Sampling variability
Three random samples of \(25\) fish give \(5\), \(8\) and \(6\) tagged. Find each sample proportion.
Solution

Divide each count by \(25\); the values differ.

\(\hat p_1\)\(=\)\(\dfrac{5}{25}=0.2\)
\(\hat p_2\)\(=\)\(\dfrac{8}{25}=0.32\)
\(\hat p_3\)\(=\)\(\dfrac{6}{25}=0.24\)

Different samples give different \(\hat p\): sampling variability.

Three sample proportions 0.2, 0.32, 0.24A number line for p-hat from 0 to 0.5 with dots at 0.2, 0.24 and 0.32, showing the three sample proportions differ. 0 0.1 0.2 0.3 0.4 ĥ
0.2,0.32,0.24

Common pitfalls

Do not swap \(p\) and \(\hat p\). \(p\) is the fixed population parameter (usually unknown); \(\hat p\) is the sample statistic worked out from data, and it varies between samples. \(\hat p\) is not the parameter.
A large sample does not fix bias. If the method is not random — e.g.\ surveying only a club's members — the estimate stays biased no matter how many people are surveyed.
Use \(\hat p=\dfrac{X}{n}\). Divide the successes by the sample size, not by the number of failures, and do not invert the fraction.
\(\hat p\) need not equal \(p\). Because of sampling variability, a single sample rarely gives \(\hat p\) exactly equal to \(p\); it is an estimate that scatters around \(p\).

Frequently asked questions

What is the difference between a parameter and a statistic?

A parameter describes the whole population and is fixed (the true proportion \(p\)); a statistic is calculated from a sample (the sample proportion \(\hat p\)) and varies between samples.

How do you calculate a sample proportion?

\(\hat p=\dfrac{X}{n}\), the number of successes divided by the sample size. For \(80\) of \(200\), \(\hat p=0.4\).

What is a random sample and why does it matter?

One in which every member has an equal chance of selection. Randomness avoids bias, so \(\hat p\) is likely to be close to \(p\).

What are sources of bias?

Convenience samples (only easy-to-reach members), voluntary response samples (only those who choose to reply) and other non-random methods all bias \(\hat p\).

What is sampling variability?

Different random samples give different \(\hat p\). The parameter \(p\) is fixed; only \(\hat p\) changes, scattered around \(p\).

What is a simple random sample?

One where every member of the population is equally likely to be chosen — e.g.\ number the population and use a random number generator.

Create a free accountTrack your progress and save your work as you go.
Create free account