Sampling Design Workbook: 10 Exercises with Full Solutions

Ten sampling design exercises with full worked solutions, from choosing the right method and calculating intervals to classifying survey errors and handling non-response. Try them, then check your reasoning against the standard answers.

This workbook accompanies the Sampling Design article and is built for active practice. Because sampling design is largely a subject of judgement rather than calculation, most of these exercises ask you to choose a method, classify an error, or justify a decision, with a few that involve the small calculations the topic does require, such as a systematic sampling interval and a proportional stratified allocation. Each exercise targets one idea from the lesson so you can see exactly where your understanding needs work.

Attempt every exercise before reading the solutions, and write out your reasoning in full sentences rather than one-word answers, because in this subject the justification is the marks. The solutions name the relevant principle, work through any arithmetic, and explain the choice, so the goal is to compare your reasoning against the standard one.

Part One: The Exercises

Exercise 1 (Census or sample). Explain the main advantage and the main disadvantage of conducting a census rather than a sample. Then give one situation where a sample would clearly be preferred and briefly say why.

Exercise 2 (Primary or secondary data). Classify each as primary or secondary data, with a one-line reason. (a) A company runs a new questionnaire on its own customers to test a product idea. (b) A researcher analyses national census figures published by the government. (c) A student uses a charity’s previously published survey results for a dissertation.

Exercise 3 (Identify the sampling method). Name the sampling method described in each case. (a) An interviewer stops shoppers in a town centre on a weekday until enough people have been questioned. (b) Every unit in the population has an equal, known chance of selection from a complete list. (c) The population is divided into year groups, and a random sample is drawn from each. (d) The country is split into regions, local areas are sampled within regions, and individuals are then sampled within those areas.

Exercise 4 (Systematic sampling interval). A sampling frame lists 2,000 employees, and a sample of 100 is required using systematic random sampling. Calculate the sampling interval, describe how the first unit is chosen, and explain one situation in which this method could give a misleading sample.

Exercise 5 (Proportional stratified allocation). A university has 12,000 students: 5,000 in year one, 4,000 in year two, and 3,000 in year three. A stratified sample of 300 students is to be drawn with proportional allocation. Calculate how many students should be sampled from each year group.

Exercise 6 (Quota or random). For each survey, state whether quota or random sampling is more appropriate and justify it. (a) A company surveys its own employees about a proposed change to working hours. (b) A market researcher wants a quick read on whether the public likes a new chocolate bar, with no list of relevant consumers available.

Exercise 7 (Classify the error). Classify each as sampling error, selection bias, or response bias. (a) A sample mean differs slightly from the true population mean purely due to random selection. (b) A survey of internet users overrepresents younger people because the frame excludes those without internet access. (c) Respondents understate their alcohol consumption because the topic is sensitive.

Exercise 8 (Handling non-response). A postal survey achieves only a 30% response rate. Explain why simply increasing the sample size will not fix the underlying problem, and list three measures that would genuinely help.

Exercise 9 (Choosing a contact method). Recommend a method of contact for each survey and justify it. (a) A survey of household shopping habits among the general public. (b) A government survey asking teachers for their exact pay and tax figures.

Exercise 10 (Stratified versus cluster). Explain the key difference between stratified and cluster sampling in terms of which groups are selected and which units within them. Then identify which design is being used: a researcher with no national student list selects six universities at random and surveys a random sample of students within each.

Part Two: Worked Solutions

Solution 1. The main advantage of a census is that it measures every unit in the population, so it carries no sampling error. The main disadvantage is that it is expensive and slow, and under budget pressure it can even suffer worse non-sampling errors when cheaper, less reliable interviewers are used. A sample would be preferred when resources are limited and the population is large, for example a national survey of consumer attitudes, because focusing resources on a smaller number of units improves data quality while keeping cost and time manageable.

Solution 2. Part (a) is primary data, because it is collected fresh for a specific purpose. Part (b) is secondary data, because the figures already exist, having been collected by the government. Part (c) is secondary data, because the student is reusing results collected earlier by someone else for another purpose.

Solution 3. Part (a) is convenience sampling, since the interviewer simply questions whoever is available. Part (b) is simple random sampling, since every unit has an equal, known, non-zero probability of selection from a complete frame. Part (c) is stratified random sampling, since the population is divided into homogeneous groups and a random sample is taken from each. Part (d) is multistage sampling, since selection happens at several successive stages from larger units down to individuals.

Solution 4. The sampling interval is the population size divided by the sample size.

x=Nn=2000100=20x = \frac{N}{n} = \frac{2000}{100} = 20

This gives a one-in-twenty sample. You choose the first unit by selecting one at random from the first twenty on the list, then take every twentieth unit after that. The method could mislead if the frame has a periodic pattern that aligns with the interval, because a regular cycle in the data matching the sampling interval would cause the sample to repeatedly land on the same kind of unit and miss the real variation.

Solution 5. Proportional allocation samples each stratum in proportion to its size, using the overall sampling fraction.

nN=30012000=0.025\frac{n}{N} = \frac{300}{12000} = 0.025

Applying this fraction to each year group gives the allocation.

n1=5000×0.025=125n_1 = 5000 \times 0.025 = 125
n2=4000×0.025=100n_2 = 4000 \times 0.025 = 100
n3=3000×0.025=75n_3 = 3000 \times 0.025 = 75

So 125 students are sampled from year one, 100 from year two, and 75 from year three, which sum to the required 300.

Solution 6. In part (a), random sampling is more appropriate, because the company holds a personnel list, so a sampling frame exists, accuracy is achievable, and employees contacted through official channels are likelier to respond seriously. In part (b), quota sampling is more appropriate, because no list of relevant consumers is available, speed matters, and only a rough read on public preference is needed rather than statistical precision, so an interviewer can quickly fill quotas by characteristics such as age and gender.

Solution 7. Part (a) is sampling error, since the difference arises purely from the random variation of taking a sample rather than a census. Part (b) is selection bias, since it comes from the sampling frame not matching the target population. Part (c) is response bias, since the measurements themselves are wrong because of the sensitive subject matter.

Solution 8. Increasing the sample size does not fix the problem because it only gathers more responses from the same kind of willing respondents, while the non-respondents, who may differ systematically, remain just as absent. A larger sample of a biased group is still biased. Three measures that would genuinely help are following up non-respondents with callbacks or an alternative contact method, improving the data collection procedures and interviewer training, and offering an incentive such as a cash payment or prize draw entry to raise the response rate.

Solution 9. In part (a), face-to-face contact is best, because a postal survey on such a low-salience topic would get a poor response rate, and a telephone survey would underrepresent groups such as the elderly and the poor. In part (b), self-completion by post or online is best, because respondents cannot recall exact pay and tax figures on the spot and need to look them up, the response rate is likely to be high given teachers’ motivation, and they are comfortable completing forms without an interviewer.

Solution 10. The key difference is that stratified sampling selects from every group, taking some units from each stratum, whereas cluster sampling selects only some of the groups and then takes either all units (one-stage) or some units (two-stage) from those chosen clusters. The design described is two-stage cluster sampling, because the researcher selects only some of the groups, here six universities, and then samples some of the students within each selected one.

How to Get the Most From This Workbook

The thread running through these exercises is that every sampling decision is a trade-off between three things: whether a sampling frame exists, how much accuracy you need, and how much time and money you have. Quota and the other non-probability methods win when no frame exists and speed or cost dominates, at the price of being unable to quantify your error. Probability methods win when a frame exists and you need real inference, with stratified sampling buying accuracy, cluster and multistage designs buying lower cost, and systematic sampling buying convenience as long as you watch for periodicity. Practise naming that trade-off out loud for each scenario, and pair it with an awareness of where non-sampling error can creep in, and you will be able to justify any sampling choice an exam puts in front of you.

See you soon.

View Comments (1)

Leave a Reply

Subscribe to My Newsletter

Subscribe to my email newsletter to get the latest posts delivered right to your email. Pure inspiration, zero spam.

Discover more from Discuss Data Science, Machine Learning and Analytics

Subscribe now to keep reading and get access to the full archive.

Continue reading