Statistics & Probability · Grades 7
Samples and Populations: Drawing Conclusions from Data
Quick answer
A population is everyone you want to know about; a sample is the part you actually measure. A random sample lets you draw conclusions about the whole population, and a biased one does not — which is decided by how the sample was chosen, not by its size. Larger samples give estimates that vary less, and two groups are compared by looking at both their centers and their spreads.
What you'll learn
- Tell a random sample from a biased one
- Estimate a population value from a sample
- Compare two populations using center and spread
Population and sample
The population is everyone or everything you want to know about. The sample is the part you actually measure.
| Question | Population | A possible sample |
|---|---|---|
| How do students at this school travel in? | all students | of them |
| How long do these batteries last? | every battery made | tested |
| Who will win the election? | all voters | polled |
Measuring the whole population is usually impossible, too slow, or — with the batteries — destroys the thing being measured. So you measure a part and reason about the whole, which is called inference.
What makes a sample random
A sample is random when every member of the population has an equal chance of being chosen.
| Method | Random? |
|---|---|
| draw names from all | yes |
| number every student and pick with a computer | yes |
| ask the first students off the bus | no |
| ask students in the library | no |
The last two are not random. Students who take the bus are certain to be in the first sample, and students who walk cannot be. That is not bad luck. It is a systematic tilt built into the method.
Why bias cannot be fixed by asking more people
This is the point the whole topic turns on.
Suppose you survey travel habits by asking students getting off the bus, and you ask of them instead of .
The bus riders were already certain to be included and the walkers already certain to be excluded. Asking ten times as many bus riders does not introduce a single walker. The estimate does not get closer to the truth. It gets more precise about the wrong group.
Sample size and bias fix different problems:
| Problem | Caused by | Fixed by |
|---|---|---|
| the estimate wobbles | small sample | more people |
| the estimate is systematically off | biased method | a better method |
A famous case: in 1936 a magazine polled over two million Americans and predicted the wrong winner of a presidential election, while a poll of a few thousand got it right. The huge sample had been drawn from car and telephone owners, who in 1936 were richer than the country as a whole. Two million of the wrong people lost to a few thousand of the right ones.
Estimating from a sample
Scale the sample proportion up to the population.
In a random sample of students, walk to school:
For all students:
About . “About” is doing necessary work — a different random sample of would give a slightly different figure, and the true number is near rather than exactly it.
That variation is why several samples are better than one. If four samples give , , and , the clustering around the mid-twenties is itself the evidence, and it tells you roughly how much to trust the estimate.
Comparing two populations
To compare two groups, compare both their center and their spread.
Two classes take the same test:
Class B scored higher and more consistently. Both facts matter, and a report giving only the medians would have missed the second.
How much the two overlap decides how meaningful the gap is. These boxes overlap between and , so plenty of Class A students outscored plenty of Class B students. B is better on average, not in a different league. Two box plots with no overlap at all would be a much stronger result.
Worked examples
Common mistakes
Practice problems
-
A school of students is surveyed by asking of them. What is the population?
Answer
All students
Full solution
The population is everyone the conclusion is meant to describe; the are the sample.
-
Is picking names out of a hat a random sample?
Answer
Yes
Full solution
Every name has an equal chance of being drawn.
-
Is surveying people outside a gym about exercise habits random?
Answer
No
Full solution
People at a gym exercise more than average, so the sample is biased towards the answer being sought.
-
In a sample of , students own a bike. What proportion is that?
Answer
Full solution
.
-
Using that proportion, estimate how many of students own a bike.
Answer
About
Full solution
.
-
A sample of boxes has mean mass kg. Estimate the mean mass of all the boxes.
Answer
About kg
Full solution
The sample mean is the estimate for the population mean.
-
Two classes have medians and on the same test. What else is needed to compare them properly?
Answer
A measure of spread
Full solution
Centers alone do not say how consistent each class was. An IQR or MAD for each is needed.
-
Sample P has IQR and sample Q has IQR , on the same measurement. Which is more consistent?
Answer
P
Full solution
A smaller IQR means the middle half of the data is packed into a narrower range.
-
Two box plots on one axis overlap almost completely, though one median is slightly higher. What can be concluded?
Hint
How much do the two groups differ compared with how much they vary?
Answer
Very little. The groups are close to indistinguishable.
Full solution
Heavy overlap means most members of each group have counterparts in the other with similar values.
A small difference in medians against that much variation within each group is weak evidence of a real difference. Two distributions that barely overlap would be a much stronger result.
-
To find the average time students spend on homework, Mia surveys students in the library after school. She says the large sample makes her result reliable. Assess her reasoning.
Hint
Who is in a library after school?
Answer
The sample is biased, and does not fix that.
Full solution
Students who stay in the library after school are more likely than average to be doing homework. Students who go straight home cannot appear in her sample at all.
That tilt is built into where she stood, so every one of her responses is drawn from the same skewed group. Surveying there would not include a single student who went home.
Her estimate of homework time will be too high, and the extra numbers make it a more precise overestimate rather than a better one.
The fix is the method, not the size. Number every student and select at random, so those who leave immediately have the same chance of being asked as those who stay.
Frequently asked questions
What is the difference between a population and a sample?
The population is everyone you want to know about. The sample is the part you actually measure, chosen because measuring everyone is usually impossible.
What makes a sample random?
Every member of the population has an equal chance of being chosen. Drawing names from a hat is random; asking whoever is nearby is not.
What is a biased sample?
One where some members were more likely to be picked than others, so the sample systematically differs from the population. A large biased sample is still biased.
How do I estimate a population value from a sample?
Scale up. If 12 of 50 sampled students walk to school, that is 24%, so about 24% of the whole school likely walks.
Does a bigger sample fix bias?
No. Bias comes from how the sample was chosen, not how big it is. A bigger biased sample gives a more precise wrong answer.
Standards alignment
This lesson covers the following Common Core State Standards for Mathematics.
- CCSS.MATH.CONTENT.7.SP.A.1Statistics and ProbabilityUnderstand that statistics can be used to gain information about a population by examining a sample of the population; generalizations about a population from a sample are valid only if the sample is representative of that population. Understand that random sampling tends to produce representative samples and support valid inferences.
- CCSS.MATH.CONTENT.7.SP.A.2Statistics and ProbabilityUse data from a random sample to draw inferences about a population with an unknown characteristic of interest. Generate multiple samples (or simulated samples) of the same size to gauge the variation in estimates or predictions.
- CCSS.MATH.CONTENT.7.SP.B.3Statistics and ProbabilityInformally assess the degree of visual overlap of two numerical data distributions with similar variabilities, measuring the difference between the centers by expressing it as a multiple of a measure of variability.
- CCSS.MATH.CONTENT.7.SP.B.4Statistics and ProbabilityUse measures of center and measures of variability for numerical data from random samples to draw informal comparative inferences about two populations.