Statistics & Probability · Grades 7

Samples and Populations: Drawing Conclusions from Data

Quick answer

A population is everyone you want to know about; a sample is the part you actually measure. A random sample lets you draw conclusions about the whole population, and a biased one does not — which is decided by how the sample was chosen, not by its size. Larger samples give estimates that vary less, and two groups are compared by looking at both their centers and their spreads.

What you'll learn

  • Tell a random sample from a biased one
  • Estimate a population value from a sample
  • Compare two populations using center and spread

Population and sample

The population is everyone or everything you want to know about. The sample is the part you actually measure.

QuestionPopulationA possible sample
How do students at this school travel in?all 900900 students6060 of them
How long do these batteries last?every battery made200200 tested
Who will win the election?all voters15001500 polled

Measuring the whole population is usually impossible, too slow, or — with the batteries — destroys the thing being measured. So you measure a part and reason about the whole, which is called inference.

What makes a sample random

A sample is random when every member of the population has an equal chance of being chosen.

MethodRandom?
draw 6060 names from all 900900yes
number every student and pick 6060 with a computeryes
ask the first 6060 students off the busno
ask 6060 students in the libraryno

The last two are not random. Students who take the bus are certain to be in the first sample, and students who walk cannot be. That is not bad luck. It is a systematic tilt built into the method.

Why bias cannot be fixed by asking more people

This is the point the whole topic turns on.

Suppose you survey travel habits by asking students getting off the bus, and you ask 600600 of them instead of 6060.

The bus riders were already certain to be included and the walkers already certain to be excluded. Asking ten times as many bus riders does not introduce a single walker. The estimate does not get closer to the truth. It gets more precise about the wrong group.

Sample size and bias fix different problems:

ProblemCaused byFixed by
the estimate wobblessmall samplemore people
the estimate is systematically offbiased methoda better method

A famous case: in 1936 a magazine polled over two million Americans and predicted the wrong winner of a presidential election, while a poll of a few thousand got it right. The huge sample had been drawn from car and telephone owners, who in 1936 were richer than the country as a whole. Two million of the wrong people lost to a few thousand of the right ones.

Estimating from a sample

Scale the sample proportion up to the population.

In a random sample of 5050 students, 1212 walk to school:

1250=0.24=24%\frac{12}{50} = 0.24 = 24\%

For all 900900 students:

0.24×900=216 students0.24 \times 900 = 216 \text{ students}

About 216216. “About” is doing necessary work — a different random sample of 5050 would give a slightly different figure, and the true number is near 216216 rather than exactly it.

That variation is why several samples are better than one. If four samples give 22%22\%, 24%24\%, 27%27\% and 23%23\%, the clustering around the mid-twenties is itself the evidence, and it tells you roughly how much to trust the estimate.

Comparing two populations

To compare two groups, compare both their center and their spread.

Two classes take the same test:

Class A scores A box plot for Class A on an axis from 40 to 100, with a box from 62 to 78 and a median at 70. 40 50 60 70 80 90 100 Min 52 Q1 62 Med 70 Q3 78 Max 88
Class A scores
Class B scores, on the same scale A box plot for Class B on the identical 40 to 100 axis, with a narrower box from 73 to 85 and a median at 79. 40 50 60 70 80 90 100 Min 66 Q1 73 Med 79 Q3 85 Max 90
Class B scores, on the same scale
A: median 70,  IQR 16B: median 79,  IQR 12\text{A: median } 70, \; \text{IQR } 16 \qquad \text{B: median } 79, \; \text{IQR } 12

Class B scored higher and more consistently. Both facts matter, and a report giving only the medians would have missed the second.

How much the two overlap decides how meaningful the gap is. These boxes overlap between 7373 and 7878, so plenty of Class A students outscored plenty of Class B students. B is better on average, not in a different league. Two box plots with no overlap at all would be a much stronger result.

Worked examples

Common mistakes

Practice problems

  1. A school of 600600 students is surveyed by asking 5050 of them. What is the population?

    Answer

    All 600600 students

    Full solution

    The population is everyone the conclusion is meant to describe; the 5050 are the sample.

  2. Is picking names out of a hat a random sample?

    Answer

    Yes

    Full solution

    Every name has an equal chance of being drawn.

  3. Is surveying people outside a gym about exercise habits random?

    Answer

    No

    Full solution

    People at a gym exercise more than average, so the sample is biased towards the answer being sought.

  4. In a sample of 4040, 1010 students own a bike. What proportion is that?

    Answer

    25%25\%

    Full solution

    1040=0.25=25%\tfrac{10}{40} = 0.25 = 25\%.

  5. Using that proportion, estimate how many of 800800 students own a bike.

    Answer

    About 200200

    Full solution

    0.25×800=2000.25 \times 800 = 200.

  6. A sample of 2525 boxes has mean mass 2.42.4 kg. Estimate the mean mass of all the boxes.

    Answer

    About 2.42.4 kg

    Full solution

    The sample mean is the estimate for the population mean.

  7. Two classes have medians 6464 and 7171 on the same test. What else is needed to compare them properly?

    Answer

    A measure of spread

    Full solution

    Centers alone do not say how consistent each class was. An IQR or MAD for each is needed.

  8. Sample P has IQR 44 and sample Q has IQR 1919, on the same measurement. Which is more consistent?

    Answer

    P

    Full solution

    A smaller IQR means the middle half of the data is packed into a narrower range.

  9. Two box plots on one axis overlap almost completely, though one median is slightly higher. What can be concluded?

    Hint

    How much do the two groups differ compared with how much they vary?

    Answer

    Very little. The groups are close to indistinguishable.

    Full solution

    Heavy overlap means most members of each group have counterparts in the other with similar values.

    A small difference in medians against that much variation within each group is weak evidence of a real difference. Two distributions that barely overlap would be a much stronger result.

  10. To find the average time students spend on homework, Mia surveys 500500 students in the library after school. She says the large sample makes her result reliable. Assess her reasoning.

    Hint

    Who is in a library after school?

    Answer

    The sample is biased, and 500500 does not fix that.

    Full solution

    Students who stay in the library after school are more likely than average to be doing homework. Students who go straight home cannot appear in her sample at all.

    That tilt is built into where she stood, so every one of her 500500 responses is drawn from the same skewed group. Surveying 50005000 there would not include a single student who went home.

    Her estimate of homework time will be too high, and the extra numbers make it a more precise overestimate rather than a better one.

    The fix is the method, not the size. Number every student and select at random, so those who leave immediately have the same chance of being asked as those who stay.

Frequently asked questions

What is the difference between a population and a sample?

The population is everyone you want to know about. The sample is the part you actually measure, chosen because measuring everyone is usually impossible.

What makes a sample random?

Every member of the population has an equal chance of being chosen. Drawing names from a hat is random; asking whoever is nearby is not.

What is a biased sample?

One where some members were more likely to be picked than others, so the sample systematically differs from the population. A large biased sample is still biased.

How do I estimate a population value from a sample?

Scale up. If 12 of 50 sampled students walk to school, that is 24%, so about 24% of the whole school likely walks.

Does a bigger sample fix bias?

No. Bias comes from how the sample was chosen, not how big it is. A bigger biased sample gives a more precise wrong answer.

What to learn next

Standards alignment

This lesson covers the following Common Core State Standards for Mathematics.

  • CCSS.MATH.CONTENT.7.SP.A.1Statistics and ProbabilityUnderstand that statistics can be used to gain information about a population by examining a sample of the population; generalizations about a population from a sample are valid only if the sample is representative of that population. Understand that random sampling tends to produce representative samples and support valid inferences.
  • CCSS.MATH.CONTENT.7.SP.A.2Statistics and ProbabilityUse data from a random sample to draw inferences about a population with an unknown characteristic of interest. Generate multiple samples (or simulated samples) of the same size to gauge the variation in estimates or predictions.
  • CCSS.MATH.CONTENT.7.SP.B.3Statistics and ProbabilityInformally assess the degree of visual overlap of two numerical data distributions with similar variabilities, measuring the difference between the centers by expressing it as a multiple of a measure of variability.
  • CCSS.MATH.CONTENT.7.SP.B.4Statistics and ProbabilityUse measures of center and measures of variability for numerical data from random samples to draw informal comparative inferences about two populations.