Statistics & Probability · Grades 10, 11
Comparing Two Treatments: Is the Difference Real?
Quick answer
Two treatment groups almost never tie, so a gap by itself proves nothing. The test is to ask how big a gap random assignment could have produced on its own. Reshuffling the results into new groups over and over builds exactly that distribution, and a gap that the shuffles rarely reach is a gap chance struggles to explain.
What you'll learn
- Use re-randomization to model differences produced by chance
- Decide whether a difference between two treatments is significant
- Evaluate a report based on data
Two groups, one gap
A teacher tests standing desks. Twelve students volunteer, and a random draw sends six to standing desks and six to regular desks. Everyone then gets twenty minutes of practice problems, and the count finished is recorded.
| Group | Problems finished | Mean |
|---|---|---|
| Standing | ||
| Regular |
The standing group finished more problems on average.
Here is the difficulty. Two groups of six students never tie. Even if standing desks do nothing whatever, the random draw would have put some students in one group and some in the other, and their natural differences would produce a gap.
So the gap of is not the question. The question is whether a gap of is larger than the draw alone would have produced.
Modeling the chance explanation
Suppose for a moment that the desks did nothing. Then every student would have finished the same number of problems either way — the belongs to that student, not to the standing desk.
Under that assumption, the only thing the random draw decided was which six results got labeled “standing.”
That is something anyone can redo. Take the twelve results, deal them into two groups of six at random, and compute the gap. Then do it again. Each shuffle is an experiment in a world where the treatment has no effect.
With twelve results there are possible splits, few enough to check every one rather than sample them.
The pile sits at zero, which is what “the treatment does nothing” should look like. Chance produces gaps of or constantly, gaps of sometimes, and gaps near almost never.
Exactly of the splits reach a gap of or more.
Why reshuffling tests the right thing
The shuffle is not a trick for making numbers. It is the assumption “the treatment did nothing” written down so it can be checked.
If the desks have no effect, each student’s count is a fixed property of that student. The experiment’s randomness came entirely from the draw, and the draw could have gone ways, each equally likely.
Redoing those ways therefore produces exactly the set of gaps this experiment could have shown if the desks did nothing. It is not an approximation, and not a model borrowed from elsewhere.
So the has a precise meaning: if the desks do nothing, a gap this large turns up about once in ninety. The gap that actually happened is one that the chance explanation struggles to produce.
Calling it significant
A difference is called statistically significant when chance alone reaches it only rarely. The usual cutoff is : less often than one time in twenty.
So the standing-desk difference is significant. Because students were assigned at random, the remaining explanation is the desks.
What the two outcomes allow
| Result | What it supports | What it does not support |
|---|---|---|
| Significant | the treatment caused the difference | a claim about size without a margin of error |
| Not significant | nothing about a cause | the claim that the treatment has no effect |
The bottom-right cell catches people out. Failing to find an effect is not the same as finding no effect. A small study can miss a real difference for no reason beyond its size.
Evaluating a report
Most data reaches people as a sentence in an article. These questions decide how much weight the sentence carries.
| Question | Why it matters |
|---|---|
| Were subjects assigned at random? | without it, the report describes an association |
| How many subjects? | a small study misses real effects |
| Is a margin of error given? | a number with no margin hides its own precision |
| How were subjects chosen? | volunteers may not represent anyone else |
| Does the conclusion match the design? | an observational study cannot support a cause claim |
| Who funded it? | not disqualifying, and worth knowing |
Worked examples
Common mistakes
Practice problems
-
In the standing-desk experiment, what is the observed difference in means?
Answer
problems
Full solution
The standing group averaged and the regular group , and .
-
What assumption does a re-randomization make about the treatment?
Answer
That it has no effect, so each subject’s result would be the same in either group.
Full solution
Under that assumption the only random step was which subjects got which label, so reshuffling the labels reproduces what chance alone could have done.
-
Ten of splits reached a gap of or more. Write that as a percentage.
Answer
About
Full solution
, which is to one decimal place.
-
Is a result that chance reaches of the time significant at the cutoff?
Answer
Yes.
Full solution
is below , so chance alone rarely produces a gap this large.
-
A re-randomization reaches the observed gap in of shuffles. Is the result significant?
Answer
No. That is , well above .
Full solution
. Chance produces a gap this size about one time in five, so the data does not separate the treatment from chance.
-
Group A averages and Group B averages . State the observed difference and one thing still needed before calling it significant.
Answer
The difference is . A re-randomization is needed to see how often chance reaches .
Full solution
.
A raw gap carries no verdict on its own. With widely varying results, chance reaches often; with tightly clustered results, it may never. Only the reshuffled distribution settles it.
-
A news report says patients who took a supplement had fewer colds than patients who did not, in a study of adults. What single question decides whether this supports a cause claim?
Answer
Were the adults assigned to take the supplement at random?
Full solution
If a coin decided who took it, the groups started alike and the supplement is the remaining explanation.
If adults chose for themselves, the groups differ in other health habits, and the report supports an association only. The sample size of does not decide this.
-
Find the difference in means for Group A: and Group B: .
Answer
Full solution
and .
The difference is .
-
Explain why a study that finds no significant difference cannot claim the two treatments are equally good.
Hint
What would a study of four people find?
Answer
A small study can miss a real difference, so failing to detect one is not evidence there is none.
Full solution
Significance asks whether the data is too extreme for chance. A study with few subjects produces gaps so variable that almost nothing looks extreme.
A study of four people would call nearly every difference insignificant, including a large real one. The verdict would describe the study’s size rather than the treatments.
Reporting “no significant difference” is therefore a statement about what this study could detect, not a finding of equality.
-
Told that chance reached the observed gap in of shuffles, Leo says there is a chance the desks make no difference. Find his error.
Hint
What was assumed before the was computed?
Answer
He reversed the conditions. The assumes the desks do nothing and measures the data, not the other way around.
Full solution
The calculation began by assuming the desks have no effect. Only under that assumption do the shuffles make sense, since they depend on each result belonging to a student rather than to a desk.
What came out was: given that the desks do nothing, gaps of or more occur of the time.
Leo’s sentence swaps the two halves: given the data, the chance the desks do nothing is . That is a different quantity, and this method does not produce it — it would also require knowing how plausible the no-effect claim was before the experiment.
The correct reading is the first one, and it is enough. A result that the chance explanation reaches once in ninety tries is strong evidence against that explanation.
Frequently asked questions
What does re-randomization do?
It reshuffles the observed results into new groups at random, over and over, to show the range of differences that random assignment alone can produce.
What does statistically significant mean?
The observed difference is larger than chance alone produces except rarely. A common cutoff is 5 percent, meaning chance reaches that gap less than one time in twenty.
Why shuffle the results instead of collecting more data?
The shuffle models the assumption that the treatment did nothing. If it did nothing, each subject's result would be the same in either group, so only the labels were random.
Does a significant result prove the treatment worked?
It rules out chance as a comfortable explanation, and in a randomized experiment that leaves the treatment. It is strong evidence rather than proof.
What should I check in a report based on data?
Whether subjects were assigned at random, how many there were, whether a margin of error is given, and whether the conclusion matches what the design can support.
Standards alignment
This lesson covers the following Common Core State Standards for Mathematics.
- CCSS.MATH.CONTENT.HSS.IC.B.5Making Inferences and Justifying ConclusionsUse data from a randomized experiment to compare two treatments; use simulations to decide if differences between parameters are significant.
- CCSS.MATH.CONTENT.HSS.IC.B.6Making Inferences and Justifying ConclusionsEvaluate reports based on data.