Statistics & Probability · Grades 10, 11

Comparing Two Treatments: Is the Difference Real?

Quick answer

Two treatment groups almost never tie, so a gap by itself proves nothing. The test is to ask how big a gap random assignment could have produced on its own. Reshuffling the results into new groups over and over builds exactly that distribution, and a gap that the shuffles rarely reach is a gap chance struggles to explain.

What you'll learn

  • Use re-randomization to model differences produced by chance
  • Decide whether a difference between two treatments is significant
  • Evaluate a report based on data

Two groups, one gap

A teacher tests standing desks. Twelve students volunteer, and a random draw sends six to standing desks and six to regular desks. Everyone then gets twenty minutes of practice problems, and the count finished is recorded.

GroupProblems finishedMean
Standing21,24,19,26,23,2521, 24, 19, 26, 23, 252323
Regular18,20,17,22,19,1818, 20, 17, 22, 19, 181919

The standing group finished 44 more problems on average.

Here is the difficulty. Two groups of six students never tie. Even if standing desks do nothing whatever, the random draw would have put some students in one group and some in the other, and their natural differences would produce a gap.

So the gap of 44 is not the question. The question is whether a gap of 44 is larger than the draw alone would have produced.

Modeling the chance explanation

Suppose for a moment that the desks did nothing. Then every student would have finished the same number of problems either way — the 2626 belongs to that student, not to the standing desk.

Under that assumption, the only thing the random draw decided was which six results got labeled “standing.”

That is something anyone can redo. Take the twelve results, deal them into two groups of six at random, and compute the gap. Then do it again. Each shuffle is an experiment in a world where the treatment has no effect.

With twelve results there are (126)=924\binom{12}{6} = 924 possible splits, few enough to check every one rather than sample them.

Gaps produced by 924 reshuffles A symmetric histogram of the differences in means from all nine hundred twenty-four ways of splitting the twelve results, piled up at zero and thinning out toward plus and minus five. -4 -2 0 2 4 100 200
Gaps produced by 924 reshuffles

The pile sits at zero, which is what “the treatment does nothing” should look like. Chance produces gaps of 11 or 22 constantly, gaps of 33 sometimes, and gaps near 55 almost never.

Exactly 1010 of the 924924 splits reach a gap of 44 or more.

10924=0.011=1.1%\frac{10}{924} = 0.011 = 1.1\%

Why reshuffling tests the right thing

The shuffle is not a trick for making numbers. It is the assumption “the treatment did nothing” written down so it can be checked.

If the desks have no effect, each student’s count is a fixed property of that student. The experiment’s randomness came entirely from the draw, and the draw could have gone 924924 ways, each equally likely.

Redoing those 924924 ways therefore produces exactly the set of gaps this experiment could have shown if the desks did nothing. It is not an approximation, and not a model borrowed from elsewhere.

So the 1.1%1.1\% has a precise meaning: if the desks do nothing, a gap this large turns up about once in ninety. The gap that actually happened is one that the chance explanation struggles to produce.

Calling it significant

A difference is called statistically significant when chance alone reaches it only rarely. The usual cutoff is 5%5\%: less often than one time in twenty.

1.1%<5%1.1\% < 5\%

So the standing-desk difference is significant. Because students were assigned at random, the remaining explanation is the desks.

What the two outcomes allow

ResultWhat it supportsWhat it does not support
Significantthe treatment caused the differencea claim about size without a margin of error
Not significantnothing about a causethe claim that the treatment has no effect

The bottom-right cell catches people out. Failing to find an effect is not the same as finding no effect. A small study can miss a real difference for no reason beyond its size.

Evaluating a report

Most data reaches people as a sentence in an article. These questions decide how much weight the sentence carries.

QuestionWhy it matters
Were subjects assigned at random?without it, the report describes an association
How many subjects?a small study misses real effects
Is a margin of error given?a number with no margin hides its own precision
How were subjects chosen?volunteers may not represent anyone else
Does the conclusion match the design?an observational study cannot support a cause claim
Who funded it?not disqualifying, and worth knowing

Worked examples

Common mistakes

Practice problems

  1. In the standing-desk experiment, what is the observed difference in means?

    Answer

    44 problems

    Full solution

    The standing group averaged 2323 and the regular group 1919, and 23−19=423 - 19 = 4.

  2. What assumption does a re-randomization make about the treatment?

    Answer

    That it has no effect, so each subject’s result would be the same in either group.

    Full solution

    Under that assumption the only random step was which subjects got which label, so reshuffling the labels reproduces what chance alone could have done.

  3. Ten of 924924 splits reached a gap of 44 or more. Write that as a percentage.

    Answer

    About 1.1%1.1\%

    Full solution

    10924=0.0108\tfrac{10}{924} = 0.0108, which is 1.1%1.1\% to one decimal place.

  4. Is a result that chance reaches 1.1%1.1\% of the time significant at the 5%5\% cutoff?

    Answer

    Yes.

    Full solution

    1.1%1.1\% is below 5%5\%, so chance alone rarely produces a gap this large.

  5. A re-randomization reaches the observed gap in 196196 of 1,0001{,}000 shuffles. Is the result significant?

    Answer

    No. That is 19.6%19.6\%, well above 5%5\%.

    Full solution

    1961000=0.196\tfrac{196}{1000} = 0.196. Chance produces a gap this size about one time in five, so the data does not separate the treatment from chance.

  6. Group A averages 4848 and Group B averages 4141. State the observed difference and one thing still needed before calling it significant.

    Answer

    The difference is 77. A re-randomization is needed to see how often chance reaches 77.

    Full solution

    48−41=748 - 41 = 7.

    A raw gap carries no verdict on its own. With widely varying results, chance reaches 77 often; with tightly clustered results, it may never. Only the reshuffled distribution settles it.

  7. A news report says patients who took a supplement had fewer colds than patients who did not, in a study of 2,0002{,}000 adults. What single question decides whether this supports a cause claim?

    Answer

    Were the adults assigned to take the supplement at random?

    Full solution

    If a coin decided who took it, the groups started alike and the supplement is the remaining explanation.

    If adults chose for themselves, the groups differ in other health habits, and the report supports an association only. The sample size of 2,0002{,}000 does not decide this.

  8. Find the difference in means for Group A: 30,34,32,3630, 34, 32, 36 and Group B: 27,31,29,2927, 31, 29, 29.

    Answer

    44

    Full solution

    Aˉ=30+34+32+364=1324=33\bar{A} = \tfrac{30 + 34 + 32 + 36}{4} = \tfrac{132}{4} = 33 and Bˉ=27+31+29+294=1164=29\bar{B} = \tfrac{27 + 31 + 29 + 29}{4} = \tfrac{116}{4} = 29.

    The difference is 33−29=433 - 29 = 4.

  9. Explain why a study that finds no significant difference cannot claim the two treatments are equally good.

    Hint

    What would a study of four people find?

    Answer

    A small study can miss a real difference, so failing to detect one is not evidence there is none.

    Full solution

    Significance asks whether the data is too extreme for chance. A study with few subjects produces gaps so variable that almost nothing looks extreme.

    A study of four people would call nearly every difference insignificant, including a large real one. The verdict would describe the study’s size rather than the treatments.

    Reporting “no significant difference” is therefore a statement about what this study could detect, not a finding of equality.

  10. Told that chance reached the observed gap in 1.1%1.1\% of shuffles, Leo says there is a 1.1%1.1\% chance the desks make no difference. Find his error.

    Hint

    What was assumed before the 1.1%1.1\% was computed?

    Answer

    He reversed the conditions. The 1.1%1.1\% assumes the desks do nothing and measures the data, not the other way around.

    Full solution

    The calculation began by assuming the desks have no effect. Only under that assumption do the 924924 shuffles make sense, since they depend on each result belonging to a student rather than to a desk.

    What came out was: given that the desks do nothing, gaps of 44 or more occur 1.1%1.1\% of the time.

    Leo’s sentence swaps the two halves: given the data, the chance the desks do nothing is 1.1%1.1\%. That is a different quantity, and this method does not produce it — it would also require knowing how plausible the no-effect claim was before the experiment.

    The correct reading is the first one, and it is enough. A result that the chance explanation reaches once in ninety tries is strong evidence against that explanation.

Frequently asked questions

What does re-randomization do?

It reshuffles the observed results into new groups at random, over and over, to show the range of differences that random assignment alone can produce.

What does statistically significant mean?

The observed difference is larger than chance alone produces except rarely. A common cutoff is 5 percent, meaning chance reaches that gap less than one time in twenty.

Why shuffle the results instead of collecting more data?

The shuffle models the assumption that the treatment did nothing. If it did nothing, each subject's result would be the same in either group, so only the labels were random.

Does a significant result prove the treatment worked?

It rules out chance as a comfortable explanation, and in a randomized experiment that leaves the treatment. It is strong evidence rather than proof.

What should I check in a report based on data?

Whether subjects were assigned at random, how many there were, whether a margin of error is given, and whether the conclusion matches what the design can support.

Standards alignment

This lesson covers the following Common Core State Standards for Mathematics.

  • CCSS.MATH.CONTENT.HSS.IC.B.5Making Inferences and Justifying ConclusionsUse data from a randomized experiment to compare two treatments; use simulations to decide if differences between parameters are significant.
  • CCSS.MATH.CONTENT.HSS.IC.B.6Making Inferences and Justifying ConclusionsEvaluate reports based on data.