"Second experiment," Comet says. "Start cue. A whistle or a clap. The deck picks six and six again."
Laps are run. Nova reports: "Whistle mean 93.3, clap mean 93. Difference 0.3 seconds."
"Tiny," Wren says. "But small gaps can still matter. What do you notice when we shuffle?"
Twenty shuffles later: "8 of 20 reached 0.3 or more," Nova says. "That is 40%."
"Way over 5%," Comet says. "Chance makes a gap like ours almost half the time."
"So the cue made no difference we can see," Wren says. "The whistle and the clap are a tie, for these runners."
Nova glows. "Would you like a hint? Not significant is a real answer, not a failure."
"Then the Run Log says: warm-up, yes. Start cue, no," Comet says. "Both are useful to know."
The start-cue gap was 0.3 seconds. In 20 shuffles, 8 reached it: 40%. That is far above 5%.
Chance alone makes a gap this big often. The crew cannot tell the cue apart from luck, so it says the difference is not significant.
Not significant does not mean the cue has zero effect. It means this experiment could not see one. A bigger experiment might.
| Experiment | Observed difference | Shuffles at or beyond | Percent | Judgment |
|---|---|---|---|---|
| warm-up: stretch minus jog | 4 seconds | 1 of 25 | 4% | significant |
| start cue: whistle minus clap | 0.3 seconds | 8 of 20 | 40% | not significant |
The crew's procedure, written out. Read it first, then put it in order.
| Statement | True or false? |
|---|---|
| 8 ÷ 20 × 100 = 40 | ? |
| Not significant means the treatment is proven to have no effect. | ? |
| A significant result points to a cause only because the groups were randomly assigned. | ? |
| The re-randomizing procedure changes when the treatments change. | ? |
Careful thinking. Tomorrow you take the method into everyday comparisons and review the week.