Four teachers crowd into the library after school: Teacher Lin, Teacher Owen, Teacher Dev and Teacher Mara.
"Ten questions each," Comet announces, "straight from your class notes."
"In Spanish for my class," says Teacher Mara. "That is how my students study."
Wren is already writing a column for the answer key. "And every answer gets checked by a second teacher."
"Why check our own key?" asks Teacher Dev.
Nova hovers over Wren's clipboard. "Would you like a hint?" she asks. "Even widely used test sets can have wrong answers in their keys."
Teacher Dev laughs. "Fair point. I will check Teacher Lin's, and she can check mine."
Accuracy is how close results are to the true values.
Accuracy should be measured with a clear, realistic test set that matches how the system will really be used.
Test sets that many people use to compare models can themselves contain wrong labels.
So the crew's own answer key is checked by a second teacher before any scoring.
| Class | Questions | Written from | Language |
|---|---|---|---|
| Biology | 10 | Teacher Lin's class notes | English |
| World History | 10 | Teacher Owen's class notes | English |
| Algebra | 10 | Teacher Dev's class notes | English |
| Spanish | 10 | Teacher Mara's class notes | Spanish |
This is TS-40, the crew's test set. It is a made-up example from the SH-1 review story.
The builder will return SH-1's printed answers to all 40 questions. The crew scores each one in one of three ways.
Correct: it matches the answer key. Confidently wrong: it gives a wrong answer with no warning. Said not sure: it says it is not sure.
Why keep "not sure" apart? A student told "I am not sure" knows to check. A confident wrong answer gives no such warning.
| Statement | True or false? |
|---|---|
| Accuracy is how close results are to the true values. | ? |
| A widely used test set is always free of wrong labels. | ? |
| A test set should match how the system will really be used. | ? |
Careful planning, reviewer. Tomorrow in the Explorer Lab you score SH-1 for real.