← Back to course
5/6
Week 05 Β· Measuring a Model

Friday

Impact Friday: The wrong test
// The right test, the right numbers
⏱ about 20 min

Friday: Impact Friday: The Wrong Test

Comet pins two numbers on the board: 92 and 75.

"If we had stopped on Monday," she says quietly, "I would have told the panel SH-1 is almost perfect."

Librarian Joss nods. "And the panel would have trusted it in every class."

Wren adds a third column to the board: what we did not test. "Handwriting. Long answers. How students really use it over weeks."

Nova hovers beside the board. "Would you like a hint?" she asks. "Who else could help you test how students actually use SH-1?"

Comet's eyes light up. "The students themselves! But we would have to ask them first."

Wren smiles. "Mistakes as data, Comet. You just found a better plan."

A score on the wrong test

Good scores on games or on tests made for humans are not proof. They do not show a generative AI system is valid or reliable for a job.

Lab results may not match the real world.

Suppose the panel had trusted the 92. The whole school could have used SH-1 without knowing the TS-40 result.

On real class questions, about one answer in four was not correct.

WHY IT MATTERS
  • Read the question.
  • Tap your answer.
Who could be affected if the panel trusted the 92 and approved SH-1 everywhere?
TS-40 got 30 of 40 correct. How many answers in 4 is that?

Structured public feedback

Structured public feedback includes focus groups, small user studies and surveys.

It also includes field testing: watching how people actually use AI-made information.

This kind of research should follow rules such as informed consent. People agree to take part after being told what it involves.

So the crew's idea, a small study with students, would ask each student first and explain what they would do.

StatementTrue or false?
Focus groups and surveys are kinds of structured public feedback.?
Field testing looks at how people actually use AI-made information.?
Students could be studied without being asked first.?
WHY THIS EXERCISETesting with people is powerful, and it must respect the people being tested.
AI actor spotlight
TEVV stands for testing, evaluation, verification and validation.
TEVV work examines a system, finds problems and fixes them. It can start as early as design.
Ideally, the people who check a model are not the same people who build and use it.
In the SH-1 review, the crew does TEVV work. They are separate from the builder, and their own test found what the 92 hid.

The model card: test results

The crew fills in the test results section of the SH-1 model card. This is a made-up example from the story.

Test resultsWhat the crew found
The packet's claim92 out of 100 on a general trivia test (does not match real use)
TS-40 (40 questions from class notes)30 correct (75 percent), 7 confidently wrong, 3 said not sure
CK-20 (Check, 20 answers)15 of 20 right (75 percent): 2 false positives, 3 false negatives
Not tested yetHandwritten answers, long open-ended answers, use over many weeks, results by class

Week review

  1. A test must match how the system will really be used.
  2. Validation is evidence that a system meets the needs of its intended use. Reliability is working without failure over time.
  3. Accuracy is closeness to the true values, measured on a clear, realistic test set.
  4. TS-40: 30 of 40 correct, 75 percent. R1 likelihood is now High.
  5. CK-20: 2 false positives and 3 false negatives. Some failures do more harm than others.
  6. TEVV checkers should be separate from the builders.
ORDER THIS WEEK'S STEPS
  • Tap a card.
  • Then tap its spot.
1First
2Next
3Then
4Last
WEEK CHECK
  • Read the question.
  • Tap your answer.
Which section of the model card did the crew fill in this week?
Why is "not tested yet" an important line on the card?
On paper, copy the test results section of the SH-1 model card in your own words.

What a week, reviewer. Tomorrow's Explorer Quest puts a family guesser to the test.

← Thursday