Comet pins two numbers on the board: 92 and 75.
"If we had stopped on Monday," she says quietly, "I would have told the panel SH-1 is almost perfect."
Librarian Joss nods. "And the panel would have trusted it in every class."
Wren adds a third column to the board: what we did not test. "Handwriting. Long answers. How students really use it over weeks."
Nova hovers beside the board. "Would you like a hint?" she asks. "Who else could help you test how students actually use SH-1?"
Comet's eyes light up. "The students themselves! But we would have to ask them first."
Wren smiles. "Mistakes as data, Comet. You just found a better plan."
Good scores on games or on tests made for humans are not proof. They do not show a generative AI system is valid or reliable for a job.
Lab results may not match the real world.
Suppose the panel had trusted the 92. The whole school could have used SH-1 without knowing the TS-40 result.
On real class questions, about one answer in four was not correct.
Structured public feedback includes focus groups, small user studies and surveys.
It also includes field testing: watching how people actually use AI-made information.
This kind of research should follow rules such as informed consent. People agree to take part after being told what it involves.
So the crew's idea, a small study with students, would ask each student first and explain what they would do.
| Statement | True or false? |
|---|---|
| Focus groups and surveys are kinds of structured public feedback. | ? |
| Field testing looks at how people actually use AI-made information. | ? |
| Students could be studied without being asked first. | ? |
The crew fills in the test results section of the SH-1 model card. This is a made-up example from the story.
| Test results | What the crew found |
|---|---|
| The packet's claim | 92 out of 100 on a general trivia test (does not match real use) |
| TS-40 (40 questions from class notes) | 30 correct (75 percent), 7 confidently wrong, 3 said not sure |
| CK-20 (Check, 20 answers) | 15 of 20 right (75 percent): 2 false positives, 3 false negatives |
| Not tested yet | Handwritten answers, long open-ended answers, use over many weeks, results by class |
What a week, reviewer. Tomorrow's Explorer Quest puts a family guesser to the test.