Comet slaps the SH-1 packet onto the table, open to section P4. "Ninety-two out of a hundred! That is an A. Can we tell the panel yes?"
Wren reads the line twice. "92 out of 100 on what?"
Librarian Joss slides over a note the builder sent. "Their benchmark is a general trivia test. Rivers, animals, old inventions."
Comet's grin fades a little. "But our students would ask about Biology notes and Spanish verbs."
Nova hovers over the packet. "Would you like a hint?" she asks. "Put one trivia question next to one question from your class notes. Are they the same kind of job?"
Wren clicks his pen. "Then we need our own test."
How do we test whether an AI system works for our purpose?
This week you will plan a fair test, score SH-1 and count two kinds of mistakes.
Then you will fill in the test results on the SH-1 model card.
The packet says SH-1 got "92 out of 100 right on our benchmark." That is 92 percent.
This week the crew learns what the benchmark was: a general trivia test.
Trivia is not how Harbor Point students would use SH-1. They would ask about their own class notes.
| A trivia benchmark question (made up) | A Harbor Point study question (made up) |
|---|---|
| Which animal is known for building dams? | In Teacher Lin's notes on cell parts, what does the cell wall do? |
| Name a famous river. | Solve 2x plus 3 equals 11. |
| What is a compass used for? | Conjugate a verb from Teacher Mara's list. |
These questions are made-up examples for our story.
Validation means objective evidence that a system meets the requirements for its intended use.
Reliability means it keeps working as required, without failure, for a given time under given conditions.
A test must match how the system will really be used. It should be a clear, realistic test set.
Good scores on games or on tests made for humans are not proof. They do not show a generative AI system is valid or reliable for that job.
Results in a lab may not match the real world either.
| Statement | True or false? |
|---|---|
| Validation means evidence that a system meets the needs of its intended use. | ? |
| A high score on a test made for humans proves an AI system fits a job. | ? |
| Reliability is about working without failure over time. | ? |
| Lab results always match what happens in the real world. | ? |
Strong start, reviewer. Tomorrow the crew plans its own test with the teachers.