Comet slaps the finished register on the table. "Five High rows, sure. But Biology and Algebra got 9 of 10 each. Let's say go, everywhere, tomorrow!"
Wren shakes his head. "What does the evidence say? Spanish got 4 of 10. Privacy and bias are High and High. I say no-go."
They both look at the register, then at each other.
"Is there anything between go and no-go?" Comet asks.
Nova splits the word "go" on the wall and writes a long "with..." after it. "Would you like a hint?" she asks. "Turn each High row into a rule the pilot must keep."
Wren taps his pencil, thinking. Then he turns to you. "Reviewer, what would make a pilot safe enough to try?"
All AI actors share the job of deciding whether AI is the right tool for a purpose at all. They also share deciding how to use it responsibly.
A decision to use an AI system should weigh its trustworthiness, risks and benefits, with input from many people.
Clear go or no-go decisions are one goal of NIST's framework.
The trustworthy characteristics pull against each other, so teams must balance trade-offs.
A secure but unfair system, or an accurate system nobody can understand, is still not good enough.
Trustworthiness is only as strong as its weakest characteristic. For SH-1, the Spanish results are a weak point.
| Option (made up) | What it gains | What it risks |
|---|---|---|
| Go everywhere now | SH-1's help in every class | Spanish learners worse off; privacy and bias rows still High |
| No-go | No new risks | Losing help in Biology and Algebra, where SH-1 got 9 of 10 each |
| Go with conditions | A small pilot in the best-scoring classes | Some risks remain, so rules and checks are needed |
The crew's recommendation: go, with conditions. A six-week pilot in Biology and Algebra only. Not in Spanish until a new test shows SH-1 serves Spanish learners as well as the others.
| ID | Condition (made up for the story) | From |
|---|---|---|
| C1 | No names in logs, IDs change weekly, raw log deleted after 14 days, no training on student questions | R3 |
| C2 | Teachers check every uploaded notes page and every SH-1 quiz before students see it | R1, R4, R6 |
| C3 | Every SH-1 output carries the "Made with SH-1" mark and footer | R5 |
| C4 | Check is a second opinion only; the teacher has the final say | R7 |
| C5 | A report button and an incident log the panel reviews every week | R1, R4, R7 |
| C6 | A paper option for every task, and library laptops for anyone without home access | Access |
| C7 | Re-run TS-40 in weeks 3 and 6 of the pilot; stop if results drop | R1, R2 |
| C8 | The builder must name what it knows about the base model and its training data, and share any energy information | R8, R9 |
| Statement | True or false? |
|---|---|
| Deciding whether AI is the right tool is only the builder's job. | ? |
| A decision to use AI should weigh trustworthiness, risks and benefits. | ? |
| A secure but unfair system is good enough. | ? |
| Trustworthiness is only as strong as its weakest characteristic. | ? |
| The pilot includes Spanish from day one. | ? |
Brave, careful thinking, reviewer. Tomorrow the crew presents to the panel.