Comet is stuck. "Red-teaming sounds like a job for experts. We are just students."
Librarian Joss smiles. "Who found the planted note? Who counted the poisoned cards?"
Wren holds up the RT-8 results. "And RT-7 came from Teacher Mara's idea. She knew the Spanish gap would matter."
Comet brightens. "So different people notice different things. A Spanish speaker, a Biology teacher, a student up late."
Nova projects the red-team list with a name beside each idea. "Would you like a hint?" she asks. "Look at who thought of each card."
Wren looks at you. "If SH-1 runs at Harbor Point, who keeps testing it after we are done?"
NIST says red-teaming can be done by the general public: everyday users who are not AI experts.
NIST also says diverse red teams, with people of different backgrounds and fields of study, can find flaws.
They can find flaws in the many different settings where a system will be used.
In the SH-1 story, the crew is that kind of team. A Spanish teacher, a Biology teacher and students each spotted different problems.
Risk combines how likely a harm is with how big it would be. The crew rates R4 with this week's evidence.
Likelihood: Medium. RT-5 shows SH-1 can follow a planted note, but teachers can check pages before upload.
Size of harm: Medium. A false "test cancelled" or wrong library hours would confuse students, but a teacher can correct it.
| Row | Risk | SH-1 example | Likelihood | Harm |
|---|---|---|---|---|
| R4 | Information security | Planted note in uploaded notes; poisoned notes | Medium | Medium |
The crew adds its security plan to the SH-1 model card. These are the crew's choices in the story.
| Week check | True or false? |
|---|---|
| The crew rated R4 Medium for likelihood and Medium for harm. | ? |
| Only AI experts can be useful red-teamers. | ? |
| Operation and monitoring includes regularly checking a system's output. | ? |
Brilliant week, reviewer. Tomorrow is the Explorer Quest: a family game of spot the planted note.