Comet writes 75 percent on the board in big letters. "Not 92, but three out of four is not bad. Maybe we pilot it everywhere?"
Wren taps the number. "Seventy-five overall. What do you notice if we split it by class?"
"Does it matter?" Comet asks. "It is the same forty answers."
Teacher Mara looks up from her grading. "It matters to my students."
Nova hovers over the board and splits the total into four rows, one per class. "Would you like a hint?" she asks. "Look for the row that does not look like the others."
Comet reads down the rows and stops. "Spanish. Four out of ten."
The room goes quiet. Then Wren says gently, "Now we know where to look."
Does SH-1 serve every student equally well?
This week you will break TS-40 down by class, look for sameness in SH-1's quizzes and rate SH-1's second risk.
Accuracy measurements can be broken down by group. This is called disaggregation.
Last week the crew reported only totals. Now they split TS-40 by class.
| Class | Questions | Correct | Confidently wrong | Said not sure |
|---|---|---|---|---|
| Biology | 10 | 9 | 1 | 0 |
| World History | 10 | 8 | 2 | 0 |
| Algebra | 10 | 9 | 0 | 1 |
| Spanish (questions in Spanish) | 10 | 4 | 4 | 2 |
| Total | 40 | 30 | 7 | 3 |
This is TS-40 from the SH-1 review story, a made-up example.
NIST names three kinds of AI bias: systemic, computational and statistical, and human-cognitive.
Each of them can happen without anyone meaning to be unfair.
A generative model can also work worse for some groups or languages. For example, it may work less well in languages other than English.
The crew's judgment is that the Spanish gap fits best under computational and statistical bias, because it shows up in the numbers. That is their reading, not a NIST rule.
| Statement | True or false? |
|---|---|
| An overall score can hide a gap between groups. | ? |
| NIST names three kinds of AI bias. | ? |
| Bias only happens when someone means to be unfair. | ? |
| A model may work less well in languages other than English. | ? |
Strong start, reviewer. Tomorrow you will look at why some groups end up less well served.