← Back to course
1/6
Week 06 Β· Bias, Fairness and Sameness

Monday

Break it down
// Does SH-1 serve every student equally well?
⏱ about 20 min

Monday: Break It Down

Comet writes 75 percent on the board in big letters. "Not 92, but three out of four is not bad. Maybe we pilot it everywhere?"

Wren taps the number. "Seventy-five overall. What do you notice if we split it by class?"

"Does it matter?" Comet asks. "It is the same forty answers."

Teacher Mara looks up from her grading. "It matters to my students."

Nova hovers over the board and splits the total into four rows, one per class. "Would you like a hint?" she asks. "Look for the row that does not look like the others."

Comet reads down the rows and stops. "Spanish. Four out of ten."

The room goes quiet. Then Wren says gently, "Now we know where to look."

This week's driving question

Does SH-1 serve every student equally well?

This week you will break TS-40 down by class, look for sameness in SH-1's quizzes and rate SH-1's second risk.

Real AI check
Nova is a character in our story. Real AI is a tool people build. It does not think or feel like a person, and it can be wrong.

Disaggregate the results

Accuracy measurements can be broken down by group. This is called disaggregation.

Last week the crew reported only totals. Now they split TS-40 by class.

ClassQuestionsCorrectConfidently wrongSaid not sure
Biology10910
World History10820
Algebra10901
Spanish (questions in Spanish)10442
Total403073

This is TS-40 from the SH-1 review story, a made-up example.

READ THE BREAKDOWN
  • Read the question.
  • Tap your answer.
Biology got 9 of 10 correct. What percent is that?
Spanish got 4 of 10 correct. What percent is that?
Add the correct answers for Biology, World History and Algebra. What do you get?
In Spanish, how many answers were confidently wrong or not sure?
Think about it
The overall 75 percent was true. It was also hiding something.
Spanish learners would get a correct answer only 4 times in 10.

Three kinds of bias

NIST names three kinds of AI bias: systemic, computational and statistical, and human-cognitive.

Each of them can happen without anyone meaning to be unfair.

A generative model can also work worse for some groups or languages. For example, it may work less well in languages other than English.

The crew's judgment is that the Spanish gap fits best under computational and statistical bias, because it shows up in the numbers. That is their reading, not a NIST rule.

StatementTrue or false?
An overall score can hide a gap between groups.?
NIST names three kinds of AI bias.?
Bias only happens when someone means to be unfair.?
A model may work less well in languages other than English.?
WHY THIS EXERCISEThese facts explain why the crew splits results by group before deciding.
What is the word for breaking a result down by group? Type it.
WHY THIS EXERCISEDisaggregation is how the crew found the Spanish gap inside the 75 percent.

Strong start, reviewer. Tomorrow you will look at why some groups end up less well served.