← Back to course
5/6
Week 06 · Bias, Fairness and Sameness

Friday

Impact Friday: Worse off than with no AI
// Does SH-1 serve every student equally well?
⏱ about 20 min

Friday: Impact Friday: Worse Off Than With No AI

The crew meets Teacher Mara and Librarian Joss to decide what to do about the Spanish gap.

"We could just leave it out of Spanish," Comet says. "Problem solved."

"That protects my students," Teacher Mara says. "But they lose the help everyone else gets."

Wren looks at his clipboard. "Then we need two things: keep it out for now, and ask for a fix we can test."

Nova hovers over the risk register and lights up row R2. "Would you like a hint?" she asks. "Ask who was missing when SH-1 was designed."

Teacher Mara smiles. "Someone who teaches Spanish, for a start."

Comet writes it down first.

A gap that leaves some students worse off

A generative model can work worse for some groups or languages.

People may wrongly trust it to work the same for everyone. That can leave a group worse off than with no AI at all.

For Harbor Point, that group is Spanish learners. A study helper that is right only 4 times in 10 can teach wrong answers.

Refining the proposal

The crew suggests three changes to reduce bias. These are story choices, not facts.

ChangeWhy
Keep SH-1 out of Spanish for nowProtects Spanish learners from a 40 percent helper
Ask the builder to add Spanish study materialTargets the gap the crew found
Test again with a new Spanish test before any useOnly new evidence can show the gap is closed
WEIGH THE CHANGES
  • Read the question.
  • Tap your answer.
Why not stop at "keep SH-1 out of Spanish"?
Why test again after the builder adds Spanish material?

Rating R2: harmful bias and homogenization

The crew rates R2 using this week's evidence. Spanish questions were right only 4 of 10 times, and all 12 SH-1 quizzes were alike.

Likelihood: High, because the gap and the sameness showed up in the crew's own tests.

Size of harm: High, because one group could end up worse off than with no AI.

Also, SH-1's quizzes would practice only one kind of thinking.

RowRiskSH-1 exampleLikelihoodSize of harm
R1ConfabulationConfident wrong answers (7 of 40 in TS-40)HighMedium
R2Harmful bias and homogenizationSpanish questions 4 of 10 right; quizzes all alikeHighHigh

These ratings are the crew's judgment in the story, not facts.

AI actor spotlight
AI impact assessment work checks accountability, works against harmful bias, and examines a system's effects, safety and security.
Diverse teams share ideas and assumptions more openly, which helps them find problems and risks.
In the SH-1 review, the crew does impact assessment when it rates R2. Teacher Mara's view helped them see what the numbers meant for her students.
SPOTLIGHT CHECK
  • Read the question.
  • Tap your answer.
Which task is part of AI impact assessment?
How do diverse teams help?

Week review

  1. Disaggregating TS-40 shows the gap: 26 of 30 in the English classes, 4 of 10 in Spanish.
  2. NIST names three kinds of bias, and each can happen without anyone meaning to be unfair.
  3. A model can work worse for some languages, and trusting it equally can leave a group worse off than with no AI.
  4. SH-1's quizzes show homogenization: one kind of question in all 12.
  5. Fixing bias does not by itself make a system fair. Training on AI-made data can cause model collapse.
  6. R2 is rated High likelihood and High harm.
ORDER THIS WEEK'S STEPS
  • Tap a card.
  • Then tap its spot.
1First
2Next
3Then
4Last
Week checkTrue or false?
The overall 75 percent showed that SH-1 works equally well in every class.?
The crew rated R2 High for both likelihood and size of harm.?
Keeping SH-1 out of Spanish for now is the crew's only change.?
WHY THIS EXERCISEThese are the decisions the crew will carry into the final recommendation.
On paper, add row R2 to your SH-1 risk register, with its ratings and one piece of evidence.

What a week, reviewer. Tomorrow's Explorer Quest looks at who a helper might leave out.

← Thursday