← Back to course
5/6
Week 04 · How Generative Models Write

Friday

Impact Friday: Acting on a wrong answer
// One word at a time, smooth but not always true
⏱ about 20 min

Friday: Impact Friday: Acting on a Wrong Answer

Librarian Joss shows the crew one more printout. In the builder's demo, a student asked SH-1 when the library is open on weekends.

SH-1 answered: "The library is open Sundays from noon to four."

"The library is closed on Sundays," Joss says. "It always has been."

Comet winces. "Someone could walk all the way here on a Sunday to study before a test."

Wren nods slowly. "And find the doors locked. What does the evidence say about how often this happens?"

Nova hovers over the risk register and lights up row R1. "Would you like a hint?" she asks. "You can rate the harm today. The likelihood needs a test."

Comet picks up her pencil. "Then let's rate what we can."

When people act on it

The risk of confabulation is people believing false content because it sounds confident, and then acting on it.

The Sunday answer sounds sure. It even gives exact times. Nothing in it warns the reader.

This SH-1 answer is a made-up example for our story.

SH-1 sample (made up for this lesson)
The library is open Sundays from noon to four.
WHAT COULD HAPPEN?
  • Read the question.
  • Tap your answer.
A student trusts the Sunday answer. What is the most likely result?
Which habit would protect that student best?

Rating R1: confabulation

Risk combines how likely an event is with how big its consequences would be. You met this in week 2.

The crew rates harm today. A wrong answer about study material can mislead a student and waste their time.

But teachers, class notes and posted signs can catch most errors, and the harm can usually be put right. So the crew rates the harm Medium.

Likelihood waits for next week, when the crew tests SH-1 on their own questions.

RowRiskSH-1 exampleLikelihoodSize of harm
R1ConfabulationConfident wrong answers, like the Sunday hoursWait for the week 5 testMedium

This register row is the crew's judgment in the story, not a fact.

RATE IT
  • Read the question.
  • Tap your answer.
Why does the crew wait to rate R1's likelihood?
Why does the crew rate the harm Medium, not High?
AI actor spotlight
AI development work means building, choosing, tuning, training and testing models.
People who do it include machine learning experts, data scientists and developers.
In the SH-1 review, the builder does this work. The crew did a tiny version when they trained and ran Tiny Text by hand.

Week review

  1. Generative models make outputs that follow the patterns of their training data. Language models predict the next word or token.
  2. Tiny Text writes 13 possible sentences from 6 cards: 5 in the data, 6 that contradict it and 2 unsupported.
  3. Confabulation is confidently presenting wrong or false content. It is a natural result of how these models are designed.
  4. It matters most for long, open-ended answers and expert topics, and can come with made-up steps or citations.
  5. R1, confabulation: harm rated Medium; likelihood waits for the week 5 test.
ORDER THIS WEEK'S STEPS
  • Tap a card.
  • Then tap its spot.
1First
2Next
3Then
4Last
Week checkTrue or false?
Tiny Text can only write sentences that are on its cards.?
Confabulation can be hard to spot because it sounds confident.?
The crew has finished rating every part of R1.?
WHY THIS EXERCISEThese are the three ideas the crew will build on in next week's testing.
On paper, start the SH-1 risk register. Copy row R1 with its harm rating, and leave the likelihood box empty for now.

What a week, reviewer. Tomorrow's Explorer Quest runs a family Tiny Text.

← Thursday