← Back to course
2/6
Week 09 Β· Confident and Wrong

Tuesday

Why prediction can go wrong
// Why smooth answers still need checking
⏱ about 20 min

Tuesday: Why Prediction Can Go Wrong

Raven spreads the tally table next to the mission log. "Let us trace exactly how the cups made that false sentence."

"After Raven, the cup holds one spots slip and one logs slip," she says. "Rocket drew spots."

"Then the a cup held three comet slips. Comet was the most likely word."

Rocket groans. "So every step followed the pattern, and the whole sentence was still false."

Nova circles the table slowly. "Every step was likely," she says. "But likely is not the same as true."

"What would happen if the cups wrote a whole page instead of one sentence?"

Likely, step by step, but not checked

Generative models predict outputs from patterns in their training data.

NIST explains that this statistical prediction can produce accurate outputs, and it can also produce false or inconsistent ones.

Each word the cups picked was a common pattern. Nothing in the cups checked whether the whole sentence matched the real night.

StepCupSlips insideSlip drawn
1Raven1 spots, 1 logsspots
2spots4 aa
3a3 comet, 1 planet, 2 starcomet
How many slips were in the a cup?
Which word had the most slips in the a cup?
Did any cup check the real mission log? Type yes or no.

Where the risk grows

NIST says confabulation is a bigger risk for open-ended questions that ask for long answers.

It is also a bigger risk on topics that need expert knowledge.

And people may believe false content because the answer sounds so confident.

WHERE WOULD YOU CHECK MOST CAREFULLY?
  • Read the question.
  • Tap your answer.
Which request carries more confabulation risk, according to NIST?
Why might a reader believe a false answer?
Every word the cups chose was likely. Is the sentence surely true?
HOW A LIKELY SENTENCE CAN TURN OUT FALSE
  • ?Nobody checked it against the real facts.
  • ?It picks a likely next word from its patterns.
  • ?The finished sentence sounds smooth and sure.
  • ?It repeats, one word at a time.
  • ?The model starts with a word.
WHY THIS EXERCISEEach step follows the pattern, but the patterns never compare the sentence with reality.
StatementTrue or false?
Statistical prediction can produce true outputs and false outputs.?
Long, open-ended answers carry less confabulation risk than short ones.?
Expert topics carry more confabulation risk.?
WHY THIS EXERCISENIST names open-ended long answers and expert topics as higher-risk places.
The cups write a page of 10 sentences, and a fact-check finds 4 are false. How many are true? Type a number.
WHY THIS EXERCISE10 minus 4 is 6. A long page gave four false sentences room to slip in.
Try it
Use your Week 8 cups to draw a five-sentence paragraph.
Underline any sentence that sounds sure. Circle any you could prove true from the mission log.

Clear reasoning. Tomorrow in the Explorer Lab, you become a fact-checker.

← Monday