Raven spreads the tally table next to the mission log. "Let us trace exactly how the cups made that false sentence."
"After Raven, the cup holds one spots slip and one logs slip," she says. "Rocket drew spots."
"Then the a cup held three comet slips. Comet was the most likely word."
Rocket groans. "So every step followed the pattern, and the whole sentence was still false."
Nova circles the table slowly. "Every step was likely," she says. "But likely is not the same as true."
"What would happen if the cups wrote a whole page instead of one sentence?"
Generative models predict outputs from patterns in their training data.
NIST explains that this statistical prediction can produce accurate outputs, and it can also produce false or inconsistent ones.
Each word the cups picked was a common pattern. Nothing in the cups checked whether the whole sentence matched the real night.
| Step | Cup | Slips inside | Slip drawn |
|---|---|---|---|
| 1 | Raven | 1 spots, 1 logs | spots |
| 2 | spots | 4 a | a |
| 3 | a | 3 comet, 1 planet, 2 star | comet |
NIST says confabulation is a bigger risk for open-ended questions that ask for long answers.
It is also a bigger risk on topics that need expert knowledge.
And people may believe false content because the answer sounds so confident.
| Statement | True or false? |
|---|---|
| Statistical prediction can produce true outputs and false outputs. | ? |
| Long, open-ended answers carry less confabulation risk than short ones. | ? |
| Expert topics carry more confabulation risk. | ? |
Clear reasoning. Tomorrow in the Explorer Lab, you become a fact-checker.