Raven tips the new training box onto the table, and forty labeled cards slide across it.
"Let us sort them by label and count," she says.
Rocket builds four stacks. The star stack grows tall, and the planet stack follows close behind.
The streak stack is short. The comet stack is only two cards high.
"Two comets," Rocket says. "Sky Sorter barely saw a comet before the test."
Nova lowers herself to eye level with the tiny stack. "So which kind of bias is hiding in this box?" she asks.
"And how could you have spotted it before testing?"
NIST says computational and statistical bias can be in datasets and algorithms.
It often comes from non-representative samples. That means the examples do not include every kind of case in fair amounts.
Sky Sorter learned its cut-offs from its training cards. With only two comets, it had very little evidence about how wide comets are.
| Label | Training cards | Test cards | Test right |
|---|---|---|---|
| star | 18 | 8 | 7 |
| planet | 14 | 6 | 5 |
| satellite streak | 6 | 4 | 3 |
| comet | 2 | 2 | 0 |
| total | 40 | 20 | 15 |
Last week you measured accuracy as right answers out of total answers.
One overall number can hide a weak spot. So reviewers also check accuracy for each group of cases.
Here a group is one label. Sky Sorter scored 7 out of 8 on stars but 0 out of 2 on comets.
| Statement | True or false? |
|---|---|
| Checking accuracy for each label can reveal a gap that the total hides. | ? |
| A sample with 18 stars and 2 comets represents comets fairly. | ? |
| You can spot a lopsided training set by counting before you test. | ? |
| Sky Sorter scored better on planets than on comets. | ? |
Sharp counting. Tomorrow in the Explorer Lab, you build a lopsided deck on purpose and watch what it does.