← Back to course
2/6
Week 07 Β· Three Kinds of Bias

Tuesday

Count the training set
// When a fair-looking score hides a gap
⏱ about 20 min

Tuesday: Count the Training Set

Raven tips the new training box onto the table, and forty labeled cards slide across it.

"Let us sort them by label and count," she says.

Rocket builds four stacks. The star stack grows tall, and the planet stack follows close behind.

The streak stack is short. The comet stack is only two cards high.

"Two comets," Rocket says. "Sky Sorter barely saw a comet before the test."

Nova lowers herself to eye level with the tiny stack. "So which kind of bias is hiding in this box?" she asks.

"And how could you have spotted it before testing?"

A sample that leaves cases out

NIST says computational and statistical bias can be in datasets and algorithms.

It often comes from non-representative samples. That means the examples do not include every kind of case in fair amounts.

Sky Sorter learned its cut-offs from its training cards. With only two comets, it had very little evidence about how wide comets are.

LabelTraining cardsTest cardsTest right
star1887
planet1465
satellite streak643
comet220
total402015
How many training cards were comets?
How many comet test cards did Sky Sorter get right?
How many star test cards did it get right?
READ THE TABLE LIKE A REVIEWER
  • Read the question.
  • Tap your answer.
Which label had the fewest training cards?
Which label had the lowest share right on the test?
Which kind of bias does the tiny comet stack show?

Accuracy for each group

Last week you measured accuracy as right answers out of total answers.

One overall number can hide a weak spot. So reviewers also check accuracy for each group of cases.

Here a group is one label. Sky Sorter scored 7 out of 8 on stars but 0 out of 2 on comets.

Out of 40 training cards, how many were NOT comets? Type a number.
WHY THIS EXERCISE40 minus 2 is 38, so almost every training card showed something else.
StatementTrue or false?
Checking accuracy for each label can reveal a gap that the total hides.?
A sample with 18 stars and 2 comets represents comets fairly.?
You can spot a lopsided training set by counting before you test.?
Sky Sorter scored better on planets than on comets.?
WHY THIS EXERCISECounting each group, before and after testing, is how reviewers catch a non-representative sample.
STEPS TO CHECK A TRAINING SET FOR GAPS
  • ?List every label the model must sort.
  • ?Count the training cards for each label.
  • ?Mark any label with far fewer cards.
  • ?After testing, find the accuracy for each label.
  • ?Compare the weak labels with the small stacks.
WHY THIS EXERCISECounting first warns you; testing each group then confirms where the gap is.
Try it
Look around your room and pick a group of things, such as books or pens.
Count each kind. Which kind would a model barely see if it learned only from this group?

Sharp counting. Tomorrow in the Explorer Lab, you build a lopsided deck on purpose and watch what it does.

← Monday