← Back to course
AI Explorers 6-8 / Week 05 / Thursday
4/6
Week 05 Β· Rules or Learning?

Thursday

Head to head on Sky Sorter
// Hand-written rules versus patterns found in data
⏱ about 20 min

Thursday: Head to Head on Sky Sorter

The crew writes Sky Sorter version 2. It keeps the streak and comet tests, but swaps in the learned rule: size 4 or more means planet.

"Rematch," Rocket says, shuffling the six practice cards from last week.

Raven runs version 1. Rocket runs version 2. Nova keeps score on the wall.

"Same cards, same order," she says. "Otherwise the comparison is not fair."

When card E comes up, Rocket grins. Version 2 calls it a star, and Raven checks the true label.

"Star," she confirms. "One point to the data."

A fair comparison

To compare two methods fairly, run both on the same cards and count the right answers.

Version 1 uses the hand-written brightness rule. Version 2 uses the size rule found in the data.

Everything else stays the same, so any difference comes from that one rule.

SKY SORTER, VERSION 2
IF streak is yes THEN label = "satellite streak"
ELSE IF size >= 8 THEN label = "comet"
ELSE IF size >= 4 THEN label = "planet"
ELSE label = "star"
CardBrightnessSizeStreak?True labelVersion 1Version 2
A1202nostarstarstar
B2305noplanetplanetplanet
C9012nocometcometcomet
D1801yessatellite streaksatellite streaksatellite streak
E2103nostarplanetstar
F1509nocometcometcomet
Version 1: right out of 6
Version 2: right out of 6
Version 2 as a percent
WHAT DOES THE COMPARISON SHOW?
  • Read the question.
  • Tap your answer.
Which part of version 2 came from the data?
Version 2 got all six right. Does that prove it will never make a mistake?
Why did the crew use the same cards for both versions?

How much data does learning need?

Our lab used ten cards. Real machine learning needs tremendous amounts of training data.

That training data usually has to be supplied by people. Sometimes the machine gathers it itself.

With only a few examples, a pattern can look true by luck.

StatementTrue or false?
Machine learning needs tremendous amounts of training data.?
Training data is usually supplied by people.?
A fair comparison runs both methods on different cards.?
Version 2 is guaranteed never to make a mistake.?
WHY THIS EXERCISEFair tests and lots of data are what make results trustworthy.
WHICH CARD CHANGED?
  • Read the question.
  • Tap your answer.
Which card did version 2 fix?
What was card E's size, the feature that let version 2 get it right?
WHY THIS EXERCISELearning the cut-off from data fixed the card the hand-written rule missed.
RUN A FAIR COMPARISON
  • ?Choose one set of test cards.
  • ?Run version 1 on every card and record its labels.
  • ?Run version 2 on the same cards and record its labels.
  • ?Count the right answers for each version.
  • ?Compare the totals.
WHY THIS EXERCISEKeeping the cards the same is what makes the comparison fair.

Sharp analysis! Tomorrow we look for rules and learning in everyday life.

← Wednesday