← Back to course
AI and You 9-12 / Week 07 / Tuesday
2/6
Week 07 · Privacy in the Machine

Tuesday

Collection you do not see
// What a question log reveals
⏱ about 20 min

Tuesday: Collection You Do Not See

Teacher Mara stops by the library with a stack of Spanish worksheets. "I heard SH-1 keeps every question. Do my students know that?"

Comet shrugs. "The log runs quietly in the background. Nobody even notices it. That is the clever part."

Wren raises an eyebrow. "Nobody notices. What do you notice about that?"

Comet slows down. "Oh. If nobody notices, nobody can say no."

Nova projects the words of section P5 again. "Would you like a hint?" she asks. "Look at the last four words of the section. Then ask who agreed to them."

Teacher Mara reads them. "To improve SH-1. So my students' questions could end up teaching the system?"

Wren writes on his clipboard. "Then we need to know what that could lead to."

Collection you cannot see

The CSTA standards point out that data can be collected and combined across millions of people.

That can happen even when people are not actively using a device or standing near it.

This automated collection that people do not notice can raise privacy concerns.

NIST also lists a privacy risk from AI systems being better at combining data.

SPOT THE HIDDEN COLLECTION
  • Read the question.
  • Tap your answer.
Why might a Harbor Point student not notice the P5 log?
What makes a log of every question across many students a bigger privacy concern than one note?

Training on personal data

NIST says using personal data to train generative AI puts pressure on widely accepted privacy principles.

Three of them are transparency, consent, and using data only for its stated purpose.

P5 says the log will be used "to improve SH-1." That could mean training SH-1 on student questions.

WHICH PRINCIPLE IS UNDER PRESSURE?
  • Read the question.
  • Tap your answer.
Students were never told the log exists.
Students were never asked whether they agreed.
A question asked for homework help is later used to train the system.

What a model can give away

NIST reports that language models have revealed sensitive information that was in their training data. This is called data memorization.

Models may also correctly guess private facts nobody gave them, by stitching pieces of information together.

Even a wrong guess can hurt someone.

SH-1 sample (made up for this lesson)
Question: Give me a practice question about verb endings.
SH-1: Here is one a Spanish student asked late on Monday night: How do the endings change for "we"?
Answer key: this sample shows the risk. If SH-1 were trained on the log, it could repeat another student's question and when they asked it.
StatementTrue or false?
Language models have revealed sensitive information from their training data.?
A model can only share facts that someone typed into it directly.?
A wrong guess about a person can still hurt them.?
Data can be collected without people noticing.?
WHY THIS EXERCISEMemorization and guessing are why training on student questions is risky.
When a language model reveals information that was in its training data, NIST calls it data ____. Type one word.
WHY THIS EXERCISEData memorization is one reason the crew does not want SH-1 trained on student questions.

Careful thinking, reviewer. Tomorrow in the Explorer Lab you will read a sample of the log yourself.

← Monday