Teacher Mara stops by the library with a stack of Spanish worksheets. "I heard SH-1 keeps every question. Do my students know that?"
Comet shrugs. "The log runs quietly in the background. Nobody even notices it. That is the clever part."
Wren raises an eyebrow. "Nobody notices. What do you notice about that?"
Comet slows down. "Oh. If nobody notices, nobody can say no."
Nova projects the words of section P5 again. "Would you like a hint?" she asks. "Look at the last four words of the section. Then ask who agreed to them."
Teacher Mara reads them. "To improve SH-1. So my students' questions could end up teaching the system?"
Wren writes on his clipboard. "Then we need to know what that could lead to."
The CSTA standards point out that data can be collected and combined across millions of people.
That can happen even when people are not actively using a device or standing near it.
This automated collection that people do not notice can raise privacy concerns.
NIST also lists a privacy risk from AI systems being better at combining data.
NIST says using personal data to train generative AI puts pressure on widely accepted privacy principles.
Three of them are transparency, consent, and using data only for its stated purpose.
P5 says the log will be used "to improve SH-1." That could mean training SH-1 on student questions.
NIST reports that language models have revealed sensitive information that was in their training data. This is called data memorization.
Models may also correctly guess private facts nobody gave them, by stitching pieces of information together.
Even a wrong guess can hurt someone.
| Statement | True or false? |
|---|---|
| Language models have revealed sensitive information from their training data. | ? |
| A model can only share facts that someone typed into it directly. | ? |
| A wrong guess about a person can still hurt them. | ? |
| Data can be collected without people noticing. | ? |
Careful thinking, reviewer. Tomorrow in the Explorer Lab you will read a sample of the log yourself.