Every AI that makes art, faces, or words learned from a giant pile of examples — its training data. And here's the rule that decides if a model is fair or a disaster: whatever you put IN shapes what comes OUT. Time to cook a dataset and taste the consequences.
A model eats whatever you feed it. Tap cards to add them to your training data. Every card has hidden properties — some are clean and consented, some are stolen or biased. Watch three meters react live, then read your model's report card. Goal: a dataset that's both useful (enough data) and ethical (low risk, consented, diverse).
Green light needs ALL four: enough data, copyright clear, consent given, and a balanced mix. A dataset can be perfectly legal and still be biased. 🌍
Here's a picture generator. First, train it on a lazy dataset: all scraped from one place, one type, no consent. Then look at the six things it invents. Then fix the diet and look again. Same model — the only thing that changed is what went in.
You're the ethics lead. For each job, pick ingredient cards until the whole checklist passes — licensed, consent, credited, AND diverse — then hit Green-light it. Careful: a set can be totally legal and still fail the diversity check. Clear 3 jobs to win the badge!
Datasets aren't a school exercise — they're the biggest fight in AI today. Artists, photographers, and whole newsrooms are pushing back on how models were fed.
Artists and writers have sued big AI labs, arguing their work was used to train models without permission or payment. Courts are deciding the rules right now.
Sites like "Have I Been Trained?" let creators check if their work is in a dataset and ask to be removed — a real-world opt-out button.
The clean path: pay for or use openly-licensed collections (public domain, Creative Commons, stock libraries) so every example is allowed.
Honest teams publish a datasheet (what's in the data, where it came from) and a model card (what it's good/bad at) so nothing is hidden.
You learned the idea that decides whether an AI is fair or a mess: what goes in shapes what comes out. Copyright, consent, credit, and bias all start at the dataset — long before the model runs. Cook clean, and the model can be trusted.