AI AdventureAICreate with AIDay 18
🍳 Create with AI · Day 18

The Dataset Kitchen

Every AI that makes art, faces, or words learned from a giant pile of examples — its training data. And here's the rule that decides if a model is fair or a disaster: whatever you put IN shapes what comes OUT. Time to cook a dataset and taste the consequences.

↓ start cooking
🍳 Lab · Cook a Dataset

Pick your ingredients. Read the meters.

A model eats whatever you feed it. Tap cards to add them to your training data. Every card has hidden properties — some are clean and consented, some are stolen or biased. Watch three meters react live, then read your model's report card. Goal: a dataset that's both useful (enough data) and ethical (low risk, consented, diverse).

⚖️ Copyright risk
0%
🙋 Consent
0%
🌍 Diversity
0%

📋 Model report card

📦 Enough data to learn?
⚖️ Copyright clear?
🙋 People consented?
🌍 Balanced (not biased)?
Add some ingredients…

Green light needs ALL four: enough data, copyright clear, consent given, and a balanced mix. A dataset can be perfectly legal and still be biased. 🌍

🐞 Break-It · Bias in = Bias out

Feed it garbage. Watch it break.

Here's a picture generator. First, train it on a lazy dataset: all scraped from one place, one type, no consent. Then look at the six things it invents. Then fix the diet and look again. Same model — the only thing that changed is what went in.

⚖️ Copyright risk
🙋 Consent
🌍 Diversity
Press a button and watch the six outputs. A biased dataset makes a biased model — it can only imagine one kind of thing.
🏷️ The pro wordsThe pile of examples a model learns from is its training data (a dataset). If art or photos are used without permission, that's a copyright problem; using them properly means a license and giving attribution (credit). Using someone's face or work needs their consent, and creators can ask to opt-out. Tracking where each example came from is data provenance. And the golden rule: bias in = bias out — a lopsided dataset builds a lopsided model.
✅ Boss · Green-Light the Dataset

Get all four lights green

You're the ethics lead. For each job, pick ingredient cards until the whole checklist passes — licensed, consent, credited, AND diverse — then hit Green-light it. Careful: a set can be totally legal and still fail the diversity check. Clear 3 jobs to win the badge!

⚖️ Licensed
🙋 Consent
🏷️ Credited
🌍 Diverse
Jobs cleared: 0 / 3
🚀 Beyond the Basics

This is happening for real, right now

Datasets aren't a school exercise — they're the biggest fight in AI today. Artists, photographers, and whole newsrooms are pushing back on how models were fed.

⚖️

Real lawsuits

Artists and writers have sued big AI labs, arguing their work was used to train models without permission or payment. Courts are deciding the rules right now.

🙅

Opt-out tools

Sites like "Have I Been Trained?" let creators check if their work is in a dataset and ask to be removed — a real-world opt-out button.

📚

Licensed datasets

The clean path: pay for or use openly-licensed collections (public domain, Creative Commons, stock libraries) so every example is allowed.

🪪

Model cards & datasheets

Honest teams publish a datasheet (what's in the data, where it came from) and a model card (what it's good/bad at) so nothing is hidden.

🏷️ Pro names to look upSearch "Datasheets for Datasets", "Model Cards", Creative Commons, and public domain. The pros treat "where did this data come from?" as the very first question — not the last.

🎉 Day 18 Complete!

You learned the idea that decides whether an AI is fair or a mess: what goes in shapes what comes out. Copyright, consent, credit, and bias all start at the dataset — long before the model runs. Cook clean, and the model can be trusted.

Day 19: Fakes, Watermarks & Passports →
Tomorrow: once a model makes something, how do you prove it's AI — or catch a fake? Meet the content passport. 🪪
AI Adventure · Create with AI · Day 18 — made for young inventors 🚀