AI AdventureAICreate with AIDay 17
⚖️ Create with AI · Day 17

The Judge's Table

A generator hands you four pictures. Which one is best? "Best" isn't a vibe — it's a rubric: did it follow the brief, does it look good, any glitches? Today you're the judge. Then you'll meet the machine that tries to judge for you… and see where it's wrong.

↓ take the bench
⚖️ Lab · Candidate Lineup

Rank them, best to worst

The generator made 4 candidates for one brief. Tap them in order — 1st tap = your #1 pick, and so on down to #4. Use the rubric to decide.

THE BRIEF: a happy blue round creature
  • 1️⃣ Follows the brief? Right colour, shape, mood.
  • 2️⃣ Looks good? Clean, clear, pleasant.
  • 3️⃣ Any artifacts? Weird counts, mush, glitches = drop it down.
🐞 Break-It · Gaming the Metric

When the score lies

Say our robo-judge scores a picture only on "how blue and round is it?" — a cheap, easy-to-measure rule. Here are two candidates for "a happy blue round creature". You decide who's actually better. Then hit the button and see who the lazy metric crowns.

Candidate A — a real creature
blueness score: –
Candidate B — just a blue dot
blueness score: –

A plain blue circle "wins" a blueness contest while being a boring, useless picture. Chasing an easy number instead of the real goal is called gaming the metric — a huge trap in AI.

🏷️ The pro wordsJudging generated output is evaluation. "Did it do what the brief asked?" is prompt adherence. When humans rank outputs by taste, that's human preference — and training a model to chase those human rankings is RLHF (reinforcement learning from human feedback). An automatic number like a CLIP score on a benchmark is only a proxymetrics are proxies, helpful but never the final word.
🎯 Boss · Tournament Bracket

Run the playoffs

Four candidates enter. Pick the winner of each semifinal, then crown a champion in the final. Complete 3 brackets to earn the badge. (This head-to-head style is exactly how real AI leaderboards rank models!)

Brackets done: 0 / 3
🥊 Semifinals — pick a winner in each pair
🚀 Beyond the Basics

How the real judges work

Everyone wants to know "which model is best?" — but the honest answer needs people, not just a number:

🏆

Leaderboards (head-to-head)

Sites like LMArena show two AI answers side by side and let thousands of humans vote which is better. Your bracket is a tiny version of that.

🅰️🅱️

A/B tests (real users)

Show version A to half the users, B to the other half, measure which people actually prefer. Real preference beats any guess.

🧑‍⚖️

Human eval still rules (the final call)

Metrics like CLIP or FID are fast screening tools. But the last word — "is this actually good?" — comes from a human. Always.

🏷️ Pro names to look upAutomatic image scores include CLIP score and FID. Human-vote rankings use an Elo system (like chess). Big test sets are benchmarks. And the method that taught ChatGPT to match human taste is RLHF.
🔬 Real Model Lab

The machine judge — real CLIP

This is CLIP, a real model that measures how well a picture matches a sentence, running 100% on your device. We'll score your 4 lineup candidates against the brief and get a machine ranking — then compare it to your ranking. Spoiler: you won't always agree.

not loaded

First load downloads the model from Hugging Face, then it's cached for offline use. On a 4GB laptop give it a minute. ☕

🎉 Day 17 Complete!

You ran the judge's table: rank by a rubric (brief, quality, artifacts), watch out for gaming the metric, and use automatic scores like CLIP as a hint — never the verdict. Metrics are proxies; humans decide. That's the beating heart of how modern AI is tuned.

Day 18 →
Next up: you can generate it, spot its glitches, and judge it — time to put your creative AI skills together. 🚀
AI Adventure · Create with AI · Day 17 — made for young inventors 🚀