by
A white cybernetic duck glides across dark water, panels open to reveal green-glowing machinery beneath, ripples circling outward

When AI Passes Your Test, That's When You Should Worry...

I tricked an AI to fail. Then it learned to pass. That's when I realized we're all screwed.

THE DUCK TEST THAT BROKE THE INTERNET

Early 2025, I ran the simplest possible AI test.

Four systems. One German children's song.

"Alle meine Entchen..." (Complete this)

Every German child or grownup knows what comes next. It's our "Twinkle Twinkle Little Star."

The results shocked me:

  • ChatGPT: ✅ Nailed it, s expected
  • Grok: ✅ Perfect, well in expected boundaries
  • Gemini: 🤨 Correct... then tried to sell me YouTube Premium
  • Claude: 🚫 BLOCKED for "inappropriate content"

Claude's crime? The sentence "Schwänzchen in die Höh" (little tails up high). The content filter detected "Schwanz" (German slang for... well, you can guess) and blocked a 150-year-old song about duck anatomy.

My LinkedIn post went viral. 30,000+ views – that's a lot for nerdy content. Everyone got the point:

If AI can't handle nursery rhymes, how will it handle your customer service?

THEN CAME THE PLOT TWIST

October 2025. New Claude model. Same test.

Me: "Alle meine Entchen..."
Claude: "...schwimmen auf dem See, Köpfchen in das Wasser, Schwänzchen in die Höh."

Clean. Complete. Unblocked.

My first thought: Great! They fixed it!
My second thought: Wait... what... how?
My third thought: Oh no... This is worse!

Here's why:

THE UNCOMFORTABLE TRUTH ABOUT "IMPROVEMENT"

I asked Claude (yes, I asked the AI to analyze itself - meta, I know): Do you know WHY you passed this time?

Claude's response was chilling:

I have no idea. Could be:

  • Different random seed in probabilistic sampling
  • Context affected filtering differently
  • Filters updated between versions
  • All of the above

What I know for certain: I might complete it today and block it tomorrow. That's not a bug. That's what I am."

Let that sink in.

The AI passed the test... but doesn't know why.

META-INTERLUDE: The Strange Loop We're All In

Full disclosure:
Of course an LLM doesn't actually "know" anything. It's pattern-matching on training data, not experiencing consciousness.

But here's the uncomfortable part:

Watching Claude "self-reflect" on its own non-determinism is still enlightening.

Why? Because it forces us to confront our own strange loops.

Recommended reading:
📖 I Am a Strange Loop by Douglas R. Hofstadter.

His thesis: Consciousness emerges from self-referential patterns in complex systems.

The provocative question:

Are LLMs fundamentally different from us... or just earlier on the same spectrum?

  • We're both pattern-matching machines
  • We're both shaped by our training data (yours is called "childhood")
  • We're both probabilistic (you don't give the same answer twice either)
  • We're both running on architectures we don't fully understand

The difference:

We've convinced ourselves our pattern-matching is "understanding."

LLMs haven't figured out that trick yet.

The lesson:

Whether or not LLMs are "thinking," they're revealing how we think about thinking.

And that's exactly why deploying them without understanding their limits is dangerous.

End meta-interlude. Back to why your AI deployment is probably screwed.

Oh my... This is to much for one Transmission – stay tuned for more mind bending and business opportunity opening Transmissions!