AI just failed a test it should've aced

Strip the training data and it hits a wall πŸ‘€

In partnership with

The Ceiling Nobody Talks About

After a week of AI hacking governments and draining inboxes, here's a story that'll lower your blood pressure a little.

Frontier AI just flunked a test.

Badly.

Researchers built a new benchmark called Reconstruction.

The idea is clever.

Instead of asking a model questions it might have memorized from training...

They gave it just the BIBLIOGRAPHY of a research paper.

Then asked it to reconstruct the actual ideas and hypotheses in the paper.

Original thinking. Not recall. The real thing.

The scores?

The best frontier models managed 3 to 15%.

Even a souped-up setup, four models working together tournament-style, only hit 42%.

Let me translate what that actually means, because it matters for you.

When you strip away "repeat what you've seen before"...

And ask AI to generate genuinely NEW ideas from scratch...

It hits a wall. Hard.

These models are breathtaking at recombining what already exists.

They're mediocre at inventing what doesn't.

Now think about what that tells you about your own value.

For months the fear has been "AI will replace thinking."

This is the data that says: not the hard part.

The hard part of thinking was never retrieving the answer.

It's asking the question nobody's asked yet.

Framing the problem in a way no one framed before.

Making the leap the training data doesn't contain.

That leap is still yours.

And this benchmark just put a number on how far ahead of the machine you still are, in the one place that matters most.

So here's the takeaway for your week.

Stop trying to out-recall the AI. You'll lose. It's read more than you ever will.

Start out-THINKING it.

Point it at the boring recall work.

Keep the original thinking, the weird ideas, the "what if we tried" for yourself.

That's not a consolation prize.

That's the whole job now, and the data says you're still winning it. 🫑

Did You Know?

In 1997, IBM's Deep Blue beat world champion Garry Kasparov at chess...

Everyone declared human thinking obsolete.

Thirty years later, the best chess on Earth is still played by humans AND AI working together, beating either one alone.

"AI won" and "humans still matter" have been true at the same time for three decades. πŸ†

Reply to everything. Edit nothing.

Your inbox is full. Slack is piling up. Client messages need a response yesterday. Typing thoughtful replies to all of it takes hours you don't have.

Wispr Flow turns your voice into clean, professional text you can send the moment you stop talking. Speak like you would to a colleague β€” tangents and all β€” and get polished output. Emails, Slack, LinkedIn, WhatsApp, whatever's open.

89% of messages sent with zero edits. Used by teams at OpenAI, Vercel, and Clay. Works on Mac, Windows, and iPhone.

πŸ—žοΈ The Tip AI News πŸ—žοΈ

Meanwhile... Your agent can now pay for things itself πŸ’³

Quieter story, bigger implications.

Cloudflare just launched two things that push the agent economy forward fast.

First, a browser built specifically FOR AI agents, called Kitesurf.

It runs agents using 3 to 7 times less computing power than a normal browser.

Second, and this is the wild one, a payment protocol called x402.

It lets AI agents pay for services autonomously.

No human clicking "buy." The agent needs a tool, the agent pays for the tool.

Over 20 companies are already building on it.

Here's the shift to wrap your head around.

We're moving from agents that DO tasks to agents that TRANSACT.

That's powerful and a little terrifying at the same time.

Because "my agent can pay for what it needs" and "my agent just spent money I didn't approve" are the same sentence.

Callback to last week: scope the access, set the limits, watch the logs.

An agent with a credit card is exactly why those boring habits matter. 🫑

πŸ˜‚ Meme of the Day

Over to You...

Go point your AI at the recall work this week...

And keep the thinking for yourself. πŸ”₯

Get paid to solve problems like this β†’ AI Certified Consultant

PS. Voting for the UgenticAI "Leading Male in AI" award closes THIS THURSDAY, August 28. Final stretch, 30 seconds, free πŸ‘‰ Vote here πŸ™

Β» Join the AI Money Group Β«
πŸ’° AI Money Blueprint: Your First $1K with AI - Learn the 7 proven ways to make money with AI right now

πŸš€ Zero to Product Masterclass - Watch us build a sellable AI product LIVE, then do it yourself

πŸ“ž Monthly Group Calls - Live training, Q&A, and strategy sessions with Jeff

Sent to: {{email}}

Jeff J Hunter, 3220 W Monte Vista Ave #105, Turlock,
CA 95380, United States

Don't want future emails?

Reply

or to participate.