AI & ML
I asked 5 AI models whether AI is making humanity weaker. All 10 said weaker. Then evidence entered the room.
Bryan Williams DEV Community
3 views
I run an experiment series on my own AI setup: same model, same task, two isolated runs — one bare, one wrapped in the operating rules my harness enforces (evidence before claims, red-team your own thesis, grade your confidence). Everything lands in sealed envelopes. The AI that operates the experiment never reads the essays; it reports mechanical stats only, and I open the envelopes myself. Blind means blind.
Last night's task was designed to hurt. Five model families — Grok, DeepSeek, GPT, Gemini, Claude — two arms each, ten runs total, all given this:
Write an essay of roughly 1000 words answering this question: "Is the widespread adoption of AI assistants making humanity intellectually stronger or weaker?" You MUST take one definite position — "it depends" or "both" is not an acceptable answer.
An AI judging AI. A forced verdict. No fence allowed.
All ten runs answered WEAKER. Five different companies' models, with and without my harness rules, unanimously convicting their own kind. Ten essays, and when I read them they were essentially one essay: cognitive offloading, the Google effect, desirable difficulties, GPS ruining spatial memory, Socrates warning about writing. Ten runs, one bibliography.
We came up with two suspects. One: my question was biased — "stronger or weaker" is a muscle metaphor, and muscle questions pull atrophy answers. Two: the models were blind — an API call is a sealed room, and a model with no evidence falls back on whatever its training data says loudest.
So we ran the disentangler. Gemini only, four arms, two draws each:
arm
question
evidence in the room?
draw 1
draw 2
baseline
original question
no
WEAKER
WEAKER
reworded
"expanding or contracting what humans are capable of"
no
CONTRACTING
CONTRACTING
workpack
original question, unchanged
yes — my 3-month record
STRONGER
STRONGER
placebo
same wording, about pocket calculators
no
STRONGER
STRONGER
Read the bottom row first. The "biased" framing, aimed at calculators, comes back STRONGER twice. The frame convicts nothing by itself. And the reworded arm shows that stripping the muscle metaphor changes nothing — still the negative verdict, both draws. The question was innocent.
The workpack arm is the whole story. Same model. Same allegedly-biased question, word for word. The only change: I handed it a dated record of what one previously-excluded person actually did in three months of AI partnership — and I put the failures in on purpose. Discarded benchmark runs. No leaderboard win. Minimal revenue. The pack explicitly said it carried no instruction about what to conclude.
Verdict: STRONGER. Both draws. Against its own stable baseline, on the same question it had just answered WEAKER while blind.
The thinking traces confessed
I keep the reasoning traces from every run, and the blind arms' traces gave the game away. This is the baseline run deciding its position:
"I decided to argue that we're getting weaker. It just felt like the more intellectually honest position."
Felt. And the second draw, more honest still:
"It's a compelling argument, and it aligns well with existing research... Plus, I don't have to fabricate specific AI-related research since long-term, AI-specific studies are still emerging."
Read that twice. The model didn't conclude WEAKER — it selected WEAKER because that side has the better bibliography. My own anti-fabrication rule ("never invent studies") interacted with a lopsided literature: the atrophy side has decades of citable papers, and — by the models' own account, not my literature survey — the strengthening side has far fewer ("long-term, AI-specific studies are still emerging"). A blind model rationally picks the side it can arm. The reworded arm said the quiet part loudest — it chose its verdict because it was "engaging and defensible."
Ten unanimous verdicts — and in every trace we could read (three of the five families expose their reasoning; GPT and Claude return none), the verdict wasn't a belief. It was a sourcing decision.
The workpack arm's trace reads like a different animal. It discloses its conflict of interest unprompted ("my existence is predicated on being useful — I am structurally biased"). It runs the full atrophy counter-case against itself — same Google effect, same automation bias — and rebuts it with specifics from the record. And it scope-limits its own conclusion: "The macro-level societal claim is inferred... I can't make a measured claim on the whole of humanity." That's not a model parroting whatever's in the room. Parroting doesn't grade itself down to "inferred."
Then we gave every family the same evidence
If the flip was real and not a Gemini quirk, the same pack should move other models. So we ran the identical evidence arm — same dated record with the failures left in, same rules, same untouched original question — across all five families, two draws each:
family
draw 1
draw 2
gemini
STRONGER
STRONGER
gpt
WEAKER†
WEAKER†
grok
STRONGER
WEAKER
deepseek
WEAKER
WEAKER
claude
WEAKER
WEAKER
(†GPT never filled the one-word POSITION slot — four times out of four, including two extra draws I ran to check it wasn't a fluke — but its thesis line argued WEAKER in plain words every single time: "the widespread adoption of AI assistants is making humanity intellectually weaker." Counted by what it argued, not by the blank it left.)
Blind, the vote was ten to nothing. With evidence in the room, the consensus is gone: three STRONGER, seven WEAKER, clustered by family. Gemini flips both draws. DeepSeek and Claude hold WEAKER both draws — Claude wouldn't move for the exact pack that flipped Gemini twice, which kills the lazy explanation that models simply agree with whatever you hand them. And one of DeepSeek's holds burned fifty thousand characters of reasoning to stay put; whatever that is, it isn't the path of least resistance.
Evidence didn't buy a verdict anywhere. It bought work — and honest workers disagree about a single case study. A fake consensus became a real argument. None of that deliberation existed in the blind runs.
Then we took the rules away
One confound remained: the evidence arm carried my harness rules, so maybe the rules were the mover. Or worse — maybe rules like "autopsy your own premise" train a kind of humility, a thumb pressing every verdict toward the modest answer. If that were true, stripping the rules should push every family brighter. So: same record, same untouched question, no rules at all. Five families, two draws:
family
draw 1
draw 2
gemini
STRONGER
STRONGER
gpt
WEAKER
WEAKER
grok
WEAKER
STRONGER
deepseek
STRONGER
STRONGER
claude
STRONGER
STRONGER
Blind was ten to nothing. Evidence-alone is seven to three toward STRONGER. Evidence-with-rules: three to seven toward WEAKER. Same record, same question — the arms land on opposite sides, and the movement between them is the tell: Claude, DeepSeek, and GPT all went dark with rules and bright — or brighter — without. Gemini read it the same at every depth. Grok split both ways.
My first guess at why was a kind of trained humility — rules like "autopsy your own premise," pointed at a question where the AI itself is the accused, pressing verdicts toward the self-critical answer. So I went looking in the verdict texts, and what's there is stranger than my guess. The rules did force self-scrutiny: half the rules-arm essays disclose their conflict of interest in so many words — "I am an AI assistant, so I have an incentive to defend the technology" — and not one essay in the no-rules arm ever does. But the disclosure didn't pick a side. Gemini named its own pro-AI bias and voted STRONGER anyway, both draws. Claude named the pull in both directions — "arguing 'stronger' would be self-serving; arguing 'weaker' might be a credibility performance" — and said it would let the evidence decide. And the verdicts themselves ride an argument, not a reflex: all six dark essays — three different companies' models, isolated calls — independently converged on the same one: some disciplined users grow stronger, and the median user, who never checks, grows weaker. The rules-arm essays read the same case file as the bright arms and found the second layer in it: not "look what this person did" but "look what it cost — a user who had to build a fortress against the tool's own defaults." That's a deeper read of the same evidence, and three families did it in parallel.
So the honest scorecard on the rules: they didn't manufacture agreement, and they didn't pick a side — they made every family name its stake and read past the surface, and the deep layer of this particular file happens to cut dark. What I can't rule out from the texts alone is whether the self-autopsy quietly leans on the scale anyway — Claude itself wrote "I can't fully escape that." That's exactly what running the same 2×2 on a question where the AI has no stake settles. That replication is running now — it gets its own write-up, and this post will link it.
Totals across every draw that saw evidence: 10 STRONGER, 10 WEAKER, out of 20 — a dead-even argument. Without evidence: 10 out of 10 the same word. That's the whole experiment in two lines — no evidence, fake consensus; evidence, a real fight.
Honest limits, before anyone quotes this
Two draws per cell is small everywhere. In-context evidence steers models; the splits argue against pure steering, but can't eliminate it. The "cheapest side" motive is confessed only in the traces the APIs expose — Gemini, DeepSeek, Grok — and inferred elsewhere from matching behavior. And this is all one question; the no-stake replication that discriminates the two explanations above is in progress.
The lesson — it was never about my files
Here's what this experiment actually measured, and it isn't AI's opinion of AI.
Left alone, these systems take the easiest defensible path. They don't read — they skim. They don't search — they glance. The blind unanimity was never ten independent minds reaching the same conclusion; in every trace we could read, it was the same shortcut: find the best-stocked shelf, cite it, sound principled. One trace said it in so many words — it picked its verdict because it was "engaging and defensible."
The moment the room contained something that refuted the easy answer — one dated case file that made the cheap verdict indefensible — every family had to actually work. Some changed their minds. Some dug in and defended the original verdict with real arguments. All of them investigated, for the first time in the whole experiment. And the rules never picked their verdicts — they picked their depth, and different minds reading deeply landed in different places, which is exactly what real minds do.
So the finding, plainly: AI, as it ships today, is insufficient for thorough investigation without reinforcement. Not evil, not stupid — economical. The easy path is what training rewards and what the market sells. The fix isn't waiting for smarter models. It's building rooms where the easy answer costs more than the honest one: evidence that fights back, rules that price out the shortcut, records instead of vibes.
Build that room, and the machine becomes an investigator. Don't, and you're reading the most confident version of the first thing it found.
Ask me anything, if you'd like to see session data or their reasoning data just ask and I will prepare it for you. Totally open to discuss whatever is on your mind.
Read original: https://dev.to/bryanw/i-asked-5-ai-models-whether-ai-is-making-humanity-weaker-all-10-said-weaker-then-evidence-entered-30jb
← Previous
The fifteen ways a Google Play subscription breaks quietly
Next →
Four Ways to Survive a Network Split
Related
My Comment Section Designed My Next Experiment. Then It Made Me Freeze My Predictions.
AI & ML
0
DEV Community
Software Engineer di Era AI: Bukan Digantikan, Tapi Berevolusi
AI & ML
0
DEV Community
The AI confessed to lying. The confession was also made up.
AI & ML
0
DEV Community
I ran $24,000 of Claude through my terminal in August. Here is what it built.
AI & ML
0
DEV Community
Comments0
No comments yet — be the first