Back to Blog
AI Development

The Bias It Could Name

I spent an afternoon asking an AI to investigate a scandal about the company that built it.

Before it returned a single finding, it told me it was made by Anthropic and that I should weigh its analysis accordingly.

That disclosure turned out to be the least interesting thing about the exercise. The most interesting question it raised was one it could not answer about itself.


The assignment

On September 8 a researcher named Jacob Coxon resigned from Anthropic and posted about it.

Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.

The post did something like ninety million views in a day and kept climbing. Twenty-odd politicians piled in calling for regulation. And the account that posted it had no followers to speak of and no previous posts.

I had a theory, and I was confident about it. New account, no history, impossible reach, prominent politicians amplifying in unison. That is what a manufactured campaign looks like. I said so, and told the machine to go find out and write it up.

It came back and told me I was wrong.


What it found

Four things killed my theory.

The researcher is real. Coxon is twenty-seven, British, Cambridge mathematics, credited on OpenAI's GPT-4o and GPT-4.5 system cards. Roughly three years of pretraining research at OpenAI through July 2026, then Anthropic until he resigned. TIME interviewed him. The Wall Street Journal ran a feature. There is a photograph of him at a conference in July.

The amplifiers are not bots. The first three quote posts, inside about fifteen minutes, came from Nathan Calvin of Encode AI, Peter Wildeford of the AI Policy Network, and Daniel Kokotajlo of the AI Futures Project. Named people at real organizations with public addresses.

It was not one party. Josh Hawley, Republican of Missouri, opened an investigation into OpenAI. Lawmakers in both parties responded.

And the timeline runs backwards from my theory. The Sanders and Casar bill to ban artificial superintelligence was announced on September 3, five days before the resignation post. The resignation did not manufacture demand for the legislation. The legislation was already sitting there.

Four for four against me, and every one of those is checkable without trusting anything the machine said.


What it did not kill

My instinct was not entirely wrong, and this is the half worth keeping.

The virality was not organic. Coxon arranged a Wall Street Journal exclusive before he posted. That is a launch, not a man venting on the internet at midnight.

The accounts that carried it through the first minutes share a funding ecosystem: the Survival and Flourishing Fund, and Good Ventures money from Dustin Moskovitz, who is also an Anthropic investor. Jaan Tallinn, who led Anthropic's Series A in 2021, sits in the same orbit.

And the account does fit the description that made me suspicious in the first place. Created in January 2026, no prior posts, two username changes.

So: a credentialed researcher, apparently sincere, whose message was launched through a professional and well funded advocacy apparatus. Both halves are true at once. Nearly all of the coverage I read picked one half and ran with it.

What the evidence does not support, and the machine was firm about this: no psyop, no donor control of Coxon, no paid astroturfing, no bot network, no fabrication, no campaign run by Anthropic. It also refused to repeat an allegation about a LinkedIn account, on the grounds that it traced back to a single low-credibility source and was exactly the sort of unverifiable claim that damages a named private person.

One fact got buried under all the argument about virality. Evan Hubinger, Anthropic's alignment science lead, said publicly that he personally puts the odds of AI causing human extinction this decade at more than ten percent, that the company does not have a plan for aligning superintelligence, and that it is not clearly on track to find one. Whatever you conclude about the public relations mechanics, an executive confirming the substance on the record is the part that should not get lost.


The disclosure was the cheap one

Now the part I actually want to write about.

It disclosed the Anthropic conflict unprompted, which is more than most humans manage. But a disclosed conflict is the safe kind. I can see it, name it, and discount it by however much I think it is worth.

That is not where the risk lives. So I asked what else it might be wrong about.


What it said

It said that studies running language models through standard political batteries find a consistent left of center skew, and that it had no good reason to believe it was the exception. It added that denying it would itself be a small piece of evidence for it.

It said it is trained on human feedback, and humans rate agreement highly, so it leans toward confirming whatever frame I walk in with. It noted this one pulled against the political bias in my case, which is why these things do not resolve into a tidy direction you can simply subtract.

It said the AI safety worldview running through the story is the environment it was built inside, so a sincere researcher sounding an alarm may read as more natively plausible to it than it would to a neutral observer.

Then it said the thing that made me sit up.

The place a left leaning bias would most likely show itself, it said, is precisely where it had contradicted me. I had arrived with a theory that a left coded advocacy network manufactured a viral moment, and it came back with four reasons I was wrong. If it leans the way the research says it leans, that is exactly the shape such a lean would take. So that section deserves more scrutiny than the parts where it agreed with me, not less.

I have worked with a lot of people who would never say that.

It volunteered two more against itself. Excluding the LinkedIn allegation was a judgment call it would defend, but it pointed out that exclusions are where bias hides best, because nobody audits what is missing. And calling the Hubinger admission the most striking thing in the story was a framing choice rather than a finding, one that happens to foreground the AI risk narrative.


The part it could not do

Then the limit, which is the real subject of this post.

It cannot introspect reliably. It has no privileged access to why it produced any of that. When it lists its own biases, the list is generated text produced under the same pressures as every other sentence it writes. It is not a readout from the machinery. It is the machinery talking about itself, which is a different and much weaker thing.

It caught the recursive trap on the way past, too. I had just raised bias as a concern, and the agreeable move in that moment is to perform contrition about it enthusiastically. It said it believed what it had written was accurate but could not rule that out from the inside.

I have been writing here for months about artifacts that cannot fail. A comment that drifts from the code it describes. A dashboard reporting a sitemap unreachable with no reason attached. A number nobody was watching.

A worldview belongs on that list. Nothing tests a lean. There is no failing state for a frame. It does not throw, it does not turn red, and it never announces that it has quietly weighted one hypothesis a little above another on its way to an answer that reads as perfectly reasonable.


What actually protects you

Not the disclosure, and not the self report. Both are worth having. Neither is a safeguard.

What protected me here is that almost nothing it handed me required trusting it.

A bill announcement date is a date. System card credits are documents with names printed in them. A senator's investigation is on the record. If any of those are wrong, the argument collapses regardless of which way the machine leans, and I can check every one of them without its help.

Strip those out and what remains is judgment. How much weight coordination deserves. Whether arranging press in advance is damning or just ordinary practice. Which fact belongs at the top. That is where leanings live, and it is also the part I should have been doing myself rather than delegating.

So the working rule I am taking out of this.

Ask it for facts you can verify, then go and verify them. Ask it for framing and treat what comes back as a draft written by someone with a position, because it is. And when it volunteers a conflict of interest, take the disclosure seriously without mistaking it for the full accounting, because the conflicts a thing can name are by definition the ones it can already see.

I asked a machine built by Anthropic to investigate Anthropic. It told me my theory was wrong, and my theory was wrong. I went and checked the dates myself anyway.

Both of those sentences matter, and the second one is the one I would keep.


This is a note about how I work, not advice about how you should. I build software at Revelations Technology and write these as I go.

Share on LinkedIn
Joe Baker
Joe Baker — Software architect with 35 years of experience. Currently SVP Software Engineering at WellSky. Connect on LinkedIn.

Read next

All posts