Skip to content
← All blogs

What our AI agents get wrong most often

A red-circled 27% statistic points to a missing footnote, examined through a magnifying glass

We build AI content workflows for clients, and we run our own content through one. Our AI Agent Content Factory (AACF) passes each article through a chain of agents. One picks the topic, two research it, one writes, one critiques the draft and one edits it. Then a person decides whether it's ready.

Between June and September 2026, the Critic agent reviewed seven drafts in our system. We went back through every report to see what it caught. This is the honest list: the mistakes our agents made, the mistakes the Critic made itself, and what we changed.

The pattern in one sentence

The agents rarely failed at writing. They failed at telling the difference between what they had evidence for and what merely sounded right.

All seven reports rated research accuracy as a critical problem, with scores between 4 and 6 out of 10. In the same reports, readability scored between 6 and 8, and structure between 7 and 9. The drafts read well. That's exactly why the accuracy problems were easy to miss.

1. Numbers that sound like research

The most common problem was a specific number with nothing behind it.

One draft told readers that "most lean marketing teams spend 3–6 hours every week" on manual reporting. The same draft priced a RevOps hire at "$90,000–$130,000+ per year". Another set hard thresholds, such as open rates below 15% or more than 30% of CRM records missing key fields, without a single source.

None of these numbers is absurd. That's what makes them dangerous. A plausible number reads like a researched one, and the reader has no way to tell the difference.

2. Real statistics, stretched

Sometimes the source was real, but the claim grew on its way to the page.

In one run, a 27% figure from the research turned into a statement about lost revenue in the article's meta description. The source didn't say that. In another, a Gartner figure about the cost of poor data quality was presented to small-business readers as if it described them.

This is harder to catch than an invented number, because the link is there and it works. You only see the problem when you open the source and read what it actually measured.

3. Words and sources that didn't exist

One draft about CRM data introduced the term "polite decay" and used it as if it were established industry vocabulary. It isn't. The model coined a phrase, then treated it as something the reader should already know.

Another draft listed "AI Maturity Gap (concept)" and "Prompt Chaining and AI-Agentic Workflows (concept)" under Sources. There was no author, no publisher and no link. At a glance they looked like references. They were nothing of the kind.

4. The quiet, mechanical misses

Not every problem was about truth. Several drafts left internal links as placeholders, missed the main search phrase in the headline, or used terms like CAC and attribution without explaining them. One leaned on a framework, "Identify, Automate, Optimize", that the Critic rightly called generic.

These are the cheapest problems to fix, and the easiest to publish by accident.

The Critic got things wrong too

This was the most useful part of the review.

In July, a draft mentioned a current DeepSeek model. The Critic called it "highly suspicious, likely hallucinated" and stated that DeepSeek only had V2, V3 and R1. It told the Editor to replace the reference with "accurate, widely recognized models", and suggested ones that were already a generation old.

DeepSeek V4 Pro is the default model for most of the agents in AACF. The Critic wasn't checking facts. It was checking them against its own training data, which ended before the model was released.

In another report, the Critic correctly flagged the unsourced "3–6 hours" claim. Then it offered a fix: rephrase it as "Our internal audit of lean B2B teams often shows 3–6 hours per week." We had run no such audit. The reviewer's fix would have turned an unsourced number into a false claim about us.

An AI reviewer is a good second reader. It is not a fact-checker. It can be confidently wrong in both directions: flagging true things as false, and proposing fixes that are worse than the problem.

What we changed

Each failure led to a specific rule. The current versions of our agent instructions include rules the earlier versions didn't have.

What went wrongWhat the agents must do now
Numbers with no sourceStatistics, quotes, examples and URLs may only come from the research brief. Unsupported trends are marked "Needs validation".
Thin research padded to look completeThe research agent opens with an evidence status: how many usable sources it found and how strong they are. If evidence is thin, the brief stays short.
Sources that weren't sourcesThe Writer may only link to URLs that appear in the research or SEO brief. It is told never to create a URL.
Editor filling gapsIf a brief is missing, the Editor may keep, soften or cut a claim, but never add a new fact or link.
Critic suggestions followed blindlyThe Editor is told not to apply a criticism if it would make the article less accurate.

The research step has guards of its own. Each web search uses at most four sources. The system throws out a research brief if the same sentence appears in it four or more times.

What a person still has to do

The rules helped. They didn't make review optional. In our two most recent runs, at the end of September, the Critic still found claims like "the most common failure" stated as fact without a source, and a how-to article resting on just two sources.

So every article still ends with a person. That person checks three things the agents can't:

  1. Open every link. Does the source say what the sentence says it does?
  2. Question every number. Could you tell a client exactly where it came from?
  3. Ask whether we'd say it. Is this something we've seen in our own work, or something that only sounds like it?

That's the setup we recommend to clients too. Let agents do the research, the drafting and the first review, and keep a person at the point where the content goes public.

You can see how the agents fit together in the AI Agent Content Factory demo. If you're building something similar and want a second pair of eyes on it, request a free audit.

← All blogs