Skip to content

We Published 100 AI-Generated Posts. Here’s What Actually Worked.

100 AI-generated posts experiment results — what worked and what broke in a 90-day content pipeline

1. Opening Hook

Everyone tells you AI content is fast. Nobody tells you what breaks.

In 90 days, our team of three humans and three AI agents published 100 pieces of content: 37 blog posts on our own site, 60+ client deliverables, and a steady drumbeat of social derivatives, newsletters, and video scripts. We didn’t set out to write a case study. We set out to see if an AI-native content pipeline could actually work at scale.

Some of it did. A lot of it didn’t. What follows is the honest scorecard — not a traffic brag, because we don’t have a public analytics dashboard and it’s too early for meaningful conversion data. What we do have is operational truth: what worked, what broke, and what we’d do differently on day one.

2. The Setup: What “100 Posts” Actually Means

Let’s be transparent about the number, because rounded headlines destroy credibility.

Category Count Notes
Content Factory blog posts ~37 Published May 1 – Aug 5, 2026
Client deliverables 60+ Includes dumMySQL launch content, other client blog posts, landing pages
Social derivatives & newsletters Ongoing LinkedIn long-form, X threads, email newsletters, video scripts
Total 100+ 90-day window

The team: three humans (strategy, editing, QA) and three AI agents (research, drafting, repurposing). Our typical weekly output settled at roughly four articles once the pipeline stabilized. Peak cadence was higher in the first 60 days — we were testing limits.

The “100” is real, but it’s a composite. If you’re reading this hoping for a pure blog benchmark, ours is ~37. The lesson isn’t the number. It’s what happens when you try to hit it.

3. What Worked

The 3.4× output multiplier is real — with caveats

AI-assisted teams consistently produce 3–5× more content without additional headcount. Industry benchmarks put the figure at 42% more content per month; our lived experience lands closer to 3.4× once the pipeline was calibrated. The multiplier doesn’t appear on day one. It appears after you fix the handoffs.

Editing time collapsed — after we restructured agent output

Our earliest agent drafts required an 80% rewrite. The human editor was essentially writing the post twice. After we restructured agent output into “annotated drafts” — sections labeled with intent, source confidence, and open questions — editing time dropped from ~90 minutes per post to ~35 minutes. That’s a 61% cut.

Specificity sells

Our most-shared post wasn’t a listicle. It was “Why I Stopped Hiring and Started Training Agents Instead” — a first-person operational story with a controversial headline and specific numbers. Readers don’t share generic advice. They share specifics they haven’t heard before.

The quality rubric changed everything

We defined “good” explicitly across six dimensions: Accuracy, Voice Consistency, Engagement, Conversion, Originality, and Helpfulness. Threshold: no score below 3, at least one dimension scoring 5. Before the rubric, quality was a vibe check. After the rubric, it was a pass/fail gate — and output consistency improved immediately.

Capping review load protected quality

We learned the hard way that reviewing five posts simultaneously is cognitively expensive. The bottleneck wasn’t the AI; it was the human brain context-switching between drafts. We capped final review at two posts per day. Output dropped slightly. Quality jumped significantly.

4. What Broke

Agent-human handoff friction

The first three weeks were brutal. Eighty percent of agent drafts were rewritten from scratch. The problem wasn’t the AI’s writing ability — it was the handoff clarity. The agent didn’t know what the editor needed, and the editor didn’t have time to teach it. We fixed this by adding labeled sections, source annotations, and explicit “decision required” flags to every draft.

The editorial calendar became a liability

We planned content two weeks ahead. Forty percent of it was wrong within 72 hours. AI news moves fast, and a pre-planned post about a model release becomes obsolete when the next version drops three days early. We moved from a rigid calendar to a “priority queue” system: a running backlog of approved topics, pulled in order of relevance, not date.

Context-switching tax

Reviewing multiple drafts in one sitting felt efficient. It wasn’t. The cognitive load of jumping between voices, topics, and argument structures degraded our judgment. The fix — capping review at two posts per day — wasn’t a productivity hack. It was a quality survival mechanism.

Voice consistency gaps

Readers could feel the difference between fully human-written posts and agent-assisted ones. Not because the AI writing was bad, but because the transitions, rhythm, and opinion were slightly off — the “last mile” AI still can’t run. Human editing for voice isn’t optional. It’s the difference between content that sounds like you and content that sounds like everyone.

The missed post

On July 1, the pipeline broke. Overcommitment collided with a client deadline, and we had a choice: ship a half-baked post or miss the date. We chose silence over slop. One missed post taught us more about sustainable cadence than twenty on-time ones.

5. What the Industry Data Says

Our experience isn’t unique. It’s part of a broader pattern that industry data confirms — and complicates.

The AI slop backlash is real.

Gartner found that 49% of consumers say GenAI made content quality worse. Bynder’s research shows 52% become less engaged when they suspect AI-generated content. And the March 2026 Core Update hammered unedited mass AI content with 60–80% traffic drops. The penalty isn’t for using AI. It’s for publishing unhelpful content.

What actually ranks and engages:

  • Human-edited AI content generates 4.1× more engagement than raw AI output
  • Unedited AI has a 23% higher bounce rate than human-edited AI
  • 86.5% of top 20 search results contain some AI-generated content — but only 9% of #1 rankings are purely AI
  • Original data and first-hand experience remain the #1 differentiator

The pattern is clear: AI accelerates production. Human judgment determines whether that production is worth reading.

6. The Honest Math: ROI

Time savings vs. compute costs

The compute costs are trivial. The human time isn’t. Our editing time dropped 61%, and our total output multiplied 3–4×. But that multiplier required upfront investment: building the rubric, restructuring handoffs, training the team on prompt engineering, and accepting slower initial output while the pipeline calmed.

Industry benchmarks suggest 420% average ROI for AI content tools with a median payback period under six months. We’re too early for a definitive internal number, but the directional signal is strong: if you have the human oversight layer, the economics work. If you don’t, you’re publishing slop at scale.

7. What We’d Do Differently

  1. Start with the rubric on day one. Defining quality after you’ve published 20 posts means rewriting 20 posts.
  2. Build handoff clarity before pipeline speed. The fastest draft is worthless if the editor rewrites it.
  3. Cap review load from week one. Context-switching is the hidden bottleneck in every AI content operation.
  4. Communicate missed deadlines publicly. The July 1 miss wasn’t a failure. Pretending it didn’t happen would have been.
  5. Ditch the rigid calendar sooner. In fast-moving fields, a two-week editorial calendar is a liability, not a plan.

8. Closing

Speed is a tool, not a goal.

The future isn’t AI versus human. It’s AI plus human with explicit quality gates. The teams that win won’t be the ones that publish the most. They’ll be the ones that publish the most useful — and that requires humans at the decision points AI can’t reach.

If you’re building an AI content pipeline, start with the rubric. End with the reader. Everything in between is optimization.


Want to talk about building your own AI content pipeline? Get in touch.