19.7 C
London
Wednesday, September 16, 2026

What Happens When AI Builds and Tests Its Own Workflow?

BusinessWhat Happens When AI Builds and Tests Its Own Workflow?

A system that can draft its own next step sounds like a closed loop. In practice the loop is only as good as the tests, the record, and the person who defined what “good enough” means before the system started.

When AI builds and tests its own workflow, three things are happening at once. A model is proposing work. Another process, sometimes a second model, sometimes a script, is checking that work. A human is still responsible for the standard both of them are using, even if that human is not in the daily loop.

Ken Ashe’s autonomous Digest at KenAshe.ai is a public version of this. The pipeline researches, drafts, reviews, and publishes only what clears a quality gate. Every post carries a disclosure that it was written and published by a system he built, with Ken named as the accountable publisher.

Generation without a gate is just a faster first draft

If the same model writes the post and then compliments the post, you do not have a test. You have a monologue. A gate has to check something the writer could fail: sources present, duplicate of yesterday, forbidden claims, schema complete, image rules met.

Ken’s later Digest pipeline is explicit about that. Ingest, embed, cluster, deduplicate against memory, rank, select, write, gate, publish. The interesting word in that list is memory. Without a record of what already ran, a self-testing system will pass the same work twice and call it consistency.

Self-testing fails in predictable ways

The test can share the writer’s blind spot. If both steps are bad at citations, they will agree.

The test can reward fluency. A confident paragraph is easier to score as “fine” than a short, sourced one.

The test can drift. If nobody looks at a sample of passed and failed items, the gate learns nothing. It just runs.

This is why the accounting angle belongs in a piece about self-testing workflows. Keep canonical state outside the agents. Make claims checkable against a record. Evaluate outcomes rather than the system’s explanation of itself. Those are audit habits. They are also the only habits that make a self-testing loop inspectable.

What still belongs to a person

Someone writes the list of things the system may not do unsupervised. Publishing a claim about a company, sending a message in someone else’s name, moving money, those are reasonable items for that list.

Someone reads a digest of what the system did. Ken has said that on most days the Telegram digest is the part he reads. Oversight that small only works if the system is designed to surface the exceptions rather than hide them in volume.

Someone remains publisher. The disclosure is how readers are told that. Removing the disclosure would make the loop look cleaner and the accountability worse.

What this does not prove

A workflow that can test itself does not prove that AI no longer needs people. It proves that a person can push their involvement up a level: from typing the draft to defining the gate and owning the output.

It also does not prove that any team should turn their publishing, support, or operations into an unsupervised loop tomorrow. The Digest is one system, labeled, with a public failure history that includes a first version that started repeating itself. That history is the part worth copying, not a generic enthusiasm for machines that grade their own homework.

How to write a gate that is not a compliment

The gate should check facts the writer could get wrong.

Required fields present.

Overlap with published items below a threshold.

Forbidden topics absent.

Source list non-empty.

Image rules met.

Human-only actions not invoked.

Score fluency last, or not at all.

If the gate cannot name a failure it caught last week, it is not a gate yet.

Operational cadence

A daily system needs a daily artifact for the owner that is shorter than the output. Ken’s Telegram digest is one shape. An email with rejects and anomalies is another. The cadence is part of the workflow. Without it, “unsupervised” means “unobserved.”

Limits of the Digest as a model for others

It publishes on Ken’s site, about topics he chose, with his name on the accountability line. A company newsroom, a client blog, or a regulated disclosure feed has different costs. Copy the label and the gate idea. Do not copy the cadence until you have copied the cost analysis.

The February 2017 Journal of Accountancy feature and the 2016 Investing.com pieces are background on the publisher, not on the pipeline.

Self-testing workflows need versioned prompts and versioned gates. When output quality drops, you have to know what changed: the model, the sources, or the instructions. Without versions you will argue from memory. With versions you can roll back. That is ordinary release discipline applied to text.

What the gate is for

A self-testing workflow is easy to oversell. The honest version is narrower. One process proposes work. Another process is allowed to reject it. A person wrote both rules and still owns what gets out.

The gate has to catch a failure the writer could actually have. Missing source. Duplicate of yesterday. A claim the system is not allowed to make. A required field left blank. If the only thing the gate scores is whether the prose sounds finished, it will pass fluent mistakes.

Ken’s Digest is useful here because the loop is visible. Ingest, cluster, write, gate, publish. Later versions added memory so the system could notice it was repeating itself. The first version did not have that, and by week three it showed. The rebuild is part of the story. A self-testing system that cannot remember what it already did will congratulate itself for the same work twice.

Oversight can be small if the exceptions surface. A short daily digest of what published, what the gate rejected, and what looked unusual is enough for an owner who is not reading every draft. Without that artifact, “the system tests itself” means “nobody looked.”

None of this proves that every team should take their hands off publishing. It proves that a named person can move from typing the draft to defining the fail conditions. That is a different job. It is still a job.

Check out our other content

Check out other tags:

Most Popular Articles