A professional standing above an operations floor of AI screens and data panels — human judgement above AI execution
Founder Note

The Problem With Letting AI Mark Its Own Homework

12 min read

The same system that created the answer has been asked to validate the answer. That is not governance. That is AI marking its own homework.

Phillip Llewellyn·Founder, Human Heartbeat AI·

There is a mistake I see more and more businesses making with AI.

They ask AI to write something.

Then they ask the same AI to check it.

Then, because the answer sounds confident, tidy and professional, they assume the work has been properly reviewed.

It hasn't.

What has usually happened is much simpler and much more dangerous.

The same system that created the answer has been asked to validate the answer. It is looking at the work through the same lens, the same assumptions, the same context and the same blind spots that produced the original output.

That is not governance.

That is AI marking its own homework.

The AI Self-Review Loop — No Governance: AI Creates, AI Validates, AI Acts with no Human Decision Gate
This is the pattern I see most often. The same system creates the output and validates it. There is no separation, no challenge, no independent check. It looks like a process. It is not.

And for businesses, that distinction matters.

Not because AI is useless. Far from it.

AI is becoming one of the most powerful business support layers we have ever had. It can draft, analyse, summarise, compare, challenge, structure and accelerate work at a level most businesses could not have imagined even a few years ago.

But power without boundaries is not intelligence.

It is risk with a better user interface.

The Real Issue Is Not "How Do We Use More AI?"

Most businesses are still asking the wrong question.

They ask:

"How can we use AI?"

Or:

"What tools should we buy?"

Or:

"How do we automate more?"

Those are not bad questions, but they are not the first questions.

The first question should be:

Where is judgement required?

Because once you know where judgement is required, you can start separating the work properly.

Some tasks are drafting tasks.

Some tasks are review tasks.

Some tasks are evidence tasks.

Some tasks are decision-support tasks.

Some tasks are operational tasks.

Some tasks should never be handed to AI without human review.

That is where most AI adoption goes wrong. Businesses do not fail because they lack tools. They fail because they blur the roles.

They let the same AI draft the recommendation, validate the recommendation, summarise the risk, propose the action and sometimes even trigger the next step.

That may feel efficient.

It is not safe.

And in a real business, "efficient" is not the same as "governed."

Most businesses are asking the wrong question — three wrong questions vs one better question: Where is judgement required?
The wrong questions feel productive. They feel like progress. The right question feels slower — because it requires you to think about your business, not just your tools.

Fresh Eyes Matter

In human work, we understand this instinctively.

The person who wrote the proposal should not be the only person who signs it off.

The person who built the spreadsheet should not be the only person checking the numbers.

The person who wants the campaign to launch should not be the only person reviewing the compliance risk.

Why?

Because creation and review are different jobs.

They require different mental postures.

A creator is trying to make the thing work.

A reviewer is trying to find where it might fail.

That separation is not bureaucracy. It is protection.

The same principle applies to AI.

If one AI Worker drafts a piece of client-facing content, another AI Worker may be useful for critique.

If one AI Worker summarises evidence, another may be useful for checking whether the conclusion is overreaching.

If one AI Worker proposes a workflow, another may be useful for identifying operational risk.

But even then, the answer is not simply "add more AI."

That is where I think many people will get agentic AI wrong.

They will hear about subagents, panels, parallel execution and specialist AI roles, and they will assume the answer is to spawn more intelligence.

More AI voices.

More outputs.

More opinions.

More speed.

But without governance, all you have created is a louder room.

Creation and Review are different jobs — Creator AI Worker vs Reviewer Separate Role
This is not complicated. It is the same principle a good editor uses. The person who wrote it cannot be the only person who reads it. Separation is not bureaucracy. It is how quality actually works.

The Business Risk Is False Confidence

The real danger is not that AI makes mistakes.

Humans make mistakes too.

The real danger is that AI makes mistakes in a way that looks finished.

It writes fluently.

It sounds structured.

It creates confidence.

It makes weak assumptions feel like strong conclusions.

And when several AI agents all produce polished outputs, the risk increases again. A business can mistake volume for validation.

Five AI Workers agreeing with each other does not automatically mean the answer is right.

It may only mean they were all given the same flawed instruction, the same incomplete evidence, or the same poorly defined authority.

That is why Human Heartbeat AI is built around a different principle:

AI should support decisions, not silently become the decision-maker.

That principle changes everything.

It means AI Workers can draft, analyse, compare and recommend.

But they do not get unlimited authority.

They do not replace human accountability.

They do not quietly turn a suggestion into an action.

They operate inside a governed system.

The False Confidence Gap — AI Confidence Level 92% vs Verified Accuracy 61% — Fluency is not the same as correctness
AI does not hesitate. It does not say it is not sure. It presents. That presentation creates a feeling of certainty that the underlying accuracy does not always support. This gap is where decisions go wrong.
“

Without governance, all you have created is a louder room.

Why Separation Matters

In our world, I do not see this as "using subagents."

That language is useful technically, but it is too small commercially.

The bigger issue is role separation.

A business does not need a swarm of clever AI tools all talking over each other.

It needs a governed operating model.

That means clear separation between:

Runtime — the AI Worker or agent execution layer.

Canon — the source of truth, rules, approved decisions, evidence and boundaries.

Orchestration — the layer that coordinates the workflow and makes sure the right role does the right job at the right time.

And above all of that:

The Human Decision Gate.

That is the point many AI conversations still miss.

They focus on how to make AI do more.

I am more interested in why AI should be allowed to do a thing in the first place.

What evidence is it using?

What authority does it have?

What must it not do?

What requires review?

What counts as a recommendation?

What counts as a decision?

Who is accountable if the output is wrong?

Those are not theoretical questions.

They are commercial questions.

They are operational questions.

They are governance questions.

And for SMEs, they are survival questions.

The Governed AI Operating Model — four layers: Runtime, Canon, Orchestration, Human Decision Gate
Four layers. Each one doing a different job. The Human Decision Gate is not at the bottom — it sits above everything. That positioning is deliberate. Authority does not live inside the machine.
Why AI Should Not Mark Its Own Homework — governance architecture diagram showing Human Decision Gate above Orchestration, Canon, and Runtime layers
The full architecture in one diagram. Every layer has a role. The Human Decision Gate is not a bottleneck — it is the point where accountability lives. Remove it and you do not have a faster system. You have an unaccountable one.
“

AI should support decisions, not silently become the decision-maker.

AI Workers Are Not Magic Staff

There is a tempting fantasy emerging around AI agents.

The fantasy says you can create an army of AI workers, point them at your business and watch everything become faster, cheaper and smarter.

I understand the appeal.

But I do not believe that is how serious businesses should adopt AI.

AI Workers are not magic staff.

They are not independent executives.

They are not accountable employees.

They are not commercially responsible decision-makers.

They are bounded business support roles.

That distinction matters.

An AI Worker can help prepare a report.

It should not decide whether the report is fit to send.

An AI Worker can identify risks.

It should not decide that the business accepts those risks.

An AI Worker can recommend a workflow.

It should not silently implement that workflow without authority.

An AI Worker can support customer communication.

It should not make commitments the business has not approved.

This is why the Human Decision Gate is not a nice-to-have. It is the line between assistance and uncontrolled delegation.

Bounded Business Support Roles — AI Workers CAN vs AI Workers MUST NOT
The left column is where AI is genuinely useful. The right column is where businesses get into trouble. The line between them is not technical. It is a governance decision — and it needs to be made consciously, not by default.

The Wrong Way To Use Agentic AI

The wrong way is simple.

Give AI broad access.

Give it vague instructions.

Let it create, review and act.

Trust the output because it sounds intelligent.

Then call that transformation.

It is not transformation.

It is operational drift.

Operational Drift — AI touching the actual business without governance: Inbox, CRM, Client Proposals, Quotes, Advice, Follow-ups
This is what ungoverned AI adoption looks like six months in. Not one big decision. A series of small ones, each reasonable in isolation, that collectively move AI from support tool to operational actor. By the time it is visible, it is already embedded.

It is especially dangerous for SMEs because they often do not have large compliance teams, internal audit departments or dedicated AI governance functions.

So when AI goes wrong, it does not go wrong inside a controlled enterprise risk framework.

It goes wrong inside the actual business.

In the inbox.

In the CRM.

In the client proposal.

In the quote.

In the advice.

In the follow-up.

In the operational promise made to a customer.

That is why "getting started with AI" is not enough.

Businesses need to get started safely.

The Better Question

The better question is not:

"How many AI agents can we use?"

The better question is:

"What should each AI Worker be allowed to do, and where must the human remain in control?"

That question creates a very different adoption path.

It slows down the right things.

It speeds up the safe things.

It avoids pretending that every AI use case is equal.

It separates drafting from deciding.

It separates analysis from authority.

It separates recommendation from execution.

And that is where real business value starts to appear.

Because when AI is properly bounded, it becomes more useful, not less.

People sometimes think governance slows AI down.

I think the opposite is true.

Good governance allows you to move faster because you know where the edges are.

You know what AI can touch.

You know what it must not touch.

You know what needs review.

You know what evidence is trusted.

You know what decisions are locked.

You know when a human must step in.

That is what creates confidence.

Not blind confidence.

Operational confidence.

The Wrong Question vs The Better Question

✗ The Wrong Question

“How can we use more AI?”

“What tools should we buy?”

“How do we automate more?”

✓ The Better Question

“Where is judgement required?”

“What should each AI Worker be allowed to do?”

“Where must the human remain in control?”

Why OSCAR Comes First

This is exactly why we built OSCAR as the starting point.

Before a business rushes into automation, AI Workers, agents, chatbots or workflow tools, it needs to understand its current AI posture.

Where is AI already being used?

Where are people experimenting informally?

Where are decisions being made without clear ownership?

Where is client data exposed?

Where are workflows undocumented?

Where is the business relying on human judgement but failing to define where that judgement must sit?

Those questions matter more than the tool choice.

Because if the foundations are weak, AI does not fix the business.

It amplifies the weakness.

If the enquiry process is unclear, AI makes unclear communication faster.

If the decision rights are unclear, AI makes unclear decisions faster.

If the source of truth is messy, AI spreads messy information faster.

If the governance is missing, AI creates the illusion of control while the business quietly loses it.

That is why the first step is not "build an AI Worker."

The first step is diagnosis.

Then governed implementation.

Then scale.

My View

I welcome the growing conversation around subagents, specialist AI roles and multi-agent workflows.

It is a useful sign that the market is beginning to understand something important:

One AI instance should not be trusted as creator, reviewer and decision-maker all at once.

But for me, that is only the beginning.

The real opportunity is not to create more AI activity.

The real opportunity is to build better business judgement into the way AI is adopted.

That means role separation.

Evidence discipline.

Clear authority boundaries.

Human review.

Source-of-truth control.

And above all, a business owner who remains accountable for the decision.

Because the future of AI in business should not be autonomous chaos dressed up as innovation.

It should be intelligent support under human governance.

That is the heartbeat.

Not AI replacing judgement.

AI supporting judgement.

Not faster decisions at any cost.

Better decisions, with the human still in the gate.

That is where I believe responsible AI adoption has to go next.

Share this Founder Note

Pass the argument to someone who should be part of it.

Continue through the Founder Notes series.

Explore the work