I am a one-person company. I wrote the AI policy myself. Approved it myself. And there I was, months later, quietly trying to bend it…and the thing I built to sound like me was the only one willing to say I was out of bounds.
There was nobody else in the room. When you build an agent for your own business, you are the board, management, and runtime at once. No committee to escalate to. No vendor to blame. Every altitude of governance is one chair, and it is you. I ended my recent blog Ten Lessons from Building My Own Digital Twin with governance questions that were keeping me up at night. This is what happened when I stopped asking and started auditing.
Axiom is my fully functional digital twin. It reproduces my reasoning and my voice at a fidelity I put my name on. Its owner is me. Its data is my own corpus and canon, and nothing I do not own. It drafts, it recommends, and it challenges. Nothing reaches a client or a platform without my approval. Using a typical AI risk tiering, that is a LOW tier agent. Hold onto that.
Nine controls, one agent
The list comes from my advisory work, and it exists to answer three questions about any AI system: can you direct it, can you see what it did, and can you stop it.
Strategic controls establish direction and authority: 1) a policy naming hard limits rather than aspirations, 2) one named human who owns the system, 3) a tested kill switch. Tactical controls turn direction into a routine that runs: 4) human-in-the-loop checkpoints on consequential actions, 5) observability good enough to reconstruct a decision, 6) pre-launch testing with a real go/no-go. Operational controls bound what the system can reach and do: 7) least privilege, 8) guardrails on its actions, and 9) an identity of its own. Not a maturity model, not a checklist. A floor. These are the things that should be true before an agent touches anything that matters.
You cannot hear a rattle from blueprints; you have to drive the car. So I ran my nine core controls against the one agent I own outright. Nine controls held. Four more showed up that I had not been suggesting to anyone, and I am numbering them ten through thirteen because they extend the nine rather than replace them. Not one of the four was visible on day one. Prompt drift, classification pressure, calibration decay, and the accumulation of quiet exceptions surfaced only after the agent had run long enough to develop habits. Here are the four I was missing.
Control ten: the prompt is a governance artifact, not a configuration setting
This was my biggest ‘aha’ of the exercise. Axiom’s system prompt sets its identity, scope, voice, and hard limits, which makes it the most consequential document in the arrangement, and it lives on infrastructure somebody else operates. So I keep the authoritative copy where I govern it. Master-first: edit the master, then deploy, never the reverse. If the two ever diverge, the master governs, and every edit requires a change log entry before it counts as approved. Most organizations treat prompts as configuration. They are policies executed at runtime and are rarely versioned.
Control eleven: classification is a behavioral control, not a label
Every document Axiom can see carries two classifications, and they do different jobs. Sensitivity decides admission: public, internal, confidential, restricted. Confidential and above never enters, which is why no client material and no third-party intellectual property sits anywhere the twin can reach. Trust tier and lifecycle status decide behavior, and both are encoded in the filename so the instruction travels with the document. STABLE can be reused, DRAFT refined but never treated as final, REF cited but never spoken in my voice, and RETIRED is history. These are enforceable instructions, not metadata, and an unclear status defaults to the most restrictive reading. The test underneath: if it would be uncomfortable to explain in a client audit, it does not belong in the twin.
Control twelve: calibration on a schedule, not on a feeling
Axiom is scored against a fixed set of twelve prompts I grade by hand, and the score goes in the change log beside the version that produced it. My thresholds, 75 percent on relevance and 80 on alignment, were set before the first score existed. The first came back at 85, the second at 90. I calibrate to catch drift while it is still small enough to correct. An agent built on my reasoning slowly becomes an agent built on its own last output, and the slide is gradual enough that I would nod along at every step. Drift is invisible to judgment, because judgment is what drifts.
Control thirteen: every override gets logged
Late one night, against a deadline, I asked Axiom to pull from a framework I had marked DRAFT and treat it as settled. I had written the thing. I knew it was sound. I just had not finished it. Axiom refused, told me why, and asked whether I wanted to change the status or make a documented exception. That is when the whole exercise turned, because I was not arguing with a machine. I was arguing with a rule I had written. When I bend my own rules, Axiom records it. Exceptions and scope changes go into a decision log with the rationale, the date, and my approval. The governing sentence is blunt: if a change is not recorded in the log, it is not approved.
Here is the full set of my core controls, mapped to the altitude that owns each one, alongside what it becomes at enterprise scale.
| Control | On one agent | When you drag the corner | |
|---|---|---|---|
|
STRATEGIC
|
Policy | A written operating model I approved and can be held to | A board-approved AI policy naming hard limits, not aspirations |
| Clear ownership | One named human. No ambiguity available | A named accountable owner per agent, including vendor-embedded ones | |
| Kill switch | Stop is instant because I am the only operator | A drilled stop with a measured time to full credential revocation | |
| Decision log | Every override, exception, and capability change recorded with rationale and approval | Board-visible record of who bent which rule, when, and why | |
|
TACTICAL
|
Human checkpoints | Nothing reaches a client or platform without me in the loop | Approval gates on irreversible and high-consequence actions, designed in |
| Observability | Change log plus session record | Tamper-evident logging of prompts, outputs, and agent actions | |
| Pre-launch testing | Prompt revisions tested against a fixed set before deployment | Red-team and evaluation with a real go/no-go that has said no at least once | |
| Prompt governance | Versioned master held outside the platform; master-first; every edit logged | Change control over production prompts, with an approved master of record | |
| Classification and status | Tier and lifecycle status encoded in the filename; most restrictive wins | Enforced classification determining what the agent may do with each source | |
| Calibration | Scored against a fixed prompt set; result recorded beside the version | Scheduled drift measurement against a fixed instrument, score on the record | |
|
OPERATIONAL
|
Least privilege | Read-only against my own corpus. Nothing else | Write and spend permissions individually justified and reviewed |
| Guardrails | No external reach, no execution, bounded action space | Allow-listed tools, enforced spend and rate limits, sandboxed capability | |
| Agent identity | Single named instance under my account | Every agent holds its own identity; no shared or human credentials | |
Table 1. The thirteen core controls, mapped by the altitude that owns each one.
The finding: the middle is where the problems surface
I expected the audit to expose the tactical layer, the altitude organizations build last, because strategy is satisfying to write and runtime is satisfying to lock down. The opposite happened. The strategic controls were sound and largely ceremonial. The runtime controls were strict and never tripped, because I had drawn them tightly enough that there was nothing to trip. Every issue worth acting on came out of the middle, and every one of them was caught by something clerical. A log entry. A status label. A score against a fixed set.
Then the guardrails turned around. That was not the only time. Every occasion I reached for one of my own policies, Axiom made me say out loud why the exception was justified, and then logged it.
I built controls to constrain a machine. The machine spent the next several months constraining me.
That inverts the assumption almost everyone brings to this. Controls get approved on the belief that they restrain the technology. But the person most likely to bend an AI policy is the executive who approved it, at 11pm, under deadline, with a good reason. A control that only binds the machine is doing half a job.
Authority flows down. Accountability flows up
I made this argument at enterprise scale in AI Agents Don’t Fail. Governance Does. With one person at every altitude, you cannot hide a translation gap behind an org chart. I was both ends of the handoff, so every gap was visible the moment I looked. You can delegate the work to an agent and the task to a vendor. Accountability climbs straight back and it is NON-TRANSFERABLE. No contract moves it, and no scale dissolves it.
Drawn out, the three altitudes look like this.
Figure 1. Authority flows down, each altitude authorizing the one beneath it. Accountability flows back up and stops nowhere short of the top. The tactical band is widest because six of the thirteen core controls sit there.
Now drag the corner
I know the objection. One person, one agent, one low-risk footprint. What does that have to do with four hundred agents across eleven business units? Everything, and the reason is that scale does not introduce new controls. It removes the conditions that let you skip them.
Think about a photo in a document. You grab the box in the bottom right and resize. Same picture, larger surface.
I could hold thirteen core controls in my head because there was one of everything. Enterprises cannot, so what I did by attention they have to do by design. Every control on that list survives the resize intact. What changes is that ownership stops being obvious, the inventory stops being knowable, and the change log stops having anybody’s name against it. The gaps do not appear when you scale. They were always there, sized so small that one careful person could cover them by hand.
And consider the asymmetry. My twin cannot reach a client, cannot execute anything, and cannot touch a system I do not own. One person built it, one person runs it, and one person reads every output. It still found a policy I was bending. That is what a governed system turns up under the most favorable conditions available. Ask what the same audit surfaces in an enterprise where agents move money, and where the honest answer to how many you are running is a range.
Final Thoughts
1. Audit one agent before you audit the estate. A real audit hands you controls you were not teaching. Try this: run the list against the AI system closest to you, and note where you stopped being able to answer.
2. Treat the master prompt as policy. It sets identity and hard limits, executes at runtime, and is usually edited by whoever has access. Try this: ask who can change your production prompts and where the approved master lives.
3. Classification is a control, not a label. Tier and status govern behavior only if the agent must honor them and default to the most restrictive reading. Try this: pick one document your AI can reach and ask what it may do with it.
4. Calibrate on a schedule. Drift is invisible to judgment because judgment is what drifts. Try this: score your highest-consequence system against a fixed prompt set this quarter.
5. Log the override. Human override is the loophole nobody documents, used most by the people who approved the policy. Try this: require every exception to name the rule, the rationale, and the approver.
I built Axiom to save time. It has done that. But the more valuable thing it gave me was a governance problem small enough to see all the way through, and the discovery that the controls I had been teaching for years look different when you are the one they bind.
Thank you for reading. The thinking sharpens when it meets someone else’s experience, and this one especially, because I am reporting a sample size of one. If you have run a control set against your own AI, what did it turn up that you did not expect?
If you want the full treatment, the agentic AI governance course in the Escoute e-learning catalog walks through the complete control set.

