Agent-Readiness Audit

A straight answer on whether your team can trust an agent to ship — and the first move if it can’t.

Two weeks. $8,500 fixed. A two-week, fixed-scope diagnostic that scores your org on the six axes that actually decide whether agent work survives contact with production — plus the cheapest, highest-leverage first move if it can’t.

The problem

Between 78% and 88% of enterprise agent pilots never make it from demo to production. The demo works. Then it meets your real codebase, your CI, your review culture — and it stalls. The pilot gets shelved, and leadership quietly concludes “AI” was overhyped.

The failure is almost never the model. It’s three things nobody put in place before turning the agent loose: verification (how a run proves it’s actually done), evals (how you know a change is correct, not just green), and operational ownership (who governs what the agent is allowed to do unwatched). Pilots that scale spend proportionally more on exactly those three. Pilots that die skipped them.

This audit tells you, in plain terms, where your org stands on all three — and what the cheapest, highest-leverage first move is.

What it is

A two-week, fixed-scope diagnostic. $8,500, flat — no hourly meter, no six-figure strategy deck. About 16 hours of my time; most of the repo and CI archaeology is done by my own agent fleet, so the calendar stays short and the price stays honest. Scope is one engineering org (up to ~50 engineers) or one product group inside a larger one.

You need to give me read-only access to one or two representative repos, your CI config, and whatever agent or assistant usage data already exists. That’s it.

The 6-axis scorecard

I score your org 0–4 on each of the six axes that actually decide whether agent work survives contact with production. Every score is backed by evidence cited to your own repos — not a vibe.

1

Verification gates

Does a run have to prove it’s done (structured completion contract, CI it can’t lie to), or does “opened a PR” count as finished?

2

Merge-safety architecture

Branch/worktree isolation per run, and an explicit auto-merge policy — or is it all-or-nothing against main?

3

Trust-ladder maturity

Is autonomy graduated and revocable (what runs unsupervised vs. reviewed, per surface), or all-or-nothing?

4

Eval infrastructure

Pre-registered success criteria, blind adjudication, contamination controls — or “it looked right”?

5

Tool agent-legibility

Are your internal CLIs and APIs actually drivable by an agent, or built for humans only? (Scored against the aclig.dev conformance lens.)

6

Operator surface

Is someone babysitting the agents, or do they surface a clean inbox of decisions?

What you get

The report stands on its own. If it says you’re not ready to auto-merge agent work today — most orgs aren’t, and that’s normal — it’ll tell you straight, and name the two cheapest axes to move first.

Agent-Readiness Scorecard

The six axes scored 0–4, each with the specific evidence from your repos and the single biggest gap called out.

Stack + workflow audit

A short written map of where agents and assistants are already used, where runs would silently fail today, and where humans are the real bottleneck.

90-day roadmap

Sequenced and costed, with a named first pilot (the highest-value, lowest-blast-radius place to start) and pre-registered success metrics for it — so “did it work?” becomes a signed number instead of an argument.

Executive walkthrough

A 60–90 minute recorded session with you and your leads, plus the written report (~8–12 pages).

Who it’s for

Engineering leaders (VP Eng, CTO, or an owner) at a 20–100 engineer B2B SaaS or dev shop who’ve already rolled out Copilot or Cursor, have at least one agent effort that stalled or that they badly want, and need to know where they actually stand before spending more. If you haven’t adopted assistants at all yet, it’s too early — I’ll tell you that for free.

Why me

I’m a staff engineer at a big tech company, and on my own time I run a production fleet of autonomous coding agents on infrastructure I built. The numbers, as of 2026-07-25, live-verified from my own instrumented systems:

Agent runs
483
Completed without me intervening
340 (70%)
Orchestrator releases
103
Agent-drivable CLIs shipped
8

The fleet has logged 483 agent runs across 12 projects; 340 of them (~70%) completed end-to-end without me intervening, and I’ve cut 103 releases of the orchestration system that runs it. Three of those 12 projects run fully autonomous with auto-merge enabled — 156 runs there, 101 of them completed.

That ~70% is my fleet on my projects — a completion rate, not a promise about your codebase, and I’ll never imply a client hit it. What it proves is that the operating model exists and I run it every day: verifier gates before merge, per-run isolation, an explicit trust ladder, structured completion contracts, evals with pre-registered criteria. Most people in this space sell the strategy. I can show you the run table.

I bring 20 years of infrastructure discipline behind it — USMC data systems, a decade leading network and security consulting delivery, now staff engineering — and I’ve authored the open standard for agent-drivable CLIs (aclig.dev). The audit turns that operating model into a diagnostic for yours.

Every number here is dated and sourced from instrumented systems. The audit is a complete, valuable deliverable on its own; if the roadmap points to a build worth doing, I offer a fixed-price implementation sprint separately — but there’s no obligation, and the report stands alone.

Two weeks. $8,500 fixed. A straight answer.

Email me a paragraph about where agents are stuck in your org. I’ll reply within one business day with whether — and how — the audit fits.