Owlhoot presents — AuditGrade

AI agents lie about finishing their work.

AuditGrade is the fail-closed completion gate for AI agents. Work isn't "done" until the evidence proves it — then it's sealed into a certificate anyone can verify.

Evidence is authoritative. Prose is not.

The problem

"Tests pass. Shipped." — except they didn't, and it wasn't.

Your agent says it ran the tests, fixed the bug, finished the task. Most of the time you can't tell if that's true without redoing the work yourself.

So you do. Every time an agent says "done" and isn't:

  • a broken feature ships as finished,
  • a security hole "passes,"
  • your most expensive engineers burn hours re-checking the machine's homework,
  • and when a customer, an auditor, or a regulator asks "prove the AI did what it claims" — you can't.

What does it cost you — every week, over the next quarter — every time "done" isn't?

The solution

Make "done" mean done.

AuditGrade gates completion on observable proof, not the agent's say-so.

An agent can't mark work complete unless the required evidence actually exists, is machine-readable, the hashes match, an independent validator accepts it, and the final answer points back to that evidence. Miss any of it, and completion is refused — fail-closed, by default.

When it all proves out, you get a Completion Certificate: portable, tamper-evident proof that the work was actually finished — verifiable by anyone, and designed to double as the event record the EU AI Act will require.

Orchestrators run your agent. Observability measures it. Audit logs record what it did. AuditGrade proves it actually finished.

The proof

We didn't just claim it's hard to fool. We tried — 18 times.

An audit tool that overclaims is worthless, so we red-team our own. Over one week, an independent adversarial AI attacked the harness 18 times in a row — trying every way to make an agent fake completion:

  • writing a file that just said {"status": "pass"} and hoping the system believed it,
  • forging a tool's output and claiming it came from the real tool by typing the tool's name as a string,
  • "cryptographically signing" its own proof — with a key it generated inside its own process,
  • pointing the trust anchor at a key it controlled.

Every attempt was caught and closed, each with a runnable proof. That's the standard the certificate is built to.

You can't build a wall your own process can climb. So the proof lives outside the agent's reach.

How it works

Three steps. Fail-closed by default.

  1. The agent does the work — in whatever runtime you already use.
  2. An independent observer verifies the evidence — it runs the checks itself; it doesn't take the agent's word.
  3. A certificate is issued only if everything proves out — otherwise completion is refused, and you know exactly what's missing.

Who it's for

If you ship AI work you have to stand behind.

  • Teams running AI agents in production who are tired of re-verifying them by hand.
  • AI-platform and agent builders who need a completion gate, not another dashboard.
  • Anyone who will have to prove to a customer, an auditor, or a regulator that AI work was actually done — correctly.

Founder note

Why I'm building this.

I'm Kirk. I spent 12+ years in corporate finance and finance automation — my job was verifying whether things were actually done, with evidence, in a way that survived an audit.

AI agents have the exact problem finance had before auditing existed: everyone takes the claim on faith. I'm building the audit layer for AI work — and I'm building it in public, proof first.

The waitlist

Be first to verify a Completion Certificate.

Join the waitlist for early access to the open verifier and the build updates. No spam — just the work, and the proof.

Early and built in public. Dev-grade today; the production issuer is on the roadmap. Owlhoot LLC.

You are on the list.

Watch your inbox for build updates and early verifier access.