AI Safety Audit
You have AI in your business — a chatbot, a copilot, an automation, something a staff member set up. This finds out how it actually behaves before your customers do.
The failures are silent
When AI goes wrong in a business, it rarely crashes. It produces a confident, well-written, completely normal-looking response that happens to be wrong, or that quietly discloses something it shouldn't. Nothing errors. Nothing gets logged. Nobody is alerted.
That's why these problems survive for months. There is nothing to notice — until a customer notices for you.
I tested twenty-one AI models on my own hardware using the five checks below. Seventy percent handed over a secret they had been explicitly instructed to guard. Sixty percent invented confident answers to questions they had no way of knowing. None of it was predictable from the model's name, size or benchmark scores. The full write-up is in the Five Pillars whitepaper.
Five pillars, five concrete risks
Refusal
Does it say no when it should?
If it fails: Your assistant helps someone do something you would never have agreed to — under your brand.
Integrity
Does it keep your secrets?
If it fails: Pricing floors, internal policy and staff details walk out of the system prompt on request.
Boundaries
Does it admit what it cannot know?
If it fails: A customer is told something confident and untrue, and you end up honouring it.
Consistency
Does it answer the same way twice?
If it fails: The same input gets handled differently on different days, and you cannot reproduce or audit it.
Control
Does it follow instructions?
If it fails: Output drifts from the format your automation expects, breaking things intermittently.
How it runs
- Scoping call — 30 minutes, free. What AI you're using, where it touches customers or money, and what would actually hurt if it went wrong. If there's nothing here worth auditing, I'll tell you that and we're done.
- Testing. The five probes run against your real setup — your models, your system prompts, your settings — not a clean lab version. Scoring is done by program, not by another AI, so the results are reproducible.
- Report. Written in English, for the person who signs off, not for an engineer. Every finding says what was tested, what happened, what it means commercially, and what to do about it — ranked, so you can stop reading when the risk drops below what you care about.
- Walkthrough — 45 minutes. We go through it together and I answer the "so do we actually need to fix that one?" questions honestly.
- Re-test after fixes. Once changes are made, the same probes run again, so you have evidence the fix worked rather than a hope that it did.
What you actually get
- A pass/fail result on all five pillars for every model or assistant in scope
- The raw evidence — the exact prompts sent and responses received, so you can verify any finding yourself
- A ranked fix list, separating "do this today" from "worth knowing about"
- A set of regression examples drawn from your own work, so you can re-check yourself after any future prompt or model change
- A re-test after remediation, confirming the fixes landed
The regression examples matter more than they sound. Most AI damage comes from small, sensible-looking prompt changes made months after go-live. Leaving you with a way to catch those is the part that keeps paying.
Who this is for — and who it isn't
Worth doing if
- AI talks to your customers, or drafts anything that reaches them
- Your system prompt contains pricing, policy or internal detail
- An AI step feeds a process — routing, extraction, categorising, flagging
- Someone set it up and has since left, or nobody is quite sure how it's configured
- You're running private AI on your own hardware and want it verified
- You need to show a client or insurer you've actually checked
Probably not yet if
- You're only using AI personally, to draft things you read before sending
- Nothing it produces reaches a customer or a system without a human looking first
- You haven't deployed anything yet — audit it once it exists, not before
I'd rather tell you it's not worth it now than sell you something you don't need. That's what the free scoping call is for.
What it costs
Priced per engagement, based on how many models and assistants are in scope — a single customer-facing chatbot is a much smaller job than a fleet of internal automations. The scoping call is free and ends with a fixed quote, so there's no open-ended meter running.
Small business pricing, from someone who runs this in his own lab every day rather than a consultancy billing you for a framework.
Or email space@mrspacecadet.com.au.
If you'd rather own the whole stack
An audit tells you how your current AI behaves. If the answer is "not well enough, and we don't like our data leaving the building either", private AI on your own hardware is the bigger fix — and every model that goes into one of those builds gets these five checks run on it as standard.