What an AI audit must answer, and what it must not
The question “can AI do this” is almost never the right one. The answer is yes too often to be useful, and an audit that stops there produces a list of possibilities none of which can be decided on.
The two useful questions are elsewhere. Which of the ten ideas deserves the first three months, and what would stop that one holding up in production. An audit that does not rank, and that does not actively hunt for what will break, does not save you time.
The six questions, in order
They come in this order because the first two eliminate the most projects, and the earliest.
1. Does the data exist? Not “could it exist”, does it exist today, somewhere, in a readable form. If the information that would have to be learned was never recorded, no model will reconstruct it. It is the most frequent cause of failure and the easiest to check.
2. Does the software accept being read, and being written to? Two distinct limits. Plenty of software reads correctly and writes badly. That is not blocking, but it changes the first batch: you start with reading and answering, writing comes later. What decides is not the age of the software, it is what it exposes.
3. Is the problem tractable? A tractable problem has a right answer that a competent human can identify in reasonable time. If it takes three experts and two meetings to agree on the right answer, there is no ground truth to learn, and there will be no evaluation set.
4. What does an error cost? That is what sets the threshold, and the threshold decides everything else. An error that is visible and correctable allows automation. An error that reaches a customer, affects a person or creates legal exposure stays under approval, whatever the confidence score.
5. Do the volumes justify automation? A hundred documents a month does not pay for a project, whatever the enthusiasm. Ten thousand does. It is the simplest arithmetic in the exercise and the most often skipped.
6. Does it scale? What works on fifty cases may not hold on fifty thousand: cost per call, latency, the software’s own rate limits, context volume. This question does not come after the pilot, it comes before, because it changes the architecture.
Rank, do not merely list
Once the six questions have been put to each idea, you have to decide. We use ICE scoring, three scores from 1 to 10 whose average is taken.
| Dimension | Question asked |
|---|---|
| Impact | what business value if the case succeeds? |
| Confidence | what probability of technical success and adoption? |
| Ease | what implementation effort? |
Its simplicity is an asset: the ranking becomes readable by a board, with no technical skill needed to interpret the result. A high-impact case that is hard to ship loses points, and that is intended. For the heaviest projects, the Ease score breaks down over the first five questions above.
Doing it yourself, in a few minutes
We have put the first stage of this exercise in open access. Softscan reads what your software says about itself publicly and returns a diagnosis: where it stands against AI, what is exposed, and where to start. It is a first read, not a scoping exercise, and it is free.
What it cannot do has to be said: it does not see your data, your interfaces or your volumes. Questions 1, 2 and 5 are settled with someone who knows the software from the inside. That is what the scoping hour is for.
When an audit is worthless
When the decision is already made. If the project is going ahead regardless, the audit becomes a justification document. Better to invest that time in the evaluation set.
When the scope is too wide. “Audit our AI” means nothing. An audit covers one activity: invoice processing, support answering, application screening. Three activities and three short audits beat one forty-page report.
When nobody owns it. An audit with no name against each recommendation produces nothing. It is the part no provider can do for you, because it depends on who decides what at your company.
The bias to watch, including with us
An audit run by whoever will sell the project carries a structural bias, and we are no more exempt than anyone else. Two ways to neutralise it, and both are verifiable.
Require the criteria before the conclusions. If the criteria are written up front and verifiable, the ranking can be contested, therefore it is useful. If the conclusions come first, it is a sales proposal in disguise.
Check that a no is possible. Ask which case the provider has already concluded should not be done. A precise answer is a good sign. No answer is one too.