A shortlist arrives from procurement with four products on it, all described by the vendor as AI governance tools. One turns out to be a risk register with an AI module. One monitors model outputs for drift and bias. One automates evidence collection for certification audits. The fourth controls which documents an AI system is allowed to answer from. They are priced similarly, they use overlapping language in their marketing, and the evaluation committee has been asked to pick one.
That framing is the problem. The four products are not competing solutions to a single requirement; they govern different parts of the AI estate, and buying any one of them leaves most of the others' territory uncovered. It happens reliably because "AI governance" describes controls spanning several layers, and vendors from very different categories have all reached for the same phrase. So the useful question is rarely which AI governance platform is best. It is which layers your organisation needs to govern, and what is genuinely covering each one today.
This post maps the four categories against the layers they govern, identifies the gap most shortlists leave open, and sets out how to run the evaluation. It sits beneath our AI governance framework for regulated enterprises, which sets out the three-layer model the comparison below relies on; if the layers are unfamiliar, start there and come back.
Why "AI governance tools" describes four product categories
The categories below are drawn by what each type of product is built to control, not by vendor. Most real products sit mainly in one and reach a short distance into another.
Policy and risk-register platforms
These govern the organisation's documented positions. They maintain an inventory of AI systems in use, classify each by risk, hold the policies that apply, and produce the register a board or regulator asks to see. Their output is a defensible record of what the organisation has decided and who owns each decision.
What they do not do is reach into the systems themselves. A risk register records that a tool is in use and how it was classified; it does not constrain what that tool does next. That is the difference between governing intent and governing systems, and a complete register with no enforcement behind it reads well until someone asks what the AI actually answered from.
Model evaluation and monitoring tooling
These govern model behaviour. They run evaluations before deployment and watch outputs afterwards, tracking drift, bias, refusal rates, and unexpected changes when a model version updates. For organisations building or fine-tuning their own models, this category is essential and has no substitute.
Its scope, though, is the model rather than the corpus. Monitoring tells you the model behaved consistently; it does not tell you whether the document it drew on was current, approved, or appropriate for the person asking. A model can perform flawlessly against every evaluation while confidently answering from a superseded policy.
Compliance-automation platforms
These govern evidence for certification. They map controls to a framework such as ISO 27001, ISO 42001, or the staged obligations of the EU AI Act, monitor whether each control is in place, and assemble the evidence pack an auditor works through. For firms heading into certification, they compress months of manual collation.
The category's limit is that it evidences controls that exist elsewhere. A compliance-automation platform can demonstrate that an approval process is running; it cannot be the approval process. If the underlying control is procedural, the platform will faithfully evidence a procedure that people are quietly working around.
Governed knowledge layers
These govern what the AI is permitted to answer from. A document becomes eligible through a deliberate act of approval rather than by accident of file permissions, supersession is handled as a first-class event, and every answer resolves to the specific documents and versions behind it. The category's organising question is the one the other three cannot reach: on a given date, which sources was the AI allowed to use, who approved each of them, and what did any particular answer draw on.
Mapping the categories to the three governance layers
The table below sets each category against the three layers in our AI governance framework: the model the AI runs on, the knowledge it is allowed to draw on, and the audit record it leaves behind. Each cell describes what the category is built to do at that layer, not how well it does it. "Out of scope" is therefore not a criticism. It means the product was never designed to cover that layer, so something else will have to.
| Tool category | Built to govern | Model layer | Knowledge layer | Audit layer |
|---|---|---|---|---|
| Policy and risk-register platforms | Documented decisions and ownership | Records the decision | Out of scope | Policy-level attestations only |
| Model evaluation and monitoring | Model behaviour in production | Core purpose | Out of scope | Model-level metrics, not per-answer provenance |
| Compliance-automation platforms | Evidence for certification | Evidences stated controls | Out of scope | Control evidence, not answer reconstruction |
| Governed knowledge layers | What the AI may answer from | Tier and jurisdiction decisions | Core purpose | Per-answer provenance to document and version |
Two things are worth noticing. The first is how much of the field clusters around the model and the policy paperwork surrounding it. The second is what happens when you read down the Knowledge layer column: three of the four categories leave it out of scope by design, and only one is built to govern what the AI is allowed to answer from.
The gap most shortlists leave open
The knowledge layer is the one an evaluation is most likely to leave uncovered, and it is frequently the layer that determines whether governance is real. Three of the four categories treat the corpus as somebody else's concern, which is defensible in each individual case and produces a strange result in aggregate: an organisation can hold a complete AI inventory, continuous model monitoring, and a clean certification evidence pack, while the AI itself answers from whatever each user happens to have access to.
That default is worth stating plainly, because it is easy to mistake for a control. Permission inheritance is not governance of the knowledge layer. When a tool answers from everything a user can technically open, the most consequential governance decision, which documents are eligible to inform an answer, has been delegated to file permissions that were set years ago by people solving an unrelated problem. Nobody chose it. There is no approver, no date, and no list to produce when someone asks.
The practical test is a question a regulated firm should be able to answer in minutes: for an answer given three months ago that informed a decision you now have to defend, which documents produced it, which version of each, who had approved them, and who received the answer. The first three categories cannot answer it, for the reasons the table sets out: they never saw the query, or they watched the model rather than the corpus, or they evidence controls rather than operate them. If nothing on the shortlist can answer it, the shortlist has a hole in it regardless of how many products are on the list. Our guide to enterprise AI governance works through why this layer is usually the right place to start an implementation.
How to run the evaluation
The question set for interrogating a specific vendor's governance posture, organised by layer, is in the AI governance framework. What follows is the process around it: how to scope the exercise so those questions land against the right products.
Scope by layer before you compare products. Decide which layers you need to govern, and to what standard, before any vendor conversation. Without that, the evaluation is shaped by whichever category happened to demo first, and categories are not comparable to each other.
Establish what each layer is covered by today. For most organisations at least one layer is already addressed, often partially and by something nobody thinks of as an AI governance tool. Naming the current coverage prevents buying a second instance of it, which is the most common form of governance overspend.
Separate structural controls from procedural ones. Ask, for every control a vendor describes, whether the governed outcome is the default or whether it depends on somebody performing a step correctly each time. Procedural controls fail quietly under load, which is exactly when they are being relied on.
Ask for evidence, not roadmap. A platform built around governance answers the hard questions crisply and in writing; one that added governance later answers some of them with a delivery date.
Test against a real query. Take one question your organisation would have to defend, and run the whole reconstruction against it during the evaluation rather than after. This converts a set of claims into a demonstration, and it usually settles the shortlist faster than another round of vendor documentation.
How AnswerVault covers the knowledge and audit layers
AnswerVault is a governed AI knowledge layer that connects an organisation's existing document sources, including SharePoint, Google Drive, and Confluence, and delivers source-backed answers through web chat, Microsoft Teams, Slack, CLI, and API. It sits in the fourth category above, which for an evaluation means something specific: it is not a replacement for a risk register, a model-monitoring stack, or a compliance-automation platform. It covers the layer those three are designed to leave alone.
At the knowledge layer, a document becomes eligible for AI answers because a named subject-matter expert approves it, with the approval written into the audit trail at the moment it happens. When a document is superseded, the new version takes over and the record of which version was canonical on which date is preserved. That is what makes the date-stamped source list a matter of record rather than reconstruction. The discipline underneath it is curation, which we cover in our curated knowledge guide.
At the audit layer, citations resolve to specific documents and versions, so the reconstruction test above is a lookup rather than an investigation. At the model layer, the platform is tiered: the Enterprise sovereign tier is UK-controlled, while standard tiers run on managed AI infrastructure with EU and UK data residency. AnswerVault is ISO 27001 aligned and ISO 42001 underway, AI is included in every plan with no separate model or API-key requirement, and customer data is never used to train models. The detail procurement teams need, including subprocessors and attestations, is on our security and compliance page.
Where to start
If you are holding a shortlist of AI governance tools, the most useful first move is to map each product on it against the three layers in our AI governance framework and mark which layer each one actually governs. Where two products land in the same column, the shortlist has a duplicate. Where a column is empty, it has a gap, and the knowledge layer is the column to check first.
For the commercial view, our pricing page sets out the tiers and what governance features sit in each, and our comparison pages work through how AnswerVault sits alongside the general-purpose AI tools most organisations already run.
AnswerVault is built by Catapult CX, an enterprise technology consultancy. The product was originally developed for a global pharmaceutical company with strict data governance requirements; the same architecture now powers the SaaS platform.