Six weeks after connecting an AI tool to SharePoint and Google Drive, the head of knowledge management at a professional services firm takes a question from the general counsel. Someone in operations acted on an answer the tool gave about a retention obligation, and the answer was wrong. It was not invented. The tool had found a real internal document and quoted it accurately. The document had been superseded twice and was three years out of date.
The instinct is to treat this as an accuracy problem and start evaluating better tools. It is not an accuracy problem. Retrieval worked, the citation was genuine, and the quotation was faithful. What went wrong happened earlier, at the point where nobody decided which documents the tool was allowed to answer from. That decision is what curated knowledge means, and it is the step most AI deployments skip. This post defines the term, explains why retrieval-based AI depends on it, and sets out how to tell whether your own document set qualifies. It sits within our guide to curated knowledge for regulated organisations, which covers the discipline in full.
What is curated knowledge?
Curated knowledge is a document set that someone has deliberately selected for a defined purpose, kept current as versions change, and approved as suitable to answer from. Three properties, all of them decisions made by a person rather than states a system arrives at on its own.
That definition is deliberately narrow, because the looser readings are what cause the problem in the opening scenario. A folder is not curated because it is tidy. A repository is not curated because it has an owner. The test is whether the three properties below hold at the level of individual documents.
Selection
Somebody decided this document belongs in the set, and the decision was about this document rather than about the folder it happens to sit in. Selection inherited from a directory structure is not selection, because the structure was built to solve a filing problem, not to answer the question of what should inform decisions.
Currency
The set reflects what is true now, and when a document is superseded the replacement takes its place rather than joining it. Most repositories accumulate. Accumulation is the opposite of currency: it means the oldest version of a procedure and the newest are equally present and equally findable.
Approval
A named person judged the document suitable for the purpose and that judgement is recorded with a date. The name matters more than the mechanism. Approval attributed to a service account or a sync job is not approval, because there is nobody to ask about it later.
Why AI needs curated knowledge
Curation is an old discipline, and organisations managed without it for a long time by relying on the people who knew which document to trust. Retrieval-based AI removes that filter, which is why the question has become urgent rather than merely good practice.
Retrieval cannot judge what it retrieves
A retrieval system matches a query against text and returns what scores highest. Nothing in that operation assesses whether the matched text is still true, still authorised, or still the version in force. Relevance and correctness are different properties, and only the first is measurable at query time. A superseded procedure that closely matches the question will outrank a current one that matches it less exactly.
Every answer arrives with the same confidence
When a person finds an old document, the context usually gives them a clue: an unfamiliar template, a departed colleague's name, a date. An answer strips all of that away. The response reads identically whether it drew on the current approved policy or on a draft somebody abandoned in 2023. Uniform presentation of non-uniform sources is the specific mechanism by which retrieval quality problems become business decisions.
The corpus is the only control surface that scales
There are three places to intervene: the model, the query, and the corpus. Model choice does not help, because the failure is not in generation. Query-level guardrails do not help either, because they cannot know that a particular document is stale. That leaves the corpus, which is the one place where a decision made once applies to every future answer. Deciding what the AI may read is not preparation for governance. It is the governance.
What curated knowledge is not
Most of the confusion around the term comes from adjacent concepts that solve different problems.
| Often mistaken for | What it actually controls | Why it is not curation |
|---|---|---|
| A permissions model | Who may open a document | Records access rights set for unrelated reasons, not whether a document should inform an answer |
| A wiki or knowledge base product | Where content is stored and edited | Describes location, not eligibility; wikis accumulate stale pages as readily as file shares |
| Data ingestion or a sync job | Which files reach the index | A transport mechanism with no opinion on what deserves to travel |
| Prompt or model tuning | How an answer is phrased | Adjusts expression of the source material without changing which sources qualify |
Each of these is worth having. None of them answers the question an auditor, a regulator, or a general counsel eventually asks, which is who decided this document could be used and when.
Standards bodies have reached the same position from a different direction. ISO 30401:2018, the knowledge management systems standard, requires an organisation to define the scope of the knowledge it manages and to assign ownership of it, both of which are curation decisions rather than storage decisions. The EU AI Act similarly expects documented control rather than asserted good practice. We work through the mapping in our guide to ISO 30401 knowledge management.
How to tell whether your knowledge is curated
Three questions settle it, and they take about an hour to attempt on a real example rather than in principle.
Can you produce the list? For a given date, name the documents your AI tool was permitted to answer from. If the answer involves querying a synchronisation log or describing a folder hierarchy, there is no defined set, only a reachable one.
Can you name the approver? Pick one document in that set and identify the person who judged it suitable and the date they did so. If the honest answer is that it arrived through a connector, the approval property is absent.
Can you find the superseded version? Take a procedure revised in the past year and ask the tool a question the old version answers differently. If the old answer comes back, currency is not being enforced anywhere in the pipeline.
Most organisations pass one of the three. The exercise is more informative than a vendor demonstration, because it tests your position rather than a product's capabilities.
How AnswerVault applies curated knowledge
AnswerVault is a governed AI knowledge layer that connects an organisation's existing document sources, including SharePoint, Google Drive, and Confluence, and delivers source-backed answers through web chat, Microsoft Teams, Slack, CLI, and API. The three properties above are not settings within it. They are the shape of how it operates.
A document becomes eligible because a named subject matter expert approves it, and that approval is written into the audit trail as it happens, which satisfies selection and approval in a single act. When a version is superseded, the replacement takes over and the record of which version was canonical on which date is preserved, so the failure in the opening scenario has no route in. Answers cite the specific document and version they drew on, which means the eligibility question can be answered per answer rather than per system.
The reason the product is shaped this way is that it was originally built for a global pharmaceutical company with strict data governance requirements, where an uncurated corpus was never an option. AnswerVault is ISO 27001 aligned and ISO 42001 underway, AI is included in every plan with no separate model or API key requirement, and customer data is never used to train models. The procurement-grade detail, including subprocessors, attestations, and the certifying entity, sits on our security and compliance page.
Where to start
Run the three questions above against one real decision from the past quarter. Whichever of the three fails first tells you which property is missing, and the missing property is usually the same one across the whole estate rather than specific to that example.
If the answer is that none of the three hold, the shortest route to all three is a document set where eligibility is decided rather than inherited. Our guide to curated knowledge explains how approval and currency work day to day, and curated knowledge for compliance covers what changes when an external party will eventually ask.
AnswerVault is built by Catapult CX, an enterprise technology consultancy. The product was originally developed for a global pharmaceutical company with strict data governance requirements; the same architecture now powers the SaaS platform.