
Your AI found the company policy.
It summarized it perfectly.
It cited the source.
It gave the user a clear, confident answer.
There is just one problem.
Nobody ever approved that policy.
It was a draft.
Someone proposed it.
People reviewed it.
Comments were added.
Maybe Legal looked at it.
Maybe Finance discussed it.
Then the organization decided not to move forward with it.
The draft stayed where it was.
And months later, your AI found it.
Welcome back to Production.
The AI Didn’t Hallucinate
This is what makes this scenario interesting.
The AI didn’t invent anything.
It didn’t fabricate a policy.
It didn’t misunderstand the document.
It didn’t bypass security.
The user had permission to access the document.
The retrieval system found relevant information.
The model accurately summarized what the document said.
It may even have provided a citation so the user could see exactly where the answer came from.
Technically, quite a lot went right.
And the answer was still wrong.
The AI didn’t hallucinate. It cited its source.
The problem wasn’t the model.
The problem was authority.
Retrieval Is Not Authority
Imagine a document called Customer_Risk_Policy_DRAFT.docx
Inside it is a proposed rule:
Effective October 1, high-risk customers require executive approval before new orders can be accepted.
Six months later, someone asks the AI: “What approvals are required for a high-risk customer?”
The retrieval system searches the company’s documents.
It finds the draft.
The document contains almost exactly the language needed to answer the question.
From a search perspective, this is an excellent result.
The AI responds: “High-risk customers require executive approval before new orders can be accepted.”
Clear.
Concise.
Supported by company information.
But still wrong.
Finding relevant information answers: Can I find something about this?
It does not answer: Should I trust what I found?
Those are very different questions.
Your Company Is Full of Things That Used to Be True
This problem isn’t unusual.
Mature organizations accumulate history.
- Old policies.
- Superseded procedures.
- Draft documents.
- Previous versions of documentation.
- Archived spreadsheets.
- Meeting notes.
- Old presentations.
- Proofs of concept.
- Temporary calculations.
- Documents describing projects that were canceled.
- Documents describing projects that were completed completely differently from the original design.
All of those things can contain legitimate company information.
Some of them were authoritative once.
Some were proposals.
Some were never correct.
Some were correct only for a small period of time.
Humans often know the difference.
AI may not.
Humans Carry Context That Systems Don’t
Someone who’s been with the company for years sees an old document and says:
“Oh, we don’t do it that way anymore.”
Or: “That was just a proposal.”
Or: “Finance changed that calculation last year.”
Or: “Don’t use that spreadsheet. Use the new one.”
Or my personal favorite: “Technically that’s still out there, but nobody should be using it.”
That knowledge may never have been captured anywhere.
It’s tribal knowledge.
Humans use it constantly.
We know that a file can exist without being current.
We know the newest file isn’t necessarily approved.
We know something labeled FINAL may be anything but final.
AI doesn’t automatically inherit any of that knowledge.
Unless we capture it somewhere, AI simply sees information.
The Sherpa’s Notebook
And if you’re thinking this is only a document-management problem…
It isn’t.
Your data platform is full of drafts too.
We just don’t call them drafts.
I have a client with older Power BI reports the business intentionally wants to keep.
There’s absolutely nothing wrong with that.
The reports represent historical information.
People may occasionally need them.
The business simply doesn’t want those reports updated anymore.
Perfectly reasonable.
A human familiar with the environment can look at one and understand: “That’s the old report.“
They know not to use it for today’s numbers.
They may even know why it was replaced.
Now give AI access to the reporting environment.
Someone asks: “How do we calculate customer profitability?”
The AI may find:
- The current Power BI semantic model.
- Current documentation.
- The current report.
- And an older Power BI report that hasn’t been updated in two years.
Maybe that older report calculates profitability differently.
Now what?
Both calculations are real.
Both were used by the company.
Both may have been completely correct during the periods in which they were used.
Only one represents how the organization calculates profitability today.
The AI doesn’t need more data.
It needs to know which data represents current business truth.
Your Data Platform Is Full of Drafts Too
Once you start looking at the problem this way, you see it everywhere.
- An old Power BI report.
- A deprecated semantic model.
- A legacy stored procedure.
- An abandoned dataset.
- A previous version of a calculation.
- A table nobody is supposed to use anymore.
- A pipeline retained for historical reasons.
- A spreadsheet containing last year’s methodology.
- A database kept online for historical lookup.
- A metric definition that changed after an acquisition.
- A view someone replaced but never removed.
- None of those are literally drafts.
But they have the same architectural characteristic as our unapproved policy.
They exist without necessarily representing current authority.
And once AI can search across everything, everything becomes a potential source.
Production Is Full of Fossils
That’s not necessarily bad.
History matters.
Sometimes history is exactly what we want.
Maybe someone asks: “What was our customer risk policy in June 2024?”
Now the old policy may be the correct source.
Or: “How did we calculate customer profitability before the 2025 methodology change?”
That stale Power BI report suddenly becomes extremely valuable.
The problem isn’t that old information exists.
The problem is when the system cannot distinguish:
Current / Historical / Draft / Superseded / Never Approved / Please, for the love of everything holy, don’t use this anymore.
That last one may not be an official governance classification.
Perhaps it should be.
Search Relevance Isn’t Business Truth
This is where traditional retrieval logic starts running into business reality.
Search systems are very good at finding relevant information.
If someone asks about customer profitability and an old report contains detailed information about customer profitability, that’s relevant.
If someone asks about high-risk customer approval and the draft policy contains that exact phrase, that’s extremely relevant.
Search did its job.
But relevance doesn’t establish authority.
A source can be highly relevant and completely inappropriate for answering the question.
So production AI needs signals beyond: How closely does this match what the user asked?
It may also need to know:
- Who owns this?
- Is it approved?
- Is it current?
- When did it become effective?
- When did it stop being effective?
- Has it been superseded?
- Is it historical?
- Which business domain does it apply to?
- Is it authoritative for this particular question?
Those aren’t model questions.
They’re governance questions.
And they need to become architectural signals.
The Newest Document Doesn’t Automatically Win
Let’s make this worse.
Suppose AI retrieves three documents:
Customer_Risk_Policy_2025.pdf
Customer_Risk_Policy_2026.pdf
Customer_Risk_Policy_DRAFT_2027.docx
Which one should it use?
The newest?
No.
The draft is newest.
The one modified most recently?
Also no.
Someone may have corrected a typo in the old policy yesterday.
The most semantically relevant?
Still no.
The draft may contain exactly the terminology from the user’s question.
The most frequently accessed?
Probably not.
Everyone may be opening the old policy because they’re trying to figure out why AI keeps giving them the wrong answer.
There is no clever ranking algorithm that eliminates the need for organizational authority.
At some point, the organization has to say:
This is the version that applies.
“But It Cited the Source”
Citations are important.
I want AI systems to show people where important information came from whenever possible.
That makes answers more transparent.
It gives users something to inspect.
It helps with troubleshooting.
But we should be careful about teaching users that Citation = Truth
It doesn’t.
A citation establishes provenance.
It tells us This is where the AI got the information.
That’s valuable.
But if the citation points to Customer_Risk_Policy_DRAFT_FINAL_v7.docx, we still have a problem.
The AI may have represented the source perfectly.
The source itself didn’t deserve authority.
Authority Is Metadata
This is where metadata comes back into the conversation.
Again.
For important information, we may need metadata that says:
- Status: Approved
- Owner: Finance
- Effective Date: January 1, 2026
- Expiration Date: December 31, 2026
- Authoritative: Yes
- Superseded By: Customer Risk Policy v4
- Historical Use: Allowed
Now the system has something useful.
It doesn’t need to guess what DRAFT means.
It doesn’t need to infer authority from the filename.
It doesn’t need to assume the most recently modified file is the best one.
The organization is explicitly describing the lifecycle of its information.
And retrieval can use that context.
Historical Questions Make This More Interesting
We also can’t simply tell AI to only use the newest thing.
Because sometimes the user wants historical truth.
Suppose Jeannine asks: “What was our approval policy for high-risk customers in March 2025?”
The current policy may be wrong for that question.
The AI needs the policy that was authoritative in March 2025.
The same applies to data.
- What was the customer profitability calculation in 2024?
- What did this KPI mean before the acquisition?
- Which sales territories existed last year?
- What was the customer’s status when this transaction occurred?
Authority can be temporal.
Something can be wrong today and still be the correct answer to a question about yesterday.
That’s why lifecycle metadata matters.
What About Semantic Models?
This is particularly important as organizations connect AI to analytics platforms.
Imagine you have the following:
- Sales Model
- Sales Model New
- Sales Model v2
- Sales Model – Testing
- Sales Model – OLD
- Sales Model – DO NOT USE
- Sales Model Final
Which one represents the business?
A human Power BI developer who has worked there for some time may know instantly.
The AI doesn’t get to inherit that person’s brain.
If the platform doesn’t identify the certified or authoritative model, we’re asking the AI to infer governance from naming conventions.
That’s not architecture.
That’s archaeology.
Don’t Delete History. Label It.
The answer isn’t deleting everything old.
Please don’t leave this article and begin deleting every archived report because AI might find it.
History has value.
Auditors need it.
Analysts need it.
The business needs it.
Sometimes AI needs it too.
The better answer is lifecycle.
- This report is current.
- That report is archived.
- This policy is approved.
- That policy is draft.
- This dataset is certified.
- That one is experimental.
- This definition applies today.
- That definition applied from 2022 through 2024.
- This semantic model is authoritative.
- That one is retained for historical reference.
Then make those distinctions available to the systems retrieving information.
We don’t need less information.
We need better information about the information.
The AI Needs to Know What Went Live
That’s really what our draft-policy example comes down to.
The organization had ideas.
The organization discussed them.
The organization documented them.
But only some of those ideas became reality.
The AI needs to know the difference.
The same is true of our data platforms.
- There are things we built.
- Things we tested.
- Things we replaced.
- Things we retired.
- Things we kept.
- Things we changed.
- Things we probably should have deleted twelve years ago but everyone is afraid to touch.
Existence doesn’t establish authority.
Access doesn’t establish authority.
Relevance doesn’t establish authority.
Even a citation doesn’t establish authority.
Authority has to come from somewhere.
The Sherpa’s Lesson
Production AI doesn’t just need access to enterprise knowledge.
It needs to understand which knowledge deserves authority.
Because your company contains thousands of things that are technically real.
Some represent what the company believes today.
Some represent what it believed yesterday.
Some are valuable historical records.
Some are unfinished ideas.
And some are FINAL_FINAL_USE_THIS_ONE_v3.
AI cannot treat them equally.
So when your AI proudly says: “I found the source.”
Don’t stop there.
Ask the following questions:
- Is it current?
- Is it approved?
- Is it authoritative?
- When was it authoritative?
- Has it been superseded?
- How does the AI know?
Because retrieval tells AI what exists.
Governance tells it what to trust.
Leave a Reply
You must be logged in to post a comment.