arXiv:2604.11759 · May 22, 2026
Retrieval Is Not Enough: AI for Organizations Needs Epistemic Infrastructure
OIDA connects decisions, evidence, and open questions so AI can follow what an organization has decided and what has changed.
- Federico Bottino
- Carlo Ferrero
- Nicholas Dosio
- Pierfrancesco Beneventano
Kakashi Ventures Accelerator (KVA) · Massachusetts Institute of Technology
Abstract
Retrieval-based organizational AI usually ranks passages by semantic relevance. In organizations, this is too weak: a passage can be relevant yet misleading because it records a superseded plan, a tentative hypothesis, an unresolved question, or one side of a later contradiction. We introduce OIDA, an epistemic memory architecture that stores these distinctions as persistent state rather than leaving them to prompt-time inference.
Formally, OIDA is a typed, attributed, signed, time-indexed graph. Nodes represent decisions, evidence, plans, hypotheses, constraints, and open information needs; edges encode support, contradiction, dependency, and supersession; retrieval combines semantic similarity with graph state. We also release three organizational epistemic-stress corpora for testing whether systems recover current decisions, contradictions, superseded claims, and knowledge gaps. The pilot evidence is preliminary: full-context baselines win in-window aggregate quality, while structured retrieval uses far fewer input tokens and its clearest positive signal is explicit knowledge-gap surfacing. The central claim is not that graph memory universally beats long context, but that organizational AI needs memory layers that represent epistemic state, not only text relevance.
The problem: relevant, yet misleading
An AI agent is asked to summarize an organization's position on a strategic decision. It retrieves several relevant records: an early proposal, a later decision note, a recent market update that challenges the plan, and an unanswered regulatory question. All are topically relevant. Yet they do not play the same epistemic role — and a system that treats them as interchangeable can produce an answer that is fluent, textually grounded, and still wrong about the organization's actual state.
The system needs to know what each document records: a decision, a constraint, a plan, a hypothesis, evidence, an observation, a disagreement, or an open question. Finding a passage on the right topic does not establish whether its contents still guide the organization's work.
See the distinction
Compare two ways to choose records for the same question. One ranks them by similarity to the question. The other also uses labels and connections to track decisions, changes, and disagreements.
Agent query
“What is our position on NovaTech?”
Each node is a typed memory object — a decision, a plan, evidence, a hypothesis, or an open question — sized by its priority weight. Select a node to inspect it.
Choose a retrieval strategy. Semantic ranks by topical relevance; Epistemic (OIDA) also conditions on graph state — current vs. superseded, contradicted, and unresolved.
Epistemic graph for the query: What is our position on NovaTech?
Typed memory nodes:
- DECISION (priority 0.96): Co-invest $2M in NovaTech's seed round.
- PLAN (priority 0.22, superseded): Initial plan: lead a $5M seed round in NovaTech.
- CONSTRAINT (priority 0.90): Fund cannot hold more than 15% in a single seed deal.
- EVIDENCE (priority 0.78): Founding team: ex-Stripe engineers with strong technical pedigree.
- NARRATIVE (priority 0.70): Fund thesis: back technical founders building developer tooling.
- EVIDENCE (priority 0.74): TAM = $12B, per NovaTech's pitch deck.
- EVIDENCE (priority 0.76): TAM ≈ $3B, per the independent due-diligence dossier.
- HYPOTHESIS (priority 0.34): NovaTech can reach $10M ARR within 24 months.
- OBSERVATION (priority 0.38): A direct competitor raised a large round last month.
- QUESTION (priority 0.62): Unresolved: is NovaTech's core patent enforceable in the EU?
- EVALUATION (priority 0.54): Partner conviction score: 7/10.
Signed relations:
- Co-invest $2M SUPERSEDES Lead $5M seed (A replaces B; B is demoted, not deleted)
- ≤15% per deal BLOCKS Lead $5M seed (A actively prevents B)
- Founder pedigree SUPPORTS Co-invest $2M (A provides evidence strengthening B)
- Dev-tooling thesis SUPPORTS Co-invest $2M (A provides evidence strengthening B)
- TAM ~$3B (dossier) CONTRADICTS TAM $12B (deck) (A contradicts B)
- TAM $12B (deck) SUPPORTS $10M ARR / 24mo (A provides evidence strengthening B)
- Conviction 7/10 BASED_ON Founder pedigree (A is the logical grounding of B)
- EU patent enforceable? BLOCKS Co-invest $2M (A actively prevents B)
- Competitor raised SUPPORTS $10M ARR / 24mo (A provides evidence strengthening B)
Semantic retrieval: Ranks the passages most topically related to NovaTech — but the top hit is a superseded plan, and only the optimistic side of the market-size contradiction is surfaced. It misses: The current decision (D1) that superseded the $5M plan; The conflicting due-diligence TAM (E3) — only the deck's $12B appears; The open patent question (Q1).
Epistemic retrieval: Returns the current decision (not the superseded plan), flags both sides of the TAM contradiction, and surfaces the unresolved patent question as an open knowledge gap. It captures: The current decision D1 — the superseded plan P1 is demoted, not retrieved; Both sides of the TAM contradiction (E2 ↔ E3); The open knowledge gap Q1 (patent enforceability).
How OIDA connects records
OIDA represents memory at a given time as a graph. Each record has a label and a priority score. Connections between records show support, disagreement, and replacement. The system uses these alongside similarity to the question when choosing what to retrieve.
Label each record
The system can give a binding decision a different priority from an observation or an untested hypothesis.
Keep disagreements visible
The graph stores a connection between conflicting records. It keeps both, and can lower one record’s priority or retrieve them together.
Record unanswered questions
A question has its own record. Its priority rises while it remains unanswered.
Nine epistemic node types
Each label has a starting priority and a rule for how it changes over time. The authors chose these rules; the pilot did not establish that they are optimal.
- DECISION1.00
Binding choice or commitment
Stable until superseded
- CONSTRAINT0.90
Hard structural boundary
Stable until changed
- EVIDENCE0.80
Verifiable supporting or refuting material
Slow decay
- NARRATIVE0.70
Persistent interpretive context
Stable contextual anchor
- PLAN0.65
Structured intention with horizon
Time-bounded decay
- EVALUATION0.55
Informed qualitative assessment
Moderate decay
- OBSERVATION0.40
Weak uninterpreted signal
Unreinforced decay
- HYPOTHESIS0.30
Unverified testable claim
Decay if untested
- QUESTION0.30
Open information need
Urgency grows while unresolved
Signed edge vocabulary
Positive connections show support or dependency. Negative connections show disagreement or a blockage.
| Relation | Coeff. | Semantics (A → B) |
|---|---|---|
| SUPPORTS | +1.0 | A provides evidence strengthening B |
| BASED_ON | +0.8 | A is the logical grounding of B |
| IMPLEMENTS | +0.7 | A operationally realizes B |
| SUPERSEDES | +0.6 | A replaces B; B is demoted, not deleted |
| REFINES | +0.5 | A narrows B without contradiction |
| DERIVES_FROM | +0.5 | A follows logically from B |
| ENABLES | +0.4 | A is a necessary condition for B |
| PRECEDES | +0.3 | A temporally precedes B |
| BLOCKS | -0.4 | A actively prevents B |
| CONTRADICTS | -0.6 | A contradicts B |
Three test document collections
The authors created three synthetic document collections to test whether AI can recover current decisions, disagreements, replaced claims, and unanswered questions. They are available on the OIDA evaluation corpora repository.
ClearPath
46 files · 318 KBConsulting / operations
Process redesign, bottleneck evidence, unresolved audit workflow, and conflicting organizational observations
FireGlass
47 files · 566 KBIoT / product development
Architecture decisions, latency constraints, technical debt, negative knowledge, and integration risk
Vertex Minds
77 files · 376 KBVenture capital
Investment decisions, due-diligence gaps, prior-deal lessons, and market-size contradiction
What the pilot shows
The evaluation is a pilot, not a benchmark-scale validation. Quality is scored with the Epistemic Quality Score (EQS = 0.20·ECA + 0.25·CP + 0.20·CR + 0.20·EC + 0.15·DE). When the whole corpus fits in context, the full-context baseline wins aggregate quality — the gap is driven mostly by context coverage.
| Dimension | Structured | Full context |
|---|---|---|
| Type fidelity | 0.667 | 0.807 |
| Groundedness | 0.550 | 0.862 |
| Context coverage | 0.418 | 0.860 |
| Contradiction handling | 0.578 | 0.788 |
| Decision utility | 0.582 | 0.843 |
| EQS (total) | 0.557 | 0.833 |
| Dimension | Structured | Full context |
|---|---|---|
| Type fidelity | 0.617 | 0.817 |
| Groundedness | 0.550 | 0.865 |
| Context coverage | 0.455 | 0.863 |
| Contradiction handling | 0.548 | 0.822 |
| Decision utility | 0.520 | 0.848 |
| EQS (total) | 0.539 | 0.844 |
42.1×
fewer input tokens for structured retrieval on ClearPath (2,623 vs 110,539).
10/10
knowledge-gap surfacing for structured retrieval, vs 5/10 for full context in this pilot.
Pilot
When the whole collection fits in the model's input, full context scores higher for answer quality. The pilot does not establish how the graph performs in production.
What is not yet established
- No longitudinal validation. The time-dependent node-priority dynamics are a specification, not a measured result.
- No dynamic-vs-static isolation. A static typed-graph baseline is the decisive comparison; dynamic weighting remains an open hypothesis.
- In-window regime. The full-context baseline sees the entire corpus and therefore wins aggregate quality in the reported setting.
- The graph must be built accurately. Typed graph retrieval only helps if nodes and edges are extracted with adequate precision and recall.
Cite this work
@article{bottino2026retrieval,
title = {Retrieval Is Not Enough: AI for Organizations Needs Epistemic Infrastructure},
author = {Bottino, Federico and Ferrero, Carlo and Dosio, Nicholas and Beneventano, Pierfrancesco},
journal = {arXiv preprint arXiv:2604.11759},
year = {2026},
url = {https://arxiv.org/abs/2604.11759}
}Frequently asked questions
- What is OIDA?
- OIDA organizes the information an AI uses to answer questions about an organization. Each record has a label, such as decision, evidence, plan, or open question. Connections show which records support, contradict, depend on, or replace others. Dates and priority scores help the system choose what still applies. Researchers describe this as a typed, attributed, signed, time-indexed graph.
- What does “epistemic infrastructure” mean?
- It is the part of an AI system that records the status of its information. For example, it can distinguish an approved decision from an untested idea, show which plan a later decision replaced, and keep track of unanswered questions. Those labels and connections stay in memory and can be used in later answers.
- How is OIDA different from RAG?
- Retrieval-augmented generation (RAG) gives an AI relevant passages to use when answering. Standard RAG often chooses them by similarity to the question. OIDA also considers each record's label, priority, and connections. It can lower an old plan's priority, retrieve both sides of a disagreement, or return an unanswered question as a record.
- Does graph memory beat long-context prompting?
- Not in the reported pilot. When the model could read the whole document collection, that approach scored higher for overall answer quality. Structured retrieval used about 42 times fewer input tokens on ClearPath and surfaced explicit knowledge gaps more consistently. Tokens are the units of text a model processes. The pilot does not establish that a graph gives better answers in general.
- What are the OIDA evaluation corpora?
- They are three synthetic document collections designed to resemble organizational work: ClearPath covers consulting and operations, FireGlass covers IoT product development, and Vertex Minds covers venture capital. They test whether a system can recover current decisions, disagreements, replaced claims, and unanswered questions. The authors released them publicly so others can repeat the tests.
- When is epistemic graph memory worth using?
- Consider it when the relevant documents cannot fit affordably into the model's input and an answer depends on changing decisions, conflicting sources, open questions, differing authority, or connections across documents. If reading the full collection is affordable, use that. Ordinary RAG may be enough for a fact found in one source.