Active Working Memory: The RAM of Agentic Systems

Part 2 of the Building the AI Memory Stack series

When I published the first article in this series, I thought I was writing about context windows.

The more I wrote, the more I found myself bouncing between documents. I had the glossary open in one browser tab. The Memory as Infrastructure article was open in another. A third tab contained notes about Context Hydration. GitHub was open with the SDK specifications, and a handful of Architecture Decision Records were sitting beside my editor.

None of those documents individually contained “the answer.” Together, they formed the temporary collection of information I needed before I could make progress.

Then something unexpected happened.

I was no longer writing about context windows.

I was reconstructing a memory hierarchy, and the context window was only its first layer.

That is why this became a series.

In Part 1, I argued that the context window is best understood as the CPU cache of an AI system. It is an execution surface, not a memory system. Like CPU cache, it is optimized for fast access and exists only for the duration of the work being performed.

But caches do not populate themselves.

Neither do context windows.

The Working Set Before the Work

Before I could write, I assembled a working set.

Browser tabs and open documents

- Sovereign Systems glossary
- Memory as Infrastructure
- Context Hydration notes
- SDK specifications
- Architecture Decision Records
- GitHub repository
- Article draft

            |
            v

    Active Working Memory

            |
            v

      Context Window

            |
            v

         Reasoning

That collection was not my long-term memory. It was a temporary working set assembled for one task. My brain still had to compare ideas, notice contradictions, and produce something new. The documents simply gave the reasoning process the information it needed.

Agentic systems work in much the same way.

Before a model begins reasoning, documents have been retrieved, tools have executed, state has been restored, policies evaluated, and responses normalized. The prompt is usually the final artifact produced by an orchestration layer, not the beginning of one.

By the time the model receives its first token, dozens of retrieval, filtering, ranking, and assembly decisions may already have been made.

That assembled execution state is what I call Active Working Memory.

Active Working Memory Diagram

This article is about the second layer in that stack.

Models Reason. Applications Assemble.

One sentence captures the distinction this entire article is trying to make.

Models reason. Applications assemble.

Those are different responsibilities.

A model does not retrieve documents.

A model does not decide which Git commit matters.

A model does not know whether a tool response is stale.

An application does.

Consider a concrete case. An agent is asked to update a Python SDK.

It retrieves the ADR describing the package boundary.

It loads the glossary definition of Active Working Memory.

It checks the current implementation on GitHub.

Only then does it build the prompt.

People often talk about “putting something into the context window.”

That wording quietly suggests the context window is responsible for finding, selecting, and organizing information.

It isn’t.

By the time inference begins, the application has already decided what the model will, and will not, be allowed to see.

Working Memory Sequence Diagram

The context window does not begin the process. It receives the result of the process.

Many failures blamed on the model are actually failures of context assembly. The wrong evidence was retrieved. A stale document won. A constraint never made it into the working set. Those are architecture problems before they are model problems.

RAM for Agentic Systems

If the context window is the CPU cache, Active Working Memory is the system RAM.

Layer Primary Responsibility Typical Owner
Durable Memory Preserve knowledge Storage / Application
Active Working Memory Assemble task state Orchestrator
Context Window Present selected information Runtime
Model Perform inference LLM

Each layer has a distinct responsibility. No layer substitutes for another. A larger context window does not repair poor selection, and a better model cannot reason over evidence it never receives.

Search Finds. Context Assembly Decides.

Search answers one question.

What information exists?

Context assembly answers another.

Given this task, what information belongs together?

Those sound similar. Architecturally, they are completely different.

Active Working Memory presents that second decision to the model.

Two systems can use the same model, the same context window, and the same knowledge store yet produce very different results, because one assembles a concise, relevant working set while the other floods the model with loosely related information.

The model is identical.

The memory architecture is not.

An Architectural Boundary

Once Active Working Memory is treated as a real layer, context assembly stops looking like prompt engineering and starts looking like systems architecture.

Information crosses this boundary only after decisions have been made about relevance, authority, recency, format, and priority.

Every unnecessary document increases Context Tax. Every verbose tool response competes for attention. Every missing source creates a blind spot the model cannot recognize from inside the window.

The context window can only reason over what Active Working Memory hands it.

Looking Ahead

The working set I assembled while writing this article disappeared as soon as the article was finished.

The article remained.

That distinction turns out to matter.

Agentic systems face the same decision. What belongs only in today’s working set? What deserves to become tomorrow’s memory?

That is where Part 3 begins.

Models don’t assemble context. They inherit it.

Working memory is not where knowledge lives. It is where knowledge collaborates.

Facebooktwitterredditlinkedinmail

The Context Window Isn’t Memory. It’s the CPU Cache of AI.

Treating the context window as memory is one of the most expensive misconceptions in AI systems design. This piece reframes it as CPU cache and maps the full memory hierarchy that has to live beneath it.

One of the most common misconceptions in modern AI is that a larger context window somehow “solves” memory.

It doesn’t.

A context window increases how much information a model can consider during a single inference. It does not give the system a durable memory of what happened before or what should matter later.

There’s a cleaner way to think about this, and it uses a hierarchy every systems engineer already knows by heart.

Traditional Computer Agentic AI System
CPU Cache Context Window
RAM Active Working Memory
Filesystem Durable Memory
Git History Reasoning Ledger
Chain of Custody Write-Side Custody

Each layer exists for a different purpose, and collapsing them is where most “memory” confusion begins.

The Context Window Is CPU Cache

A CPU cache is extremely fast and intentionally temporary. Data flows through it constantly because the processor needs immediate access while work is being performed. Nothing is meant to live there.

A context window plays a remarkably similar role. It holds the information required for this reasoning step. Once inference completes, that working state effectively disappears unless another component deliberately preserves something from it.

That is why I prefer to treat the context window as an execution surface rather than a memory system. It is where thinking happens, not where knowledge lives.

Context is borrowed. Memory is curated. One exists only for the duration of reasoning. The other exists so reasoning does not have to begin again.

The Rest of the Stack

The cache analogy only works if the layers beneath it are real, so it is worth naming them.

Active Working Memory is the RAM of the system: the retrieved documents, tool results, and intermediate state assembled for the current task. It outlives a single cache line, but not the session.

Durable Memory is the filesystem: the decisions, evidence, and domain knowledge written down on purpose so they survive long after the prompt that produced them.

The Reasoning Ledger is the git history: not just what the system knows, but how it came to know it, including the revisions and corrections that accumulate over time. It is the opposite of a Digital Attic, the anti-pattern of dumping raw logs into storage and hoping search can reconstruct the reasoning later.

Write-Side Custody is the chain of custody: the guarantee that everything entering durable memory is attributable, verifiable, and hard to tamper with after the fact.

A context window touches all of these during inference. It replaces none of them. Confusing these layers is the architectural equivalent of expecting CPU cache to replace a filesystem. It works only until the process exits.

Bigger Caches Don’t Fix Poor Inputs

Modern models keep pushing context windows into the hundreds of thousands, and now millions, of tokens.

That is genuinely impressive. It also does nothing to eliminate Prose Tax.

Prose Tax is the cost of recovering intent from verbose, ambiguous, or poorly organized information. A larger window simply raises the budget you are allowed to spend. It says nothing about whether you are spending it well.

Past a certain point, the extra room actively works against you. As a window fills with weakly relevant material, signal density falls and the model’s recall degrades, a drag the specification names the Context Tax.

In practice, a carefully structured 20,000-token context often communicates intent better than an unstructured million-token dump. Capacity and communication are different optimization problems, and only one of them is solved by scale.

Memory Begins After Inference

This is where Memory as Infrastructure enters the picture.

Rather than assuming memory emerges on its own from larger prompts, the surrounding architecture decides, deliberately, what should survive.

Not every prompt deserves to become memory. Some do:

  • Decisions
  • Evidence
  • Corrections
  • Provenance
  • Domain knowledge
  • Reasoning history

These become durable assets that future reasoning can build on, instead of reconstructing them from scratch every time.

The return trip matters just as much. Hydration is the moment memory becomes voice. Information that has been compacted, verified, and preserved is expanded back into language so it can participate in reasoning once again. The knowledge never disappeared; only its representation changed. Context Hydration is where durable memory becomes working memory again, and it closes the loop the cache analogy opened.

The Architectural Shift

Most current discussions ask:

How do we fit more information into the context window?

I think the better question is:

What information deserves to survive beyond the context window?

Those are fundamentally different design problems. The first is a question about model capability. The second is a question about systems architecture, and it is the one that compounds over time.

Looking Forward

As context windows keep growing, I suspect competitive advantage will shift away from raw token capacity and toward memory architecture.

The systems that win won’t be the ones that can read the most. They will be the ones that know what to preserve, what to forget, and how to keep that memory trustworthy across months and years of operation.

A larger context window lets an AI think longer. Memory as Infrastructure lets a system learn longer.

The context window is today’s execution surface. Memory is tomorrow’s foundation.

Architectures that understand the difference will outlast those that simply buy larger windows.

Facebooktwitterredditlinkedinmail