Letting a model manage its own context is very bitter lesson-pilled
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
this is done on the model side so it is invisible to the caller. all of deepseeks kv cache compression stuff is basically the same thing. it doesn't really put the entire context into kv cache and does various inference side stuff to decide what to actually use as context.
I would be concerned with context management consuming limited attention resources.
Do you want your agent solving its own memory crisis, or do you want it solving the actual task? It can probably do both at the same time, but I suspect there is a non trivial cost associated with this.
A separate hypervisor agent that manages the main agent's context would be much better in my experience. You can run it on a different schedule and the main agent has to spend zero tokens thinking about it. This also makes it a lot easier to control when caches will be missed.
An argument against having a separate hypervisor agent.
A separate hypervisor agent at least doubles the cost because (1) the underlying model needs to be the same so that we get the same degree of intelligence, and (2) it needs to have the same context + more tokens for the work it does.
I wonder how bad the performance would be if they plain ignored the whole rotary encoding dance and just back-filled precisely the parts of the cache that changed directly. Would it break the model? Confuse the model? Or would the model internally correct for it?
The issue is this: rotary encoding currently numbers the tokens on the request level.
If you pre-process a text file, you cannot insert it into the context because even if you numbered the tokens within a file properly, it is numbered wrong in the global context of the request. So you have cached data but the request structure does not allow you to insert it.
I think one way out of this would be to split a token into its request and content components. Then you might have to recalculate the request level information but never the content level information.
Edit: Now that I think about, why even cache the RoPE'd values to begin with? If you only need the RoPE'd data during token generation, then you could just RoPE on the fly instead of baking it into the KV cache, meaning you unlock block level suffix and infix caching, not just prefix caching.
Edit 2: In case caching the RoPE'd values is necessary, it might still be possible to apply the inverse RoPE for recalculation purposes.
Edit 3: By reserving a fixed block of tokens for a summary at the front right after the system prompt, you could now update that summary for the cost of prompt processing the summary without worrying that you modified something at the front.
The issue is that when you have messages 1, 2, 3, 4, 5. And you want to update messages 1,2,3, because they are in the reserved chunk, then 4 and 5 already contain some baked in information from 1,2,3. This means you cannot modify 1, 2, 3 arbitrarily without changing the architecture, but inversely speaking, if all you are doing is generating a summary that is strictly meant to reflect the content of the messages 1, 2, 3, then replacing them with a less information dense version that still retains the same meaning for 4, 5 means you could implement this directly in the inference engine without modifying the model at all. This would just be a special form of sliding window attention where the prefix is updated and causes partial preprocessing.
> The issue is that when you have messages 1, 2, 3, 4, 5. And you want to update messages 1,2,3, because they are in the reserved chunk, then 4 and 5 already contain some baked in information from 1,2,3. This means you cannot modify 1, 2, 3 arbitrarily without changing the architecture
What does "you cannot modify without changing the architecture" mean exactly? I.e. will the runtime crash if you do it wrong, or would it merely risk confusing the LLM?
Because if it's the latter, I wonder if this isn't less of a problem in practice than it sounds.
Humans don't update their memories all at the same time, either. We often end up in inconsistent mental picture that we resolve automatically when it becomes apparent (the brief "pause to think" moments).
Trajectory: >>What is the capital of France? Let me think<<
the KV cache tokens for "let me think" will have baked in them among other things "France", "Paris", "major city", "question"
Now you compact, and reuse the later tokens (suffix)
Compacted trajectory: >>Let me think<<all right, let's see<<
Now the compacted KV cache for "Let me think" will have baked in them "France", "Paris", and the model can be "that is weird, why am I thinking about France and Paris? there is nothing related in the previous context"
So, from this follows there must be a point between "pristine cache" and "100% stale cache" where the surprise is small enough the model can recover. Perhaps with extra hint at the end, "your context cache is no fully rebuilt after a change, except confusion".
Bit like a person who just woke up, or was suddenly distracted, and now has some stray thoughts from previous context floating around. We normally dismiss them. Maybe the models can, too, and this could achieve more flexibility with cache management (therefore much lower costs of context management), at the price of slight capability reduction after context management events?
> Priming is a concept in psychology and psycholinguistics to describe how exposure to one stimulus may influence a response to a subsequent stimulus, without conscious guidance or intention
Also, we are AGI, the model isn't, it can't (re)organize its thoughts as easily as we can.
But nobody knows, what will happen is people will experiment with this, you can run the benchmarks, if it works it will be used, if not, well...
However given it's a super-obvious thing to do, drop middle tool calls from cache and keep the rest without re-prefilling, I would guess it degrades performance quite a lot, otherwise the labs would have been doing this already.
You can apply rotary encoding on the file level to keep relative positioning within a file. Then there would be no global order of files (which could be an issue) but you can now add or remove whole files as you like.
Kimi K3 is an example of a modern LLM that doesn't use positional encodings. It uses NoPE'd MLA for global attention and a variant of Gated DeltaNet (KDA) for local attention. I wonder how much the KDA would degrade if you just did a bounded replay over the last few thousand tokens to recover its approximate state when assembling your context window from chunks of known MLA KV.
Wow. Context management is one of the big remaining hassles with modern LLMs so this could be big. The obvious complication is cache busting so it's also exciting they investigated solutions for that.
The biggest issue with RLM is the loss of context as you descend the call graph. Adjacent stack frames generally don't have any issues. It's the transitive loss of context that begins to get messy. When you have more than two stack frames on top of each other you will start to see this effect. The model loses track of the original goal and has a tendency to drill into a stack overflow.
I've never seen an RLM harness work very well beyond a recursive depth of 1, which is precisely what this paper is constrained to as well.
Can't we do this trick today with any model? Just send the file as next context. Of course you pay the price for cache misses, depending how deep you make changes, while CLM just ignores the recomputation.
One approximation of this is the experimental context management Codex has been moving towards (not released yet). Rather than relying on summary compaction, the model maintains notes as it works and as it approaches the context limit. A new session is just a fresh context with those notes attached, and a pointer back to the previous session.
Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.
I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.
That is similar to what I am thinking... not just edit the context as a file or string, but have a way to evict blocks and replace them with summary notes and also be able to retrieve them on demand.
No public repo but it's not too difficult to point a LLM at the general idea. Since Codex is open source, you can even take a look at how they do it. Here is what the model is instructed to do in codex (look at "guidance_message"): https://github.com/openai/codex/blob/d91294c39edb93d204926b3...
Interesting. It matches my manual workflow with all harnesses (including vanilla web ChatGPT/Gemini/Claude) for the past year or so: when the session gets compacted, or (ideally) when I feel it's about to be, I just tell it to write a handover note, and start a new session.
With some specific workflow I use in some cases (involving leaving long-lived intermediary artifacts), this turned into me pasting a path to handover file in previous agent's session, and handover itself directs the agent to key files from that session to read, and that's it. So far, with this process, at no point I felt any quality degradation (though early on I often see "I need to check how my predecessor did ${something}", followed by surgical spelunking of past chat's history), even as I carry a single piece of complex analytical work over 5+ sessions.
> LLMs have a lot of knowledge but few competencies. If you constrain them to output knowledge and use that to further constrain results, you’ll go far. For context management, I have the system generate `log.jsonl` and `log.py` (which queries the other document). Whenever an action is processed (an error’s corrected etc.) the system adds something to `log.jsonl`. If it needs to know what happens, it uses `log.py` to query and display only the relevant/required information (like a date, errors or attempted fixes) reducing tokens.
- https://alexalejandre.com/interviews/interview-with-claude-r...
yup, I build a set of fs tools in my custom coding harness that worked like this, doing it again in another custom harness that I only expect to take one turn per request, you can still get decent caching by ordering things so most dynamic comes later
how long until cutting out the token middleman and just keep the cache constant sized but edit it and just send the model new information instead of the same thing over and over. token bad latent good.
> We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files.
So, is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come along with it?
MMUs are already emulated in engines like vLLM (paged attention).
>This allows the model to learn what is most important to maintain in context
DeepSeek's Lightning Fast Indexer already does something similar, although without context compaction. It identifies which tokens are most important to attend to, which allows the model to skip irrelevant ones. A similar idea could be used to remove unnecessary tokens from the context altogether, while somehow strengthening the representation of the important ones (increasing their attention weight, merging information from discarded tokens into them, or creating compressed summary representations)
Letting a model manage its own context is very bitter lesson-pilled
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": https://www.completeskeptic.com/p/kv-cache-rules-everything-...
is kvcache a prematural optimization?
It's pretty much required to make it work with large contexts, so probably not
Do you even understand what you're suggesting? Do you understand how the data flow would look like without a kv cache?
Not having it would be like a text editor that loads your entire document from disk and re-saves it every time you press a key. On a floppy drive.
this is done on the model side so it is invisible to the caller. all of deepseeks kv cache compression stuff is basically the same thing. it doesn't really put the entire context into kv cache and does various inference side stuff to decide what to actually use as context.
I would be concerned with context management consuming limited attention resources.
Do you want your agent solving its own memory crisis, or do you want it solving the actual task? It can probably do both at the same time, but I suspect there is a non trivial cost associated with this.
A separate hypervisor agent that manages the main agent's context would be much better in my experience. You can run it on a different schedule and the main agent has to spend zero tokens thinking about it. This also makes it a lot easier to control when caches will be missed.
Maybe have A second model do the management?
Looks like they tested that in the paper, and the code allows for it as well. https://github.com/facebookresearch/context-language-models/... Of course, this is for suggesting context management strategies rather than the actual management afaict.
Yeah maybe a second model can do the actual management
Yeah, that would be the canonical solution. Modularity is better
I was thinking this, not so dissimilar from co-training dflash drafters
An argument against having a separate hypervisor agent.
A separate hypervisor agent at least doubles the cost because (1) the underlying model needs to be the same so that we get the same degree of intelligence, and (2) it needs to have the same context + more tokens for the work it does.
The biggest discovery might actually be that they ignored regular caching rules and kept invalid cache suffixes and it didn't hurt performance
I wonder how bad the performance would be if they plain ignored the whole rotary encoding dance and just back-filled precisely the parts of the cache that changed directly. Would it break the model? Confuse the model? Or would the model internally correct for it?
The rotary issue is a simple rotation, cheap, no reason to skip it.
But this is the kind of thing you could ask your agent to test locally.
The issue is this: rotary encoding currently numbers the tokens on the request level.
If you pre-process a text file, you cannot insert it into the context because even if you numbered the tokens within a file properly, it is numbered wrong in the global context of the request. So you have cached data but the request structure does not allow you to insert it.
I think one way out of this would be to split a token into its request and content components. Then you might have to recalculate the request level information but never the content level information.
Edit: Now that I think about, why even cache the RoPE'd values to begin with? If you only need the RoPE'd data during token generation, then you could just RoPE on the fly instead of baking it into the KV cache, meaning you unlock block level suffix and infix caching, not just prefix caching.
Edit 2: In case caching the RoPE'd values is necessary, it might still be possible to apply the inverse RoPE for recalculation purposes.
Edit 3: By reserving a fixed block of tokens for a summary at the front right after the system prompt, you could now update that summary for the cost of prompt processing the summary without worrying that you modified something at the front.
The issue is that when you have messages 1, 2, 3, 4, 5. And you want to update messages 1,2,3, because they are in the reserved chunk, then 4 and 5 already contain some baked in information from 1,2,3. This means you cannot modify 1, 2, 3 arbitrarily without changing the architecture, but inversely speaking, if all you are doing is generating a summary that is strictly meant to reflect the content of the messages 1, 2, 3, then replacing them with a less information dense version that still retains the same meaning for 4, 5 means you could implement this directly in the inference engine without modifying the model at all. This would just be a special form of sliding window attention where the prefix is updated and causes partial preprocessing.
> The issue is that when you have messages 1, 2, 3, 4, 5. And you want to update messages 1,2,3, because they are in the reserved chunk, then 4 and 5 already contain some baked in information from 1,2,3. This means you cannot modify 1, 2, 3 arbitrarily without changing the architecture
What does "you cannot modify without changing the architecture" mean exactly? I.e. will the runtime crash if you do it wrong, or would it merely risk confusing the LLM?
Because if it's the latter, I wonder if this isn't less of a problem in practice than it sounds.
Humans don't update their memories all at the same time, either. We often end up in inconsistent mental picture that we resolve automatically when it becomes apparent (the brief "pause to think" moments).
Nothing will crash. It might create "surprise".
Example:
Trajectory: >>What is the capital of France? Let me think<<
the KV cache tokens for "let me think" will have baked in them among other things "France", "Paris", "major city", "question"
Now you compact, and reuse the later tokens (suffix)
Compacted trajectory: >>Let me think<<all right, let's see<<
Now the compacted KV cache for "Let me think" will have baked in them "France", "Paris", and the model can be "that is weird, why am I thinking about France and Paris? there is nothing related in the previous context"
So, from this follows there must be a point between "pristine cache" and "100% stale cache" where the surprise is small enough the model can recover. Perhaps with extra hint at the end, "your context cache is no fully rebuilt after a change, except confusion".
Bit like a person who just woke up, or was suddenly distracted, and now has some stray thoughts from previous context floating around. We normally dismiss them. Maybe the models can, too, and this could achieve more flexibility with cache management (therefore much lower costs of context management), at the price of slight capability reduction after context management events?
> Perhaps with extra hint at the end
most of this stuff will be "unconscious" the model will have a pull in a particular direction without being aware
> stray thoughts from previous context floating around. We normally dismiss them
It's not that simple, see https://en.wikipedia.org/wiki/Priming_(psychology).
> Priming is a concept in psychology and psycholinguistics to describe how exposure to one stimulus may influence a response to a subsequent stimulus, without conscious guidance or intention
Also, we are AGI, the model isn't, it can't (re)organize its thoughts as easily as we can.
But nobody knows, what will happen is people will experiment with this, you can run the benchmarks, if it works it will be used, if not, well...
However given it's a super-obvious thing to do, drop middle tool calls from cache and keep the rest without re-prefilling, I would guess it degrades performance quite a lot, otherwise the labs would have been doing this already.
You can apply rotary encoding on the file level to keep relative positioning within a file. Then there would be no global order of files (which could be an issue) but you can now add or remove whole files as you like.
Kimi K3 is an example of a modern LLM that doesn't use positional encodings. It uses NoPE'd MLA for global attention and a variant of Gated DeltaNet (KDA) for local attention. I wonder how much the KDA would degrade if you just did a bounded replay over the last few thousand tokens to recover its approximate state when assembling your context window from chunks of known MLA KV.
Wow. Context management is one of the big remaining hassles with modern LLMs so this could be big. The obvious complication is cache busting so it's also exciting they investigated solutions for that.
Related: "Recursive Language Models" https://arxiv.org/abs/2512.24601
The biggest issue with RLM is the loss of context as you descend the call graph. Adjacent stack frames generally don't have any issues. It's the transitive loss of context that begins to get messy. When you have more than two stack frames on top of each other you will start to see this effect. The model loses track of the original goal and has a tendency to drill into a stack overflow.
I've never seen an RLM harness work very well beyond a recursive depth of 1, which is precisely what this paper is constrained to as well.
Can't we do this trick today with any model? Just send the file as next context. Of course you pay the price for cache misses, depending how deep you make changes, while CLM just ignores the recomputation.
One approximation of this is the experimental context management Codex has been moving towards (not released yet). Rather than relying on summary compaction, the model maintains notes as it works and as it approaches the context limit. A new session is just a fresh context with those notes attached, and a pointer back to the previous session.
Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.
I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.
That is similar to what I am thinking... not just edit the context as a file or string, but have a way to evict blocks and replace them with summary notes and also be able to retrieve them on demand.
Do you have a public repo for your approach?
No public repo but it's not too difficult to point a LLM at the general idea. Since Codex is open source, you can even take a look at how they do it. Here is what the model is instructed to do in codex (look at "guidance_message"): https://github.com/openai/codex/blob/d91294c39edb93d204926b3...
Memento
Interesting. It matches my manual workflow with all harnesses (including vanilla web ChatGPT/Gemini/Claude) for the past year or so: when the session gets compacted, or (ideally) when I feel it's about to be, I just tell it to write a handover note, and start a new session.
With some specific workflow I use in some cases (involving leaving long-lived intermediary artifacts), this turned into me pasting a path to handover file in previous agent's session, and handover itself directs the agent to key files from that session to read, and that's it. So far, with this process, at no point I felt any quality degradation (though early on I often see "I need to check how my predecessor did ${something}", followed by surgical spelunking of past chat's history), even as I carry a single piece of complex analytical work over 5+ sessions.
Claude Roux uses a refined version:
> LLMs have a lot of knowledge but few competencies. If you constrain them to output knowledge and use that to further constrain results, you’ll go far. For context management, I have the system generate `log.jsonl` and `log.py` (which queries the other document). Whenever an action is processed (an error’s corrected etc.) the system adds something to `log.jsonl`. If it needs to know what happens, it uses `log.py` to query and display only the relevant/required information (like a date, errors or attempted fixes) reducing tokens. - https://alexalejandre.com/interviews/interview-with-claude-r...
yup, I build a set of fs tools in my custom coding harness that worked like this, doing it again in another custom harness that I only expect to take one turn per request, you can still get decent caching by ordering things so most dynamic comes later
Eventually the CLM will be a separate model co-trained with the actual model right?
And there will be multiple contexts like hot vs cold pages in DBs.
Speaking of which I am predicting a "Context as a DB" paper within one year
how long until cutting out the token middleman and just keep the cache constant sized but edit it and just send the model new information instead of the same thing over and over. token bad latent good.
Mom can we have deep equilibrium (DEQ) transformers?
No, we have DEQ at home.
DEQ at home: We just put the context into a file first and then prompt the model to edit it.
[flagged]
This was my gut reaction as well but I think this is just a poor abstract. What they actually did appears to be much deeper than reinventing agents.md
> We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files.
So, is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come along with it?
yes, and "kubernetes"
>Do we have to reinvent MMUs for LLMs
MMUs are already emulated in engines like vLLM (paged attention).
>This allows the model to learn what is most important to maintain in context
DeepSeek's Lightning Fast Indexer already does something similar, although without context compaction. It identifies which tokens are most important to attend to, which allows the model to skip irrelevant ones. A similar idea could be used to remove unnecessary tokens from the context altogether, while somehow strengthening the representation of the important ones (increasing their attention weight, merging information from discarded tokens into them, or creating compressed summary representations)
Looks Slop. Isn't this just RLMs? but instead of variable its just a file?
https://x.com/a1zhang/status/2105409935936782708?s=20
Bitter Lesson showing up yet again, this time in context management?
Looks great!
that's interesting!
Did they just rename their AI lab to superinterlligence lab!!
It is pathetic