I never hit compaction, most of my sessions are 150-300k tokens long with the longest being around 700k. Using sub-agents means that they can't use that cache, have to re-read everything and now multiply this for every sub-agent you call and it just wastes money/tokens. I also don't like how intransparent sub-agents are, I can't follow what they are doing and I can't really steer them. Claude Code has the /agents view but it's clunky and awful to use.
GPT context window is way too low for me and my last experience with it (GPT 5.6 Sol) was so awful and I hit limits way too fast that I cancelled it (and at least got my money back).
I'm no longer using Pi since it got worse IMHO and Claude subs can only be used in Claude Code but I miss the /tree feature which is perfect for first letting the model read & cache the important bits of the codebase and then start your plan from there (as long as you stay in the Cache TTL). Claude Code has /rewind but it's not as good.
I'm only using the 20$ plans.
> Using sub-agents means that they can't use that cache
They can, when they're forked off the main session instead of spawned from scratch.
I admit I'm not a heavy user of subagents but isn't one of the standard use cases for subagents to run a single command, take the output, summarize it for the main agent and pass it up instead of polluting the context?
When cargo fails to build and creates a massive amount of compile errors, you're better off having this preprocess step.
I also don't like subagents.
I usually have my main agent write a wrapper command around things like that as it hits them. The wrapper only surfaces the important info, writes the full log to a file, and the agent gets some instructions on using sed and the like to navigate the output.
It seems to work reasonably well.
>When cargo fails to build and creates a massive amount of compile errors, you're better off having this preprocess step.
Claude already greps and tails every output by itself, a sub-agent would do the same.