3.8 Flash is just quite good, and so is the Antigravity harness.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
Even if agy was the best (it's not, and is missing basic features) you wouldn't rather have a choice?
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
I use a variety of models for various subagents. I don't want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
I have used it for little more than 6 hours or so in total but I'm pretty sure it doesn't have compaction?
How else would it work? Less technical people don't even watch their context usage.
In the olden times, aka like two years ago, AI chats would just stop working or just start slicing off the oldest parts of the context to fit the model's window.
That said, compaction feels like an idea that should work reasonably well, but across all of the major providers and agent tools I've used has never actually produced compelling results, to where if I see I'm getting close to the token limit I prefer to start putting a bow on the project and readying it for a fresh start. Even when I provide a detailed compaction prompt it usually focuses on the wrong stuff.
Do you have experience with OAI's, it's been known to be to be good for a while now, going off public consensus and my experience.
Yes, and I think it has improved some, but just this week 6 Astra lost the most key details of a project across a compaction and got confused about what we were actually trying to do. I would have preferred to stop at 85%, interactively develop a next-steps prompt and continue from there when ready, rather than seeing it compact and become 5x dumber from one turn to the next.
I have a "wrap up the session" skill that I use when the session gets >50% of its token use. It commits everything, updates documentation, writes a handoff doc, makes sure the todo.md is up to date, etc.
Still works better than compaction.
I do something similar - as I approach the context limits I have a pre-compact flush skill that extracts anything useful from the context, updates the MEMORY.md and my Obsidian vaults (set up as a poor man's graph DB) and so forth. Once everything's been stored I run /compact to keep the general session flow intact. Recently I added a small embedder/vector search setup to the same skill, which seems promising so far.
On another environment I've been doing something roughly similar, but have integrated Hindsight as a kind of all-in-one of the above and am still trying to suss out the best compaction strategy.
This is what I've been working toward as well. It's interesting how having the agent do its own reasoning about what it thinks is the most relevant knowledge to carry forward into the next pieces of work is vastly more effective (and even fast sometimes) than whatever the mystery-meat "compaction" process is.
Mentioned by the author in a recent HN thread, I'm also experimenting with automating this through a tiny issue tracker called epiq [1] that basically lets the agent sessions themselves file tickets with the follow-on tasks and relevant handoff right in them, and then a dispatcher automatically launches those tickets into new agent sessions.
[1]: https://ljtn.github.io/epiq/
The idea to kill "mystery-meat compaction" and use an external handoff primitive is brilliant, but doesn't letting the agent author its own handoff tickets re-introduces the same failure mode?
You would think so, right? But as with others in the thread, I'd found directing the agent itself to prepare the handoff does deliver much better continuity.
I assume Anthropic & friends have noticed this as well and will change how they handle long running sessions, so the gap will likely close over time, but this is definitely where things stand today.
i have active disagreements with teammates on the value of compaction/months-long sessions.
Same teamates also post 'Sol deleted my git repo!' or 'Sorry, ignore those 300 PR comments i was just looking!' ~once a month.
No compelling results because summarization is really hard.
It certainly has compaction (since the public launch I assume) and I HATE it. I have some remedies but nothing perfect yet. It never retains ALL the crucial bits. If a conversation runs into two compactions it is often a sign that I have to abandon it and retain whatever I can, to form a seed prompt for an adjacent conversation.
That is really the biggest beef I have with agy over others, the forced auto compaction at the 250k token threshold (3.8-flash), while the model itself (via API) would be fine with a 1M context window. Even if the model is great, restricting context to 250k tokens (and auto compacting no matter what) limits certain applications and workflows somewhat.
There is auto compaction at 250k? Having used agy for months with multi day sessions, I have never seen this.
Wtf. Can you turn that off? On CC I have compaction completely turned off. I’d rather hit the hard out of context limit at 1M.
I find this discussion interesting. I've had huge increases in accuracy and huge reductions in token usage by capping my Claude models at 200k instead of 1M. I find 1M unusable and wasteful and feel that 200k should be the default. This also makes sense given that the whole reason people use "Ralph loops" is to keep the context window small for all tasks to get better results. Of course, clearing it yourself and manually managing it is better, but if I have 8 projects going in different terminal tabs I'm not watching any one of them that closely to effectively do that.
What are people using 1M context window for?
They only released auto mode in the last 2 weeks. Before that it was bypass permissions or manually approve every single tool call. Antigravity is permanently 6 months behind.
I have a skill that spins up worktrees and isolated services on unique ports so I can work in parallel. Antigravity queues all my prompts and makes me confirm to submit them anytime a long running process like a hot reloading UI is active.
The models are fine, the limits are generous, but the dev experience shit tier. Before they were a Codex clone, AntiGravity was an IDE and during the transition to a clone they outright deleted my IDE. It took them a week to roll out a fix.
For almost a year they didn't allow you to see usage limits. Then when they did show them, they update every ~30 minutes and require 4 clicks to navigate to. It's a little better now, but it's still painfully behind the curve.
Holy shit: the software that works is already there, it’s open source, you just have to clone it, the code writes itself, and Google still manages to fuck it up. I swear, these guys are beyond salvation.
They are now infested with Indian style middle management making them an Infosys / Cognizant / Tata clone.
Do you know a single product from Infosys / Cognizant / Tata done right?
Arguably gemini-cli was done in similar style, but claims on reasons for switching were about speed and efficiency of the internal jetski tool in comparison (antigravity toolkit wraps jetski code)
Is auto mode only in the IDE? I'm not seeing it in the CLI on my end, version 1.2.14.
Does it have /goal feature similar to Codex?
yes
Auto mode?
It absolutely has auto mode.
Via cli switch, but in process w/o fine graining? If so please tell
I think it only has "--dangerously-skip-permissions" Claude and codex auto mode will reject certain actions. No secondary check on gemini/agy AFAIK
agy cli does not have auto mode. I've tried and tried and tried to work with agy cli sandbox-mode and just failed.
in my experience is the only workable solution that doesn't ask confirmation for every step. And I hate working in YOLO mode. Seemingly the Antigravity GUI had some features added in a recent release, but a) I don't want to work with the GUI and b) it was poorly implemented as I couldn't get it to work. VS Code plugins are allowed with subscriptions, but is not the CLI experience of Claude Code I want.
gemini-cli supported 'pre-write diff tabs' (y/n) in external editors like vscode. In Claude Code I heavily use 'pre-write diff tabs' for documentation and miss it sincerely in agy cli.
IMHO Gemini 3.8 flash is fast and good enough, but the agy-suite is below par to say it nice. Someone else in this thread calls agy a terrible harness which is probably more accurate.
PSA, in the agy ui there is a button. It was very annoying until i set it. Slight downside- if I ask it to write a planning doc it will write the doc and then implement it without asking. but as long as you know that, no problem....
do you mean?
also available with shift-tab in agy cli [0]. That is not related to executing commands, only to allowing agy edit files. Unless you set Turbo-mode == yolo-mode, agy gui prompts a zillion times too.
With 'some features added' I'm referencing this but don't see change in daily work (still many prompts making getting-work-done impossible): v2.14.0 (September 15, 2026) "New Permissions System" [1]
Details: Introduced the new unified permissions system, presets (Default, Request Review, Turbo), syntax-highlighted permission requests, and restructured the settings under Global Permissions and project-level Inherit Global.
[0] https://antigravity.google/docs/cli/modes/#available-modes [1] https://antigravity.google/docs/changelog/
have you had any experiences where the agent just made unintended edits to the code as you tried to make it run auto? cuz I'm always skeptical about letting it go auto but there's not much I can do when it gets repetitive
Auto mode means that another model reviews tool calls to attempt to disallow less safe ones. It's different from bypass permissions mode which typically just doesn't filter at all.
No auto mode is the thing that bothers me, I don't trust the cli blindly nor do I trust myself to read every python script it throws at me. Auto mode is an acceptable middle ground in my experience
you could use an ai governance agent if you don't want to manually review every script. and if you already use any which ones do you think are the most recommendable?
Shift+TAB sets accept-edits and plan mode.
Emacs integration over ACP.
They've got Zed, VSCode, Jetbrains... But no Emacs or NeoVIM
agent-shell works with antigravity but I haven't tested it much yet
It works but it's a violation of the TOS to use it.
I would rather not risk my Google account.
It's not - it uses the same interface as editor extensions like the one for VScode.
It does mean however that it cannot operate as flexibly as it it could with raw API, IMO, but agent-shell is essentially designed towards wrapping the official clients
Sorry for the delay, I didn't want to drop a glib half answer on you. Using agy is like going back in time. It's better than Gemini CLI was, but that's a really low bar.
I also had that weird Youtube problem. I had to go without it for several days because signing up for Ultra hijacks your YouTube account for no reason.
1) Try to integrate agy into a workflow. It can't do standard I/O like: tail -200 app.log | claude -p "Find the problem"
2) Hard iteration limits. Preventing runaways is good. Preventing me from looping on purpose is anti-user. See also number 7.
3) Not open source so I can't fix any of these problems.
4) No skills. In 2026. Yikes.
5) No persistent memory (see Claudes auto memory)
6) No sub-agents or orchestration of any type really.
7) Weird hard coded limits and constant API errors on everything (scaling problems?)
8) No /loop command
9) /btw is weird and ephemeral. No way to merge it back to the conversation.
10) Unstable in general.
11) No way to control it via API.
I could keep going on. I would suggest taking a class on Claude Code or Codex then using it for a few months. Swapping is always painful, but it's so worth it. Then if you want try to go back to agy. Don't worry, agy won't have changed much. It improves at a snails pace.
It has both skills (https://antigravity.google/docs/skills/) and agents (https://antigravity.google/docs/subagents/)
Opinions are my own.
I'm not sure this list is correct. Number 4 is especially wrong, since Skills are available with the launch of Antigravity 2:
https://antigravity.google/blog/introducing-google-antigravi...
Isn't /btw meant to be ephemeral?
and has explicit "/copy btw" if you want to save it
I don't know when you last tried agy, but if you ever go back to try it again, you'll hopefully be happy to know it does indeed support skills, sub-agents with pretty good inter-agent communication, and probably more.
edit: removed persistent memory from list since I realized I'm using a plugin for that and it's apparently not native
>I couldn't subscribe to a YouTube Family plan while I had it active
Luckily for me both expired yesterday and I was able to subscribe back again (first Youtube family and then Google AI plan).
Why use Codex CLI if you can use the ChatGPT Linux app (which is a Codex GUI in all but name). Personally I weirdly got used to the terrible TUI stuff.
Why use some GUI when there is a TUI?
Because 'graphical' TUIs are pale imitations of GUIs. I don't have a TUI fetish, despite having grown up with them.
It is not a fetish, some people like this and others like that.
I like CLI more than TUI, and TUI more than GUI where appropriate. For working with text TUI is better, for e.g. images GUI (GIMP).
I love a GUI app but calling TUI usage "that" is utterly limiting and unfortunate to put it mildly. I moved entirely to CLI/TUI just because dumping notes, handoffs, anything, and everything I want even for a literally literary work (grammar etc; let alone coding/planning work) is infinitely better. Also, control edits/changes/shapes/etc better among many other things. GUI doesn't even come close (nope!). That' the reason I moved to almost 100% CLI/TUI.
You're a software engineer living in a middle of an AI revolution and you confuse what IS with what CAN BE? GUI can be good people just didn't do them because it took time and the foundation was shit (for the options that didn't take time). That all changes now when a new class of software can be generated with AI. Unfortunately the people generating software and their users still base the decision on what IS (with highly technical reasoning such as "infinitely better"). People literally have the power to define what IS these days. I've only recently switched from claude and codex cli to their respective apps and those apps while not the best gui apps, already infinitely better than tui. The tui is actually hurting my main usage of them which is controlling agent programmatically. The only really point I'd give for TUI is compatibility. Running coding agent directly on android/ios is nice sometimes.
> That all changes now when a new class of software can be generated with AI.
All the AI-generated UIs I've seen have been very derivative, certainly not eliminating any of the disadvantages of typical GUI interfaces.
Realistically, the whole "overlapping windows" GUI model, and everything that derives from that, was a metaphor geared towards people who'd never seen a computer before. It fit the increasing consumer focus of computing interfaces. It's no wonder that technical people often prefer TUIs.
Maybe AI will bring real advancements in GUIs, but someone's still going to have to make it happen.
GUI is not just your "overlapping windows" straw man, it's interactivity, plus graphical display. Yes via a taxonomy hack TUI is not GUI and there are technical hacks that bring graphics to TUIs these days, but the point remains that we need complex (but not overlapping windows, sure!) interactivity and rich display capability.
That is what we need, and if you're making the argument that a terminal shell is the best place to provide them then I don't know what to say.
I'm pointing out that current GUIs are pretty primitive and haven't undergone much serious thought about functional improvement since Xerox PARC in the late 1970s, and that's why TUIs can still have an edge with technically-inclined people.
It's not that GUIs are inherently worse in principle, but in practice they often are.
The point about overlapping windows is that that "desktop" model permeates the thinking about GUI design, but it's fundamentally limiting and misguided.
> we need complex (but not overlapping windows, sure!) interactivity and rich display capability.
Yep. Pity today's GUIs can't deliver that.
> what IS with what CAN BE
Despite the capitals and trying to come across as someone who SEES "what is COMING" (yay!), you have absolutely no idea what you are talking about.
No wait, I will leave you with a hint. Do whatever you wish to do with that.
Hint: So, you probably want me to use a toolset which is incredibly inferior, as of today, for "me/my usage", just because you feel "IT CAN BE". Right.
This is not about vision or anything like that it's about thinking like actual engineers and understand that what accidentally IS will always be inferior to what is designated to (can) be. TUIs will always be a hack and good by accident.
“ 39. Re graphics: A picture is worth 10K words - but only those to describe the picture. Hardly any sets of 10K words can be adequately described with pictures.”
Perlis has an aphorism for this, as he does every important problem [0].
[0] https://www.cs.yale.edu/homes/perlis-alan/quotes.html
IF you havent written your own harness - you would not understand what you can do when you're writing your own harness . the current set of harnesses - all of them are crap tier. The only clue i can give you - it's not in the model providers interest to have token efficiency - but when you are coding the harness yourself you can shoot for that .
In today's world - and idea stated stated is an idea stolen .
the only clue you can give... why? because it's just words? plenty of opensource harnesses there that invalidate the conspiracy in your only clue that you can give.
What you're saying is sus. If you have a harness that's a tier above frontier labs' offerings, link to it.
It's not sus at all. You're looking at it from the perspective of a general purpose system that is not adapted to your use case. Basically you prompt the model and then let the agent do everything. That's the use case you have in mind when you think that it's about being "a tier above frontier labs' offerings".
It's missing the point. I mean think about the basics, why open the huge bash hole only then to have to close it? If you think about it logically, the only way you can sandbox bash is by writing your own bash implementation specifically for agentic use cases.
Just about every harness is. This is common knowledge, no? Frontier lab TUI tend to steal from the OSS harnesses not the other way around.
Care to back that up with a benchmark? I've not seen that. OpenCode is on here, not exactly dominating though: https://artificialanalysis.ai/agents/coding-agents
Benchmarks are a terrible judge for this.
Why? Sounds like you don't have a way to support your viewpoint.
The process by which benchmarks are setup and run does not correspond at all to how human developers engage with a coding agent. At best it is a loose proxy, and often a bad one.
What benchmarks are usually good at is showing to what degree new models are better than old models. What they are not good at, by construction, is showing that harnesses are well adapted to how people use them.
I presume you're not talking about autonomous agents, right? Because then you could give it a set of tasks and check its success rate. But, even so, are you not giving it tasks? Are you principally interacting with it through discussions that are more difficult to quantify? Even question answering has benchmarks. I'm having trouble imagining how you use them (or "how people use them" in your words).
I don't know if you noticed, but Opus 4.6 was peak for human-computer interaction. Everything has been fairly downhill from there despite better benchmarks, at least in that one regard. Opus 5 and 5.5 are clearly a step above in capabilities than 4.6, and I don't think anyone wants to go back, but 4.7 and 4.8 were arguably worse overall. I genuinely feel I got more done with 4.6 and often switched back, prior to 5 coming out.
Why? Because 4.6 actually talked like a human being. It actually organized its thoughts well, and got the main information across without the wall of text that makes your eyes glaze over. So from the perspective of human-computer interaction and maximizing the productivity of a developer+agent team, 4.7 and 4.8 were regressions. Despite much better benchmark performance.
Even if we consider autonomous agents, that benchmark is not indicative of how well they will interpret *your* requests. Or how well they will interact with other agents in a flock/swarm situation. The benchmark just doesn't cover this. (And the difference can be nontrivial! Sakana AI's published results show two generations of uplifting potential from better harnesses.)
Would you please share some blog posts (non-AI written) where people have tried writing their own harness as a learning path and maybe practically as well? I know I can ask an agent and start, but I want to start that outside an llm really and maybe make them do the heavy lifting once I am ready (have a bit of understanding).
Man good luck with finding that. People who wrote their own harness tend to self select into the type of people that DON’T self-author blog posts.
I think it is very much in the first party providers' interest to chase token efficiency, considering that they are offering fixed price monthly plans, and users may go to a competitor when they hit limits.
> I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[)
This sort of thing alarms me. Having a $bigcorp account becomes a ""social credit"" system where they can ban you from all your personal stuff if they decide that you (or your agents!) are doing stuff they don't like.
I take that more as the user is mixing business with pleasure. When i joined a company using `gcp` heavily, I didn't attach my personal gmail to it - I created a 'business' account and used that. Slightly inconvenient - agree, but it's on the individual to draw the distinctions.
There is still anecdata that Google "knows" you are the same person and will ban all accounts associated with you if you do something they don't like. And of course they never need to explain themselves or listen to your appeal.
And that's why it's very powerful to be able to run AI locally, even if less smart.
There is a pi plugin to use agy directly from it.
You get banned if they catch you.
Yes and I've seen reports of it being an ENTIRE GOOGLE ACCOUNT BAN.
I don't want to mess with antigravity because my google account is too entrenched in my life.
That's why it's a nonstarter for me.
which makes basically any product to build with google a nonstarter.
without having an entirely separate google account with its own separated bans, theres just no ability to trust those
I've read here in previous years about bans propagating to other accounts. I think that if Google can associate you with other accounts that they're not above banning those too.
I was getting excited but thanks for reminding me of this. Not messing with this.
Yup. I don’t plan to do anything sneaky but one wrong question or query that looks like “cyber”, say me fixing a buffer overflow in library I maintain, and all of the sudden my gmail is blocked. Yeah, not worth the risk. I feel like even with a different account they’ll figure out it’s me because well, as ad sellers that’s their business to find out who is who and I will still be banned.
There's an API: you can use Gemini with other harnesses. Isn't the situation exactly like Claude vs Claude Code?
You can't without mortgaging your home to pay enterprise API rate pricing. It's prevented on the plans, and if you find a way around it they don't ban you from Gemini... They ban your entire Google account forever.
Yes it's the same with Claude. However, OpenAI allows you to use any harness you like. Which makes sense and that's the primary reason I have their plan now rather than Googles.
I'm cowboying Gemini on oh-my-pi. Been running OK so far, hopefully I won't get banned, and if so hopefully I'll only lose access to the models, not the storage -- while models are a sort of commodity, my data isn't.
Careful, if they ban your Google account, they might ban you from everything: gmail, google drive, adwords, app engine, youtube, voice, android...
I'll create another email, put it in my Google family and use that to reduce a potential blast radius. There is always the chance the whole family account might go down, but I think as long as I don't use the main account for these inference shenanigans it should be OK.
EDIT: I checked online and I couldn't find any report of complete banning for using third party harnesses, only Gemini service suspension, but It's never hurts to be careful. If my secondary email gets banned, I should be able to use my main email on agy.
It also makes me think if Google Family with Google One could be abused for extending inference limits.
No, Google family now shares limits across all your accounts.
Yes, it's just the 5TB that's per account. It's alright though, since GPT 6 Luna haven't been able to hit subscription limits anyways.
Your comment made me try agy again, but after having to approve and persist every single read tool, I got fatigued really fast. Oh-my-pi shows that these checks aren't really necessary for harness safety, it's best to invest in harness predictability.
I found 3.8 Flash in Agy to be generally better than GPT 6, and only behind from Opus 5.5.
I use Antigravity but for some reason, `agy` in the command line feels very bad/incapable of doing things. I can't quite explain it but the most common issue I run into it is just hanging on being unable to finish a tool call
I remember having this issue ALL THE TIME with Gemini CLI but personally I haven't experienced that yet with agy.
Same! 3.8 Flash does really well for writing code as long as you give it a good design and plan to follow. I use Opus for architecture/design/implementation plans and let Gemini 3.8 work using those. Even on the $20 Pro plan I've only come down to 10% before the weekly reset.
I couldn’t find a way to decline model training and reduce data retention for antigravity or Gemini. Apparently it’s only available on a business/enterprise plan, not personal. Did you manage to solve it? That’s the only reason I don’t use Gemini or antigravity.
The gemini series have also been really strong on text extraction. I've been evaluating models to replace gemini 2.5 and its been hard to find something that performs as well as other gemini models
I tried antigravity a little while ago and it was utterly useless for Objective-C code, tasks that both Claude and Codex handled just fine.
Not only could it not complete the small task, the code was obviously wrong from looking at it and did not even compile.
When I pointed that out it got pissy and insisted the code was perfect and I didn't know how to use a compiler, or the compiler was buggy. Pasting the compiler errors did not help.
Surreal.
> Antigravity
Is that a reference to https://xkcd.com/353/