ssivark 14 hours ago

Between this and the Cloudflare post, this is a lot of words and no simple system level picture. Here's what I think is going on:

1. Homoiconicity: Harness mechanics are kludgy and we need proper homoiconicity to uniformly handle code (tool calls) and data ("natural language") as token streams.

2. Actor semantics: for isolation, encapsulation and concurrency.

3. Object capabilities: injected references and no ambient authority. Capabality-based reflection/introspection is a clean way to discover interfaces and affordances.

"Code mode" or whatever is basically rediscovering this by hacking outward from LLM token streams, instead of from system design principles based on decades of computer science. It is the beginning of treating an LLM as a programming-language runtime participant (any takers for eval/apply?) rather than as a text/token generator with the harness as an ad-hoc interpreter.

If existing implementations of code mode don't already support all this, I anticipate they will keep piling on hacks till they get to this point.

----

I think it was Dan Ingalls who said "An operating system is a collection of things that don't fit into a language. There shouldn't be one.". I see the same for harnesses -- they're awkward middle children which fit neither in an LLM nor in the programming environment.

Maybe the answer is to partner LLMs with Common Lisp or Scheme fibers / Spritely Goblins or Erlang BEAM and be done!

  • searealist 14 hours ago

    It's more like this:

    Right now when a model wants to call 1 or more tools, there is a fixed json schema to describe the tool calls.

    Codemode is like: Why don't we just let the model write a program to call the tools and compose them however it wants? The key is that the programming language exposed to do this (usually javascript) will have APIs available to do some things internal to the harness (like call tools/mcp).

  • v9v 14 hours ago

    > Maybe the answer is to partner LLMs with Common Lisp

    Autolith (https://autolith.rocks/) and some other CL-based harnesses allow the LLM to modify their own harness within the session.

  • mi_lk 13 hours ago

    > no simple system level picture

    proceed to drop the word homoiconicity... Of course I know what it means without looking up

  • nextaccountic 12 hours ago

    It's actually about context management. Calling a MCP tool will load a potentially huge json on your context window, maybe triggering a compaction. Using a subagent for that is less bad, but still expensive (and you wouldn't run every tool call on a subagent)

    But if you had a mcp client cli, you can just pipe it to jq or whatever and extract what you need. or chain multiple tool calls into a single one. None of this will make the agent receive the intermediate text passed between those (the agent might want to save the intermediate results into temporary files however)

    But shell scripting sucks. Javascript or Python is better suited for handling json and things like that. That's what is being called codemode.

    Ok so.. it makes a lot of sense to integrate agents and programming languages. But right now, harnesses are more or less interchangeable and you can hop to another one very quickly. The more coupling between all those moving parts, the hard will lock-in hit. (just picture the mess that is Claude Code having a severe lock in on the ecosystem)

    • mijoharas 12 hours ago

      I agree with most of what you said, but wanted to add that another point is there is a different level of trust in code executing from the harness, and code executing from the bash shell by the agent.

      I block all network access from bash calls, but want to provide certain tools that can call certain services in a specific way. I can have a tool (or mcp) that does just that, and runs outside of the sandbox.

      • Miyamura80 5 hours ago

        Yup this is 100% the best security measure we recommend clients & others, so they get better control over that single-point-of-access

    • CuriouslyC 8 hours ago

      A little nit here, bash isn't the most elegant language, but shell languages have a very nice feature that makes them easier for Agents to write: they're executed left to right, in the same order the tokens are generated. Python/Javascript can execute out of token order (e.g. foo(bar(baz(bak(bat))))), which is more error prone. Additionally, shell is very terse and designed for duct-taping things together.

      I think something like nushell is probably a better direction than trying to use python/javascript as a shell.

      • nextaccountic 7 hours ago

        Maybe shell scripting doesn't suck (for agents). It sucks for people because it's very error prone (even though it's my favorite language actually), and my guess is that something like Typescript will measurably decrease tool call error rate just because the compiler can catch type errors etc. But, maybe bash is so "in distribution" that this doesn't matter

      • ffsm8 2 hours ago

        Well, as you're already referencing Python... Ynow that's solvable, right?

        Pretty old libs around that played around with those concepts, especially in Python

        https://pypi.org/project/plumbum/

  • lukebuehler 12 hours ago

    I agree. Harnesses are quickly becoming the next systems programming frontier. The OS analogy is apt and often made--just like browsers have been compared to becoming "platforms" and "OSes". Another analogy that is obvious is threads/processes and agents and sub-agents: all the concurrency, shared resource access, cooperative multitasking, etc problems are emerging again.

    It will take some time to shake out, but lots of old ideas are new again. Personally, I'm mining lisp machines, homoiconicity, actor models, and other ideas for my harness designs.

  • ignoramous 11 hours ago

    > Dan Ingalls who said "An operating system is a collection of things that don't fit into a language. There shouldn't be one."

    Anil Madhavapeddy's take on such setups: A compiler that just refuses to stop.

      MirageOS is a system written in pure OCaml where not only do common network protocols and file systems and high-level things like web servers and web stacks can all be expressed in OCaml but the compiler just refuses to stop ... compiler, instead of stopping and generating a binary that you then run inside Linux or Windows, will continue to specialize the application that it is compiling and ... emit a full operating system that can just boot by itself.
    

    https://signalsandthreads.com/what-is-an-operating-system / https://archive.vn/yLfkq

  • TacticalCoder 10 hours ago

    > Maybe the answer is to partner LLMs with Common Lisp or Scheme ...

    That LLMs proved to everyone that piping programs at the CLI was indeed the one true way to do things was already quite delicious.

    If now LLMs were to then make the world move to Lisp would really be the icing on the cake.

    I'm all for it!

    • CuriouslyC 8 hours ago

      In terms of token and attention efficiency, lisp's syntax is actually not quite optimal. Ironically, something like xml where you can munge the delimiters together with the tag as a single token for common cases (e.g. '<tr>' and '</tr>' vs '(' 'xyz' ')') is better.

  • otikik 9 hours ago

    Still too many words in my opinion. It's really just 2 things.

    1. Search: The server publishes its tools via a query interface and with gradual level of detail. This way an LLM can discover the tools it needs without loading the ones it doesn't into context. Saves context.

    2. The LLM sends glue code when it needs to use multiple tools together. Said code gets executed by the MCP server itself, and it returns the result. Saves tool calls, which saves tokens.

    On point 2: Imagine there's 2 tools in the server. One lists customers the other gives sales for a given customer. For a task like "add all the sales for the customers whose name begin with A".

    With traditional MCP: the LLM has to use the MCP like a regular HTTP API: first get the customers (1 tool call). Then for every customer begining with A, get their sales (N tool cals). Then add all of the sales (potentially 1 tool call).

    With code mode the LLM sends a single program. The program has code that does things like `customers = call_tool("get_customers"); for each c in customers do if customer.name.begins_with("A") then ...` etc etc. That code is sent to the server for execution. So for the LLM it counts as a single tool call.

    On the downside, the "code execution" on the Server means that it is executing untrusted code coming directly from LLMs. In order to offer this confidently your MCP servers really needs a very tight sandbox in which to execute this "glue code".

    • mongrelion 2 hours ago

      Pardon my ignorance but I truly want to understand this. Don't harnesses and LLMs already do this? Like, I have seen an LLM via pi write and execute some python code to complete a task instead of running a regular bash command. Is this the same principle only more elaborated? Is codemode being presented to the LLM as yet another tool just like read,write,bash and so on?

ohgodhelpplease 9 hours ago

If you're like me and you're still wondering what "codemode" is after reading the article and comments:

It's giving the harness a small js sandbox to compose tool calls and manipulate the data they return before reading it into context.

  • imtringued 7 hours ago

    Yes codemode is just replacing bash with sandboxed JavaScript.

    • magnusokik 6 hours ago

      Makes sense but why not use bun repl or some other repl, rather than a new sandbox.

    • fg137 5 hours ago

      Feels like a solution to the wrong problem.

      • troupo 35 minutes ago

        It kinda makes sense at least at this point. Models already drop down to writing scripts all the time (even for edits), why not make the write scripts for tools.

        It is still kinda bullshit because model training will override any "please create scripts no mistake pls" markdown files thrown at them, so they will happily ignore Codemod

ylxdzsw 17 hours ago

I'm not sure why almost all codemode implementations choose Javascript. I prototyped an agent[1] to use bash as the language for codemode, which in my opinion worked equally well and requires no teaching (there is literally 0 prompt to teach the LLM about codemode. A tool named "bash" is enough to have them know the usage).

[1] https://github.com/ylxdzsw/mu

  • searealist 17 hours ago

    Can do you make a tool or mcp call from bash?

    • clintonb 16 hours ago

      Yes. Invoking an MCP tool is just an HTTP call. You can do it with curl.

      • searealist 16 hours ago

        That's one kind of MCP. Another is a local stdio server.

        Also there are things like subagents, etc (which may be considered tools).

  • the_mitsuhiko 16 hours ago

    > I'm not sure why almost all codemode implementations choose Javascript

    Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global.

    But a big reason is that code mode runs on the harness side so bash is a tricky target in particular.

    • ylxdzsw 14 hours ago

      sandboxing is indeed an advantage (can be an important one!), but

      1. bash can also express concurrency easily, with sync and async (using standard & syntax) support for each command, and standard cancellation (kill, though crude). 2. "way fewer instructions": It requires no instruction for agents to use bash either, except for merely listing the custom commands (view_image, apply_patch, etc.). Also, bash has standard progressive disclosure mechanism (--help) that models will automatically use with no instruction.

      • the_mitsuhiko 13 hours ago

        We might be talking past each other here. The point of codemode is to orchestrate the LLM side tool calls, not to orchestrate scripts that it might execute within Bash.

        In a world where brain and hand are on different machines, getting the bash hands to reach back into the harness brain is something that requires a) putting tools in its hands that it does not know about b) are tricky to set up, usually involving some sort of socket based back channel.

        I tried this quite a bit, by having pi be always there on the hands side, but it causes a lot of complexity and the LLMs really do not understand it well at all.

  • Bonteq 15 hours ago

    > However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.

    The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.</i>

    Does your prototype overcome the limitations mentioned in the article?

    • andreypopp 15 hours ago

      There's no limitation, just have bash commands `read` or `view_image` which communicate back to agent.

      • codybontecou 15 hours ago

        But doesn't Armin mention the limitation?

        > If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.

        Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support?

        (This is Bonteq, I was just logged into the wrong account.)

        • andreypopp 15 hours ago

          it cannot use cat but you can make a command which communicates back to agent. Armin mentions that:

          > While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process

          But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.

          JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode.

          • the_mitsuhiko 15 hours ago

            > But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.

            It becomes much crummier when hands and brain are on different machines.

            • andreypopp 15 hours ago

              Agree, but this orthogonal to JS or bash question. (I mean can always run some bash locally as well, even wasm compiled one).

              • the_mitsuhiko 13 hours ago

                > Agree, but this orthogonal to JS or bash question.

                It is not from my perspective because Codemode runs in the brain, and bash necessarily runs where the hands are. So if the hands need to reach into the brain, I need to set up a communication layer from the hands to the brain.

    • alfiedotwtf 13 hours ago

      Surely LLMs understand how to decide base64, but if that’s not possible you could convert the image to ascii and then send that in

  • anilgulecha 15 hours ago

    Almost every model is fully trained on js. That does not need teaching either.

    Infact harder to sandbox bash (just-bash or brush based) than it is to js or lua, which has fantastic embedded tooling.

dnlgrmm 13 hours ago

Sadly this seems like introducing a very complex apparatus for little gain. I don't need MCP or many tool calls that the agent can program around - in fact the promise of pi was that you basically just need bash and no other tools. The minimally-invasive approach would have been to just "inject" tool calls as virtual bash commands. No code mode required, no hands vs. brains dichotomy. LLM can use language of its own choosing to interact with tools.

  • gf000 9 hours ago

    Let's say you have 40 files and you want to determine which are trash that you could delete.

    You can iterate over them one by one, potentially doing many turns for each and execute it in 40*n turns and a huge context, or you can just let the LLM write a script that is executed "in the brain" and has access to LLM calls itself.

    So you can just do

      for (var f : files) {
        if (llm.model("is this file trash?", [f]) {
          rm f;
        }
      }
    

    And you can even execute that in parallel, each call costing you minimal tokens (not needing every context just a targeted prompt plus much smaller model)

    • dnlgrmm 5 hours ago

      Thanks. I guess that's something you cannot easily do today.

      Given my limitation with Qwen3.8-Flash of about 55t/s tg and ~1kt/s pp most of the work happening in harnesses today (around subagents, excessive use of skills, tools, huge system prompts etc.) is just not very interesting. My workflow works really well without any of it, tbh it seems more of a way to sell more tokens when you realize how much compute these techniques use to achieve modest workflow improvements.

      I can really recommend running your own models to really appreciate the power and compute that goes into your prompts and tokens. Qwen3.8-Flash is good enough (Opus 4.7 level) for all of my coding needs.

planb 10 hours ago

The headline is "What is codemode". As someone who did not know what codemode is, I had to read until the section "Orchestrating The Harness", where it says "If you are not familiar with Codemode, it’s basically just a way to issue tool calls from within some language, in our case JavaScript."

I could then somehow get what is meant from the examples, but the contents of the article do not at all match the headline.

freakynit 14 hours ago

Can someone please tell me how to disable it permanently in pi? I have already disabled it in settings.json, but, it just doesn't get disabled:

    {
      "source": "npm:pi-mcp-adapter",
      "extensions": [
        "-index.ts",
        "-builtin:codemode"
      ]
    }

and

    "autoEnableCodemode": false,
  • the_mitsuhiko 13 hours ago

    MCP or codemode do not enable automatically, so not sure why they enable in the first place for you. Ask pi to figure out why it's one :)

        { "extensions": ["-builtin:codemode"] }
    

    Is all that is needed to turn off codemode entirely. And -builtin:mcp would independently get completely rid of mcp.

lemontheme 15 hours ago

I’m still figuring out codemode. I was using it in a prototype, but ended up stripping it out again, after realizing my small local LLM was using more tokens than usual. It was combining tool calls elegantly in code exactly how I hoped it would. The problem was that when any of those embedded tool calls failed (e.g. on parameter validation) the parent code execution tool call also failed. In response, the LLM kept rewriting large parts of the original code block.

Btw, Monty by the pydantic team is a joy to work with if you need a way to securely run unverified code. It’s a simplified Python dialect. You can also use it from JS, iirc.

  • hamandcheese 15 hours ago

    This somewhat validates a fear I've had (but hadn't tested) of codemode with smaller models. It works fantastically with Opus and Sol but I've always wondered how well it scales down.

WinstonSmith84 5 hours ago

I've been using Codemode with the new Dot so that Dot can coordinate my agents with:

1. Launching workflows, so it avoids some large configuration

2. Reviewing work, with reviewers restricted to read-only tools or specific tools. Or some tools for reviewers to check hashes, correctness, etc.

3. I'm using the widely used pi-extensible-workflows. So say you want to stop a workflow, you gonna call workflow_stop. Or check the status with workflow_status. Without Codemode, Pi exposes an API, that the agent can use. So yes, you can call workflow_status. Or call workflow_stop. With Codemode, the agent can do something more complex like a for loop over all the list of workflow ids and get a status and stop them all (again, it's just an example).

So basically, Codemode allows to build a sort of framework for more correctness like using a typed language vs. JavaScript without a linter. It gives boundaries essentially. And also Codemode allows an agent to do things more easily like using internal Pi extensions.

skybrian 6 hours ago

This seems to come down to object model and tooling differences between sandboxed JavaScript and sandboxed Unix. You can make Unix commands do whatever you like, including special commands to send things back to the coding agent. But the output of a Unix command is arbitrary text (by default) and the AI will have to pipe things into some other command when there’s a lot of output. In the coding harness I use, I see OpenAI’s LLMs writing a lot of pipelines, often processing the output with python. Also the whole thing needs to run in a VM.

A JavaScript sandbox works somewhat differently since the output is an object. Tools are functions that naturally return objects. Composing functions and async calls work differently. If a function returns a very large object, the coding harness could be smarter about presenting the result to the LLM. Maybe JacaScript works better than bash for calling mcp APIs?

But you can have both! The LLM can get direct access to a JavaScript sandbox, which in turn provides an API to access a Linux sandbox. This seems to be what pi is doing with code mode?

mi_lk 14 hours ago

Why is this a personal blog, not a Pi post?

https://earendil.com/posts/you-said-no-mcp/ remains difficult to read without understanding what codemode is, with some quote like "Now we talked so much about Codemode, it might be worth explaining what that even is."

  • fg137 5 hours ago

    It's a term they made up for something nobody else is interested in.

    I also got very confused. After reading this article (as well), I decided that I would just forget about all of this and wait until it gains traction, if that ever happens.

    I genuinely don't understand what problem they are trying to solve.

ferroman 4 hours ago

tbh but this article is pile of shit. Why would author expect ppl to push themselves through the whole thing to simply understand what it is all about?

pjm331 11 hours ago

I appreciated this post, had never heard of codemode and I think the name is bad but it makes perfect sense

I have wanted something like this ever since I hooked up my first database MCP

Originally I had thought to set up a python runtime with like a standard data science toolkit and the db mcps available as functions, and maybe I still will but JS has been working fine for now

injidup 15 hours ago

How is this different to Claude writing mini scripts to get jobs done which it does quite often?

  • odo1242 15 hours ago

    Claude’s scripts can’t call MCP tools, meaning everything has to be CLIs or libraries. At which point you lose the “everything has self-documenting input-output schema” that Codemode is going for.

maherbeg 8 hours ago

We just need an LLM based Lisp so we can finally close the loop of code is data and data is code.

lionkor 13 hours ago

Am I crazy for saying I like codemode? It reduces the amount of turns by about an order of magnitude.

  • mijoharas 11 hours ago

    not at all, it's nice! solves the composability problem while keeping a trust separation between harness/ai bash calls.

aidiveyt 15 hours ago

subagents differ: in my claude code logs every subagent cache write is 5-minute tier, main session 1-hour

soltanov 17 hours ago

Recovery after interrupted execution; distinguishing completed side effects from calls that can safely repeat.

Starlevel004 14 hours ago

I've noticed that code mode seems to cause progressive lobotomisation in Sol 6.1 around subagents. The more it uses it, the worse it gets at giving prompts to subagents:

  URLs, datasourceproxy accesspathonlyallowed anddefault404. Do not set authbasic onprivate ports; perplanPodmannetwork trustedinfrastructure nottenant boundary. Publicroutes internal /tinyauth protected internal, proxyGETemptybody /api/auth/nginx toprivate tinyauth3000; proxy_pass_request_headersoff, CookieonlyTinyauthSession header extracted map name/value actual runtime pattern tinyauth-session-[0-9a-f]{8}, X-Original-URL constructed

This is unedited; it's merging words together and spamming keywords.

troupo 14 hours ago

So... What exactly is it in the end? I try to parse the long prose, and couldn't.

It doesn't help that the article is titled "What is Codemode" and then goes to say "If you are not familiar with Codemode, it’s basically just...". You article is supposed to say what it is without anyone being familiar with it.

Like what does this passage even mean:

--- start quote ---

For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally.

--- end quote ---

  • lpedrosa 14 hours ago

    Unfortunately if you're not in the deep end when it comes to these tools, your brain will only be able to parse gibberish.

    I was expecting a bit more care from Armin in explaining what does codemode actually do, from first principles.

  • the_mitsuhiko 13 hours ago

    > Like what does this passage even mean:

    It means that a LLM side tool (bash) can expose larger (and structured) outputs to Codemode than it normally does when that tool is executed straight to the LLM back as text.

    • quantumwoke 8 hours ago

      I understand what codemode is from other blog posts, and I think this is a suboptimal explanation. It seems that you are explaining it at the wrong level each time, either too deep (here) or too high (elsewhere in-thread when you talk about 'hands' and 'brain').

    • troupo 51 minutes ago

      Once again it still makes no sense. I guess I need to be familiar with Codemode before trying to understand an article called "what is Codemode".

  • sqwxl 12 hours ago

    i was also confused. this introductory post by the creators of Codemode is much easier to follow: https://blog.cloudflare.com/code-mode/

  • agentdev001 10 hours ago

    The harness exposes tools to the model. The model, during its turn, can request that the harness execute a tool. It may even request multiple tools.

    However, natively, the model has no way of composing the tools it sees that it can call. It cannot say, effectively, "use the output of the get_email_id tool, as input for the get_email_Metadata tool."

    This is what codemode is. Allowing for the model to interact with the tools that the harness exposes, programmatically.

    So, TLDR; Codemode allows for composing at the harness level, rather than just at the inner-tool level.

    • troupo 50 minutes ago

      Thank you!

  • keybored 7 hours ago

    I read a bit and gave up. Same feeling as when I read “What is gas station” or whatever it was called. But Gas Station was much worse.