I've been using SlicerVM extensively - which is Firecracker MicroVMs for the regular person (and for the irregular with their platform offering) - to run local 'edge' style workloads locally and securly. Agents, local dev CI, etc. It slotted in and replaced my proxmox vm orchestrator, and now I have secure and and fast vms on my laptop wherever I go. It also supports dockerfile style builds if you're wanting a security upgrade from containers (which, you should if you're using agents).
Honestly, while I see firecracker replacing docker on the horizon I don't see firecracker replacing v8 isolates for most edge function execution. Firstly, this article's scenario is a bit unusual in that they were using someone else's isolates - so adding on a few hops; secondly isolates running JS/TS can be statically analyzed quite well, and at scale looking historically for issues and exploits, in many edge compute scenarios this is quite desirable. MicroVMs can have an awful lot more flexibility so to get the same benefit you have to really lock down what is available - the trade-offs for mid-size companies seems to benefit isolates. Obviously netlify is more than big enough and relies heavily on this that it leans in their favour.
Smolvm with it's libkrun vmm provides significantly worse security positioning than slicervms use of firecracker, which leads to slicervm for any dangerous or secure workload.
That is so effing cool — my only fear with it is whether you could turn it into a sustainable business, because I want that project to be around for a long time.
The world: "build a secure, enterprise-ready microVM automation solution - work on it full time, and pay salaries for the staff that work on it"
Also: it has to be free.
So yes you're right, people confuse VC backed companies, and vibe-coded pet-projects for sustainable software.
SlicerVM was started in 2022 and internal only, plenty of YouTube videos and such about it - written completely manually from our Actuated work.
There's a free trial for anyone who wants to play about on their Mac or Linux computer, the comment here is from a real user (unprompted) that knew and used free alternatives previously.
> I see firecracker replacing docker on the horizon
I don’t think this is going to happen because they serve different purposes. Having to boot an entire OS inside a VM is a step backwards compared to containerisation. There are definitely use cases for isolating containers by running them inside a VM (see: Kata) but it generally ‘replacing’ docker, I don’t think that is on the horizon at all
(Also: you need KVM available, you need a rootfs to boot which will be larger than a container, it has no built in support for mounts or really any kind of communication with the host unless you explicitly set up networking etc for it)
> The whole point of Firecracker over normal VM solutions is that "booting an OS" takes milliseconds.
Define "OS". A barebones OS will indeed start in microseconds. But then your app/service won't. Because it will likely need a gazillion things that a barebones OS doesn't provide. Suddenly the startup time isn't microseconds or even milliseconds.
Sadly, we never got the promise of unikernels, and everything requires the full-blown OS to run anything
The VM's I'm talking about don't run some kind of weird bespoke OS, just a normal Linux kernel. You can make Linux boot quite quickly if you strip out everything but the hardware you virtualise, and if you resume a snapshot from CoW RAM, you don't even need to do that.
I don't have any recent measurements myself, but the startup time for a full Linux kernel using Firecracker's snapshotting feature seems to be around 30ms if you optimise calls for latency: https://dev.to/adwitiya/how-i-built-sandboxes-that-boot-in-2...
The measurements in that blog post place Firecracker above Docker in terms of startup performance. Most of that delay is probably the Docker daemon rather than Linux namespace APIs, but that's beside the point.
Hmm... "resource management benefits you get from sharing a kernel" costs the security risks you incur from not having kernel-level isolation.
These days I'm looking at microvms (esp. smolvm from smolmachines.com) by default,
for anything where I'd previously have reached for docker (via colima or more recently orbstack).
If someone in Netlify is listening, could you please add support for Fetchable[1] in Netlify Edge Functions? I opened a ticket 1 week ago but it has gone unanswered[2].
The idea behind Fetchable is to have a (semi) standard across runtimes and environments to be able to do this:
Even Node.js, who are traditionally the laggards in these things, are already working on it and implementing it. This would be very useful for library writers like me to be able to standardize code across envrionments. I'm re-writing npm's `server` (currently WIP here [3]) and I have some ugly workarounds ONLY for Netlify Edge Functions.
> In the past, requests went out to a hosted execution service. Today, they run on MicroVMs inside our own edge network — roughly 5x faster at the median.
So, the execution itself might now be slower as far as we know, they just eliminated some networking from the mix? Misleading.
I'm having trouble understanding/believing this, given that Cloudflare Workers are also v8 isolates and run vastly faster than the 25-40ms that netlify says their isolates took...
from the article: "With our old infrastructure it went out over the internet, ran the edge function, and came back to us to pass on. With the new compute platform, the request is forwarded to a compute node within our network."
As far as I know, Cloudflare Workers have always executed within Cloudflare's network, not gone out to the internet and executed elsewhere (which I read as being in a hyperscaler cloud).
The next time you want to curse AWS, remember they gave us Firecracker, one of the best microvm technologies out there, and the basis of at least a few non-AWS products out there (this one being the newest entry to the list).
I don't think anybody hates AWS as a technology. It's mostly the rent-seeking (pricing almost-free bandwidth to an egregious level, nickel and diming users through fractional pricing calculations that are about as transparent as the US health insurance billing, predatory B2B vendor lock-in practices) that gets on people's nerves.
When MBA (and YC startup) schools teach to build product with a "high switching cost", they do not have consumer's (in this case the developers) interests in mind.
The other issue is that outside of the core offerings (S3, EC2/Lambdas, the logging and queuing services), everything else seems to be in a perpetual state of beta with products shipped by interns. The situation is not as bad as Cloudflare but the bar's on the ground. Pre-LLM, services like Cognito caused developers no end of pain and suffering. Poor documentation, buggy products opaque console UIs, there's no end to complaints when it comes to AWS. If their services were designed well, then companies like Vercel and Heroku shouldn't exist at all.
And what's worse, with all of their efforts at squeezing customers, they still pay the worst of out of all of the big techs for non-senior leadership day to day engineering ICs. At least with Facebook there's commensurate pay. It's like Amazon took a look at Asian "996" culture and figured that if it works for their warehouse staff it should be good for the engineering folks too.
This is highly interesting considering AWS invented the MicroVMs for lambda, yet node on lambda is dog slow (both in latency and throughput). I can traumadump on request. I bet they could use some of this tech especially since their use cases are often not too dissimilar (auth validation, rule checking etc)
It is originally AWS, though we forked Firecracker ~4 years ago now and maintain an internal version with custom hypercalls (command-and-control from within the VM, POSIX-style fork, etc.) and a number of other features. We regularly pull patches and fixes from upstream Firecracker. You can read about the differences here:
> Over the past several months, our team has rebuilt the infrastructure behind Edge Functions, working closely with the team at Unikraft, who wrote about the experience from their side.
1. What's the node + aws story? Did you stick with aws or stay but try a different serverless tech?
2. Is lambda/serverless overhyped? I'm worried about reaching for it too soon.
> When the MicroVM boots up and the JavaScript server begins to listen on a port, we take a snapshot of the MicroVM. [...] we start a new MicroVM from that snapshot.
That sounds scary, since forked RNG states can lead to catastrophic failures in UUID generators or cryptography.
You’d hope so, though it’s not unheard of to get this wrong. Fastly managed to snapshot guest seeds and use them at runtime for RNG in their WASM snapshots a few years ago https://nvd.nist.gov/vuln/detail/cve-2022-39218
v8 isolates aren't actually a great sandbox and I would not trust them implicitly in the AI era. This is probably why they wrap them in an additional sandbox.
Not only are they shared kernel... They're shared process, shared address space, shared memory pool and allocator... In fact, there is very little isolated about them at all.
I bet there are a million ways to cause side channels allowing learning about other code or data on the same machine, and just one V8 bug (of which there have historically been thousands) let's you take over or modify code in another isolate.
But those same things (and more), in the majority case of completely benign workloads, are significantly better for resource utilization and performance.
I don't think their model is "run everything in V8 isolates as the only isolation primitive", I believe it's closer to "run things with V8 isolates as the floor, dynamically trading efficiency for security in response to runtime (and I'd also assume static) analysis". They also add restrictions to make it more difficult/expensive for code to exploit side-channels (ex. changing the resolution and behavior of `performance.now ` and `Date.now` , no multithreading, no SharedArrayBuffer , etc.). Code attempting to exploit side-channels usually has a fingerprint. If you can classify it well enough, and the cost of a false positive is paying for the process isolation you'd otherwise have paid for everything all the time, you probably end up with healthier margins.
I don't disagree that there are real issues, but I don't think that Cloudflare necessarily misrepresents them (though they do perhaps fall quite a bit short of saying "don't run security critical workloads on our platform"). If you can accept the risk though, you get cheap compute with someone else managing all the infrastructure. If you can't, you probably shouldn't be using Workers (and maybe not even cloud compute in general).
It's true that V8 isolates are more risky than micro VMs*. However, it's also true that micro VMs are more risky than giving each tenant their own machine.
At some point you have to decide where along the spectrum you want to put the boundary of "acceptable" for your workload.
Some people argue that isolates are below the necessary threshold and micro VMs are above it. But this isn't really based on any rigorous mathematical analysis, it's mostly vibes. It used to be that people said VMs weren't secure enough for critical workloads, but few people say that these days.
I would argue that we (Cloudflare) have demonstrated that the isolate model can work fine if done carefully: we've been doing it this way for nearly a decade with no breaches.
* Not "strictly riskier", though. There are some risks micro VMs have that V8 isolates do not. Micro VMs that allow tenants to run arbitrary x86 code place a huge amount of trust in the hardware to be exactly correct; a trusted JIT makes it much easier to work around hardware bugs when they are found.
I'm the original poster that said v8 isolates aren't a great sandbox and I would state that I didn't say they were impossible to use as a sandbox. I know you all at cloudflare are doing a ton on top of v8 isolates to make them secure.
I think as a general rule you'd just want multiple uncorrelated layers of isolation. If you aren't willing to spend what cf does on securing v8 isolates I think process isolation + seccomp + v8 isolates might be enough, otherwise all of our browsers would be ticking time bombs.
Without commenting on v8 isolates specifically, this doesn't necessarily hold in any isolation situation; many customers are running code on behalf of their customers, which are often submitting jobs on behalf of theirs, and so on. Isolation breaches within a platform customer can result in significant cross-user data breaches.
You're right, but there are cases when the risk can be managed. Like if you're running an auth lambda in one isolate, and another isolate is running the exact same copy of the code, I'd say malicious exploitation would be low enough a risk, that I'd be comfortable running things like this.
> The likely explanation is that I'm probably getting old and senile, but maybe, just maybe, I am not the only one experiencing this.
Here it looks like you haven’t gotten interested or kept up with some of the field, none of “edge functions”, “v8 isolates”, and “microvms” / “firecracker microvms” are ultra novel or made up, just relatively specialised.
I've been using SlicerVM extensively - which is Firecracker MicroVMs for the regular person (and for the irregular with their platform offering) - to run local 'edge' style workloads locally and securly. Agents, local dev CI, etc. It slotted in and replaced my proxmox vm orchestrator, and now I have secure and and fast vms on my laptop wherever I go. It also supports dockerfile style builds if you're wanting a security upgrade from containers (which, you should if you're using agents).
Honestly, while I see firecracker replacing docker on the horizon I don't see firecracker replacing v8 isolates for most edge function execution. Firstly, this article's scenario is a bit unusual in that they were using someone else's isolates - so adding on a few hops; secondly isolates running JS/TS can be statically analyzed quite well, and at scale looking historically for issues and exploits, in many edge compute scenarios this is quite desirable. MicroVMs can have an awful lot more flexibility so to get the same benefit you have to really lock down what is available - the trade-offs for mid-size companies seems to benefit isolates. Obviously netlify is more than big enough and relies heavily on this that it leans in their favour.
25 USD/m to run a daemon on my own hardware. Yikes.
That seems off to me as well.
Fwiw, you can run this instead free and open source: https://github.com/smol-machines/smolvm
Disclaimer: Am author.
Brew tap: https://github.com/smol-machines/homebrew-tap
This looks amazing! The credential injection trick is particularly cool :)
you beat me to it (I'm singing the praises of smolmachines.com "smolvm" microvms all over the place)
Appreciate your support!
Why is this better than just using Docker?
Kernel level isolation.
Functionality of criu built in so you can get rewind, pause, in an accessible manner.
Embeddable (you can write JavaScript to programmatically use an isolated environment)
Native performance on multiplatform + consistent experience across platforms.
This sounds great. By the way, does "kernel-level isolation" mean "you have to allocate a chunk of RAM to this"? Or is that CPU-level?
EDIT: Ah, looks like it means "runs its own kernel", not "isolates at the kernel" like Docker does.
yes, separate kernels + virtualized hardware via hypervisor.
containers are built on linux primitives & so shares the kernel.
Excellent, thank you. This is definitely useful to me.
amazing!
any point of comparison with microsandbox? [0]
[0] https://github.com/superradcompany/microsandbox
I focus on building the best VM tech.
Good sandboxing is a feature of a good VM.
Outside of that I support GPU and enables something called branchable computing.
Smolvm with it's libkrun vmm provides significantly worse security positioning than slicervms use of firecracker, which leads to slicervm for any dangerous or secure workload.
https://github.com/libkrun/libkrun
Libkrun and firecracker had similar foundations (Rust, KVM, rust-vmm).
Firecracker has a long track record but has a lot of knobs and tunings to get the security right.
smolvm's serve mode confines each VMM by default with a seccomp allowlist, Landlock, a per-VM uid and no_new_privs, much like Firecracker's jailer.
For dangerous workloads, people can do the same things such as skip host mounts and use virtio-net.
It's not a different security class just because it's libkrun vs firecracker
That is so effing cool — my only fear with it is whether you could turn it into a sustainable business, because I want that project to be around for a long time.
I'll keep it going just for you
I think Alex learned his lesson offering a free version running OpenFaaS.
The world: "build a secure, enterprise-ready microVM automation solution - work on it full time, and pay salaries for the staff that work on it"
Also: it has to be free.
So yes you're right, people confuse VC backed companies, and vibe-coded pet-projects for sustainable software.
SlicerVM was started in 2022 and internal only, plenty of YouTube videos and such about it - written completely manually from our Actuated work.
There's a free trial for anyone who wants to play about on their Mac or Linux computer, the comment here is from a real user (unprompted) that knew and used free alternatives previously.
https://unikraft.com/blog/netlify-edge-functions and https://slicervm.com/ both using the lookalike slop design template made me to believe it was a shadowy unikraft project lol.
> I see firecracker replacing docker on the horizon
I don’t think this is going to happen because they serve different purposes. Having to boot an entire OS inside a VM is a step backwards compared to containerisation. There are definitely use cases for isolating containers by running them inside a VM (see: Kata) but it generally ‘replacing’ docker, I don’t think that is on the horizon at all
(Also: you need KVM available, you need a rootfs to boot which will be larger than a container, it has no built in support for mounts or really any kind of communication with the host unless you explicitly set up networking etc for it)
> Having to boot an entire OS inside a VM
The whole point of Firecracker over normal VM solutions is that "booting an OS" takes milliseconds.
Still, I don't think Docker is at risk of being replaced just yet because of the resource management benefits you get from sharing a kernel.
Balloon devices go a long way for ram sharing, though not all the way. CPU scheduling is great.
There's definitely a tradeoff, but IME it's not as drastic as you first think.
Whether it takes milliseconds or seconds though, you are still doing it. You’re still just running a VM
> The whole point of Firecracker over normal VM solutions is that "booting an OS" takes milliseconds.
Define "OS". A barebones OS will indeed start in microseconds. But then your app/service won't. Because it will likely need a gazillion things that a barebones OS doesn't provide. Suddenly the startup time isn't microseconds or even milliseconds.
Sadly, we never got the promise of unikernels, and everything requires the full-blown OS to run anything
The VM's I'm talking about don't run some kind of weird bespoke OS, just a normal Linux kernel. You can make Linux boot quite quickly if you strip out everything but the hardware you virtualise, and if you resume a snapshot from CoW RAM, you don't even need to do that.
I don't have any recent measurements myself, but the startup time for a full Linux kernel using Firecracker's snapshotting feature seems to be around 30ms if you optimise calls for latency: https://dev.to/adwitiya/how-i-built-sandboxes-that-boot-in-2...
The measurements in that blog post place Firecracker above Docker in terms of startup performance. Most of that delay is probably the Docker daemon rather than Linux namespace APIs, but that's beside the point.
Huggingface places Firecracker startup time between 100-300ms while Docker is at 50-200ms, based on Kata: https://huggingface.co/blog/agentbox-master/firecracker-vs-d...
Either way, latency is in the same ballpark of "low enough that it only matters in edge cases".
I would say that kind of got halfway there, when doing serverless.
The gazillion things that the OS doesn't provide is taken care by language runtimes, which can perfectly run directly on top of type 1 hypervisors.
In similar way how some of those languages have bare metal implementations for embedded development, where the runtime takes the OS role.
Naturally on languages with thin runtimes and heavy reliance on POSIX like C and C++, this isn't as straightforward.
Hmm... "resource management benefits you get from sharing a kernel" costs the security risks you incur from not having kernel-level isolation.
These days I'm looking at microvms (esp. smolvm from smolmachines.com) by default, for anything where I'd previously have reached for docker (via colima or more recently orbstack).
> Having to boot an entire OS inside a VM is a step backwards compared to containerisation
If it boots faster than a container ...
Also containers don't work well with remote filesystems, or filesystems in general, because the admin always limits you.
Would love to read a blog post about this if you've every time to cobble one together.
If someone in Netlify is listening, could you please add support for Fetchable[1] in Netlify Edge Functions? I opened a ticket 1 week ago but it has gone unanswered[2].
The idea behind Fetchable is to have a (semi) standard across runtimes and environments to be able to do this:
Even Node.js, who are traditionally the laggards in these things, are already working on it and implementing it. This would be very useful for library writers like me to be able to standardize code across envrionments. I'm re-writing npm's `server` (currently WIP here [3]) and I have some ugly workarounds ONLY for Netlify Edge Functions.
[1] https://fetchable.org/ [2] https://answers.netlify.com/t/support-fetch-for-netlify-edge... [3] https://server-js.com/
Even better, catch up with everyone else and support containers on serverless infrastructure.
That's weird, considering the guy who put the fetchable.org website up works at Netlify.
That's so funny and ironic, I missed that little detail! Hey, this might be a way for adding internal pressure if he's the only one asking for it.
Alex from Unikraft here! Happy to answer any questions about the microVM part of the story from our side.
We also did a couple of technical write ups if you're interested:
- https://unikraft.com/blog/netlify-edge-functions
- https://unikraft.com/customer-stories/edge-functions-netlify
excited to see Unikraft getting some traction...!
I guess I am curious if/how exactly the Netlify announcement relates to unikernels?
> In the past, requests went out to a hosted execution service. Today, they run on MicroVMs inside our own edge network — roughly 5x faster at the median.
So, the execution itself might now be slower as far as we know, they just eliminated some networking from the mix? Misleading.
I'm having trouble understanding/believing this, given that Cloudflare Workers are also v8 isolates and run vastly faster than the 25-40ms that netlify says their isolates took...
from the article: "With our old infrastructure it went out over the internet, ran the edge function, and came back to us to pass on. With the new compute platform, the request is forwarded to a compute node within our network."
As far as I know, Cloudflare Workers have always executed within Cloudflare's network, not gone out to the internet and executed elsewhere (which I read as being in a hyperscaler cloud).
OK, so it's less about isoates vs firecracker than it is about them now doing it in-house...
Perhaps the article says that, but the framing is all wrong
> In the past, requests went out to a hosted execution service. Today, they run on MicroVMs inside our own edge network
The isolates were not being run at the edge.
They were running on the edge, and in the same datacenters but by another provider.
That provider being Deno
The next time you want to curse AWS, remember they gave us Firecracker, one of the best microvm technologies out there, and the basis of at least a few non-AWS products out there (this one being the newest entry to the list).
I don't think anybody hates AWS as a technology. It's mostly the rent-seeking (pricing almost-free bandwidth to an egregious level, nickel and diming users through fractional pricing calculations that are about as transparent as the US health insurance billing, predatory B2B vendor lock-in practices) that gets on people's nerves.
When MBA (and YC startup) schools teach to build product with a "high switching cost", they do not have consumer's (in this case the developers) interests in mind.
The other issue is that outside of the core offerings (S3, EC2/Lambdas, the logging and queuing services), everything else seems to be in a perpetual state of beta with products shipped by interns. The situation is not as bad as Cloudflare but the bar's on the ground. Pre-LLM, services like Cognito caused developers no end of pain and suffering. Poor documentation, buggy products opaque console UIs, there's no end to complaints when it comes to AWS. If their services were designed well, then companies like Vercel and Heroku shouldn't exist at all.
And what's worse, with all of their efforts at squeezing customers, they still pay the worst of out of all of the big techs for non-senior leadership day to day engineering ICs. At least with Facebook there's commensurate pay. It's like Amazon took a look at Asian "996" culture and figured that if it works for their warehouse staff it should be good for the engineering folks too.
> companies like Vercel and Heroku shouldn't exist at all
Crazy that Heroku with its head start has become abandonware. Goes to show that no incumbent is immune to change.
Blame Salesforce.
All else being equal, v8 isolates should be lower latency than firecracker.
As they mention, the v8 latency was because the v8 runtime was hosted externally.
Firecracker has better security model.
This old thread[0] with kentonv author of Cloudflare's workers platform covers some of the differences.
[0]: https://news.ycombinator.com/item?id=31740885
This is highly interesting considering AWS invented the MicroVMs for lambda, yet node on lambda is dog slow (both in latency and throughput). I can traumadump on request. I bet they could use some of this tech especially since their use cases are often not too dissimilar (auth validation, rule checking etc)
> they could use some of this tech
Firecracker is AWS tech, and was behind both Lambda & Fargate. I imagine it's all the enterprisey extras built on top that cause those issues.
It is originally AWS, though we forked Firecracker ~4 years ago now and maintain an internal version with custom hypercalls (command-and-control from within the VM, POSIX-style fork, etc.) and a number of other features. We regularly pull patches and fixes from upstream Firecracker. You can read about the differences here:
https://unikraft.com/blog/unikraft-vs-firecracker
Weird time to advertise your product.
When would you prefer they advertise? Seems entirely relevant.
When they're contributing to the conversation. It's not relevant.
How is it not relevant? In what way doesn't it contribute?
It's on you to prove it does. In what ways does it contribute?
We now know about their unikraft VM, whereas previous we did not.
If only there was an HN guideline about not using the site primarily for promotion.
You said they didn't contribute and weren't relevant. Whether their account primarily exists for promotion is a separate question.
As fly by comment, I bet very few HNers actually read those guidelines.
It's relevant because it's about TFA! I'm assuming you didn't read it: Netlify uses Unikraft MicroVMs
Netlify use Unikraft for microVMs; just explaining the variant of Firecracker they are using.
ICYMI:
> Over the past several months, our team has rebuilt the infrastructure behind Edge Functions, working closely with the team at Unikraft, who wrote about the experience from their side.
fargate is not firecracker
can confirm that this is true for the fargate product used by consumers.
as someone who worked on it
From the Firecracker README:
https://github.com/firecracker-microvm/firecracker
"Firecracker was developed at Amazon Web Services to accelerate the speed and efficiency of services like AWS Lambda and AWS Fargate."
seriously - if netlify can do 15ms warm start + 10ms cold start, wtf is lambda doing for the other 1500ms?
did you get the warm and cold numbers reversed?
that plus sign is addition. netlify claims cold start is only 10ms slower than warm start.
ah, so 15ms warm and 25ms cold. that makes much more sense.
1. What's the node + aws story? Did you stick with aws or stay but try a different serverless tech? 2. Is lambda/serverless overhyped? I'm worried about reaching for it too soon.
> When the MicroVM boots up and the JavaScript server begins to listen on a port, we take a snapshot of the MicroVM. [...] we start a new MicroVM from that snapshot.
That sounds scary, since forked RNG states can lead to catastrophic failures in UUID generators or cryptography.
Firecracker has existing solutions to this, which I assume they're using: https://github.com/firecracker-microvm/firecracker/blob/main...
You’d hope so, though it’s not unheard of to get this wrong. Fastly managed to snapshot guest seeds and use them at runtime for RNG in their WASM snapshots a few years ago https://nvd.nist.gov/vuln/detail/cve-2022-39218
Wish it explained where the v8 isolate latency is coming from compared to microvms
v8 isolates aren't actually a great sandbox and I would not trust them implicitly in the AI era. This is probably why they wrap them in an additional sandbox.
Why are v8 isolates bad, I see speculative execution hacks, but are there others?
v8 isolates are still shared kernel
while microvm's are separate kernel + hardware virtualization through hypervisor guarantees
I wouldn't call it bad either, just different tools for different things
Not only are they shared kernel... They're shared process, shared address space, shared memory pool and allocator... In fact, there is very little isolated about them at all.
I bet there are a million ways to cause side channels allowing learning about other code or data on the same machine, and just one V8 bug (of which there have historically been thousands) let's you take over or modify code in another isolate.
But those same things (and more), in the majority case of completely benign workloads, are significantly better for resource utilization and performance.
I don't think their model is "run everything in V8 isolates as the only isolation primitive", I believe it's closer to "run things with V8 isolates as the floor, dynamically trading efficiency for security in response to runtime (and I'd also assume static) analysis". They also add restrictions to make it more difficult/expensive for code to exploit side-channels (ex. changing the resolution and behavior of `performance.now ` and `Date.now` , no multithreading, no SharedArrayBuffer , etc.). Code attempting to exploit side-channels usually has a fingerprint. If you can classify it well enough, and the cost of a false positive is paying for the process isolation you'd otherwise have paid for everything all the time, you probably end up with healthier margins.
I don't disagree that there are real issues, but I don't think that Cloudflare necessarily misrepresents them (though they do perhaps fall quite a bit short of saying "don't run security critical workloads on our platform"). If you can accept the risk though, you get cheap compute with someone else managing all the infrastructure. If you can't, you probably shouldn't be using Workers (and maybe not even cloud compute in general).
Sources:
* https://gruss.cc/files/scalableisolation.pdf
* https://arxiv.org/html/2110.04751v1
* https://arxiv.org/pdf/2608.17043
* https://blog.cloudflare.com/revisiting-spectre-attacks-on-wo...
It's true that V8 isolates are more risky than micro VMs*. However, it's also true that micro VMs are more risky than giving each tenant their own machine.
At some point you have to decide where along the spectrum you want to put the boundary of "acceptable" for your workload.
Some people argue that isolates are below the necessary threshold and micro VMs are above it. But this isn't really based on any rigorous mathematical analysis, it's mostly vibes. It used to be that people said VMs weren't secure enough for critical workloads, but few people say that these days.
I would argue that we (Cloudflare) have demonstrated that the isolate model can work fine if done carefully: we've been doing it this way for nearly a decade with no breaches.
* Not "strictly riskier", though. There are some risks micro VMs have that V8 isolates do not. Micro VMs that allow tenants to run arbitrary x86 code place a huge amount of trust in the hardware to be exactly correct; a trusted JIT makes it much easier to work around hardware bugs when they are found.
I'm the original poster that said v8 isolates aren't a great sandbox and I would state that I didn't say they were impossible to use as a sandbox. I know you all at cloudflare are doing a ton on top of v8 isolates to make them secure.
I think as a general rule you'd just want multiple uncorrelated layers of isolation. If you aren't willing to spend what cf does on securing v8 isolates I think process isolation + seccomp + v8 isolates might be enough, otherwise all of our browsers would be ticking time bombs.
Despite naming them isolates, the V8 team does not consider them to be a security boundary.
The v8 JIT is very complex and can lead to sandbox escapes if there are type confusion bugs.
But I guess they are good enough to isolate multiple instances of the same code, ran by the same customer in parallel.
Without commenting on v8 isolates specifically, this doesn't necessarily hold in any isolation situation; many customers are running code on behalf of their customers, which are often submitting jobs on behalf of theirs, and so on. Isolation breaches within a platform customer can result in significant cross-user data breaches.
You're right, but there are cases when the risk can be managed. Like if you're running an auth lambda in one isolate, and another isolate is running the exact same copy of the code, I'd say malicious exploitation would be low enough a risk, that I'd be comfortable running things like this.
And I'd say this even is a majority use case.
"In the past, requests went out to a hosted execution service."
They were outsourcing to another company so there's plenty of room for overhead to creep in.
We are running Postgres and Bun on the same tech at Prisma. It’s good.
I don't think isolates are good idea. We need kernel level process isolation. Chrome itself doesn't trust isolates.
They would be even faster with AOT compiled code.
> 5x faster Edge Functions: V8 isolates to Firecracker MicroVMs
Off-topic, but ...
I've noticed over the last 5 years I find myself increasingly incapable of understanding the title of HN posts upon first reading them.
And then sometimes also, reading the article itself does not help either.
The likely explanation is that I'm probably getting old and senile, but maybe, just maybe, I am not the only one experiencing this.
This post's title is a decent example, albeit far from the worse I've read.
I think HN should have a feature to graph the level of incomprehensibility of post titles over time.
An LLM should probably be able to measure this automatically.
> The likely explanation is that I'm probably getting old and senile, but maybe, just maybe, I am not the only one experiencing this.
Here it looks like you haven’t gotten interested or kept up with some of the field, none of “edge functions”, “v8 isolates”, and “microvms” / “firecracker microvms” are ultra novel or made up, just relatively specialised.
If you're counting milliseconds why use Javascript?
V8 is extremely optimized for script startup time.
So the whole rewrite node scripts into Rust and Go isn't really needed. /s
V8 isn't written in JavaScript?
nope
The whole rewrite into compiled languages hasn't yet reached serverless circles.