I am not sure what's the current status of Wuffs-based WebP decoder, but that would be another implementation worth comparing since it shares the same goals of safety and speed.
Thanks for the heads-up, I didn't know wuffs had a WebP decoder. Just benchmarked it (b2e6da3), and wpd is about 2-3x faster on lossy stills, 5-7x faster on lossy with alpha, and 7-8x faster on lossless on my M5 Pro (even single-threaded).
Also, wuffs is not bit-identical to libwebp, and animated WebP is completely unsupported, so I'm not sure it is a real option for WebP decoding.
We got a pull request (https://github.com/google/wuffs/pull/168) in February to add lossy WebP support. Lossless WebP (which in some sense is an entirely different format, just reusing the WebP "brand") has had a Wuffs implementation for a couple of years now.
Anyway, the PR included SIMD acceleration and performance was on par with libwebp (C code).
The PR's code was, as far as I could tell, somewhat or mostly AI assisted. While that's great in terms of features, I still have more confidence in hand-crafted code.
I have since been working to manually rewrite the PR. I'd also like to add animation support, and last month I landed some Wuffs tooling changes re animated PNG, to be better able to (as a comparison baseline) decode and test animated WebP.
A lot of that manual rewrite has been committed, but the SIMD parts haven't landed yet. So yes, for what's on the main branch (not the PR), performance is not as good as libwebp yet, but landing the SIMD parts should fix that.
> wuffs is not bit-identical to libwebp
That's news to me. PR 168 says it produces pixel-identical output to libwebp. And what's in the main branch aims to be pixel-identical, e.g. for YUV to RGB conversion, it implements libwebp's formulae, not libjpeg's formulae. Both use BT.601, but libwebp uses studio range and libjpeg uses full range.
Can you link to some example .webp images that are not bit-identical?
Congratulations to the Halide folks on wpd, their new decoder. The performance numbers are impressive. If wpd (a Rust library) works for you, great, you should use it!
But since Wuffs was mentioned, I'll just drop a few selling points for why you'd still consider Wuffs.
1. Wuffs' implementation is transpiled to C code (and that in-C-form is checked into the repository, as well as into the leaner google/wuffs-mirror-release-c Github repo). If your existing project is C/C++, not a Rust one, then it's very easy to add Wuffs as a dependency. It's like adding any other third-party C library. It's just not hand-written .c code. It's hand-written .wuffs code that gets transpiled to a single-file C library, as easy to integrate as the STB libraries but memory-safe.
1.a. Similarly, if you're a Python project, or Java, or whatever, if you can wrap C code, you can wrap the Wuffs library (in its C form) and still get in-process, memory-safe image decoding without having to add a new toolchain to your build process.
2. Wuffs is a zero-capability language. It's a language for writing (safe) libraries, that only compute. It's not a language for writing applications. The Wuffs language cannot open files or write to the network. It can't even dynamically allocate memory. That means that Wuffs code can operate under `SECCOMP_MODE_STRICT` sandbox that prohibits basically everything except reading from stdin and writing to stdout.
2.a. For out-of-process, extremely memory-safe image decoding, the example/convert-to-nia/convert-to-nia.c program in the Wuffs repository reads an image (JPEG, PNG, WebP, etc) on stdin and writes NIA (a trivial image format, similar to Farbfeld) on stdout. Even if you don't want to audit the Wuffs language and toolchain itself, the security review for that convert-to-nia program is also absolutely trivial, because one of the first things that the main function does is to self-impose a `SECCOMP_MODE_STRICT` sandbox.
3. Wuffs uses intrinsics for SIMD, like C/C++, but memory safety is enforced on all loads and stores. And the toolchain also enforces that Wuffs general code can't call into Wuffs AVX2-using code unless your cpuid is AVX2-capable. The memory safety story of all of that is a bit better than just "we do SIMD via assembly in unsafe blocks". Wuffs (the language) doesn't even have an "unsafe" keyword.
These are all valid points, and I think the wuffs project makes a lot of sense for certain use cases. But, I think when characterizing wpd, "we do SIMD via assembly in unsafe blocks" is a bit of a mischaracterization. wpd doesn't implement its decoder using Rust SIMD inside broad unsafe regions. Unsafe code is denied throughout the crate and allowed only in the dedicated assembly binding module; without assembly, the crate uses forbid(unsafe_code).
The handwritten SIMD kernels sit behind safe Rust wrappers that calculate or validate the exact slice bounds before producing raw pointers, and CPU-specific routines are only installed after runtime feature detection. wpd even has guard-page tests specifically checking that the assembly kernels don't read or write beyond those wrapper-established windows. Handwritten SIMD was valuable enough that we felt this was all worth it (see some of the dav1d talks for justifications, we share these views).
Wuffs does still have a stronger formal property in the fact that its compiler checks the loads & stores inside the SIMD implementation itself, but most of the interesting memory safety attack surface (parsing and decoding attacker-controlled structure) remains entirely in safe Rust in wpd anyway, so I think wpd is still safe and incontrovertibly a significant improvement over libwebp anyway.
P.S. you can disable wpd's assembly if you feel so inclined, and get similar performance to Wuffs with safety.
Related is Signal Messenger's webpsan crate, which validates webp container syntax before it is passed to libwebp. It stops just short of decoding actual pixel data, however, as the way webp works requires the entire canvas to be allocated in order to fully decode pixel data, which would have just made webpsan a full-on decoder anyway..
Validators separate from parsers have led to many vulnerabilities in the past. There's always something the validator didn't catch that crashed the parser anyway.
As part of Google's PR for their upcoming Gemini 4 Argon LLM release they said they'd rewritten a few things in rust replacing hand written simd. But I don't think webp was mentioned.
They said:
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.
> For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
Not directly mentioned in the article: The decoder is licensed as free software and/but seems to have been developed with extensive use of Claude.
It’s neat that you can compile it without hand-written assembly for more safety, but what does the performance look like in that case? I assume the presented benchmarks don’t show this?
Why not just just have most of it in wasm? There are wasm runtimes that don’t require js, and I’m sure Google can adapt their engine. Then they could integrate the other side with rust
You can always compile to wasm, or ship a wasm decoder. You can even ship a JS decoder, but speed and resource consumption are important, ruling both out.
1. Can be compiled without assembly, and I believe SIMD can be written safely if you think about what it is actually doing
2. The only dependency is zerocopy which is used in a lot of major projects considered to be safe
I am not sure what's the current status of Wuffs-based WebP decoder, but that would be another implementation worth comparing since it shares the same goals of safety and speed.
https://github.com/google/wuffs/tree/main/std/webp
Thanks for the heads-up, I didn't know wuffs had a WebP decoder. Just benchmarked it (b2e6da3), and wpd is about 2-3x faster on lossy stills, 5-7x faster on lossy with alpha, and 7-8x faster on lossless on my M5 Pro (even single-threaded).
Also, wuffs is not bit-identical to libwebp, and animated WebP is completely unsupported, so I'm not sure it is a real option for WebP decoding.
Wuffs author here.
We got a pull request (https://github.com/google/wuffs/pull/168) in February to add lossy WebP support. Lossless WebP (which in some sense is an entirely different format, just reusing the WebP "brand") has had a Wuffs implementation for a couple of years now.
Anyway, the PR included SIMD acceleration and performance was on par with libwebp (C code).
The PR's code was, as far as I could tell, somewhat or mostly AI assisted. While that's great in terms of features, I still have more confidence in hand-crafted code.
I have since been working to manually rewrite the PR. I'd also like to add animation support, and last month I landed some Wuffs tooling changes re animated PNG, to be better able to (as a comparison baseline) decode and test animated WebP.
A lot of that manual rewrite has been committed, but the SIMD parts haven't landed yet. So yes, for what's on the main branch (not the PR), performance is not as good as libwebp yet, but landing the SIMD parts should fix that.
> wuffs is not bit-identical to libwebp
That's news to me. PR 168 says it produces pixel-identical output to libwebp. And what's in the main branch aims to be pixel-identical, e.g. for YUV to RGB conversion, it implements libwebp's formulae, not libjpeg's formulae. Both use BT.601, but libwebp uses studio range and libjpeg uses full range.
Can you link to some example .webp images that are not bit-identical?
Would love to help all implementations improve – would you be able to send Halide an email so we can take this offline? Happy to help if possible!
Wuffs author here.
Congratulations to the Halide folks on wpd, their new decoder. The performance numbers are impressive. If wpd (a Rust library) works for you, great, you should use it!
But since Wuffs was mentioned, I'll just drop a few selling points for why you'd still consider Wuffs.
1. Wuffs' implementation is transpiled to C code (and that in-C-form is checked into the repository, as well as into the leaner google/wuffs-mirror-release-c Github repo). If your existing project is C/C++, not a Rust one, then it's very easy to add Wuffs as a dependency. It's like adding any other third-party C library. It's just not hand-written .c code. It's hand-written .wuffs code that gets transpiled to a single-file C library, as easy to integrate as the STB libraries but memory-safe.
1.a. Similarly, if you're a Python project, or Java, or whatever, if you can wrap C code, you can wrap the Wuffs library (in its C form) and still get in-process, memory-safe image decoding without having to add a new toolchain to your build process.
2. Wuffs is a zero-capability language. It's a language for writing (safe) libraries, that only compute. It's not a language for writing applications. The Wuffs language cannot open files or write to the network. It can't even dynamically allocate memory. That means that Wuffs code can operate under `SECCOMP_MODE_STRICT` sandbox that prohibits basically everything except reading from stdin and writing to stdout.
2.a. For out-of-process, extremely memory-safe image decoding, the example/convert-to-nia/convert-to-nia.c program in the Wuffs repository reads an image (JPEG, PNG, WebP, etc) on stdin and writes NIA (a trivial image format, similar to Farbfeld) on stdout. Even if you don't want to audit the Wuffs language and toolchain itself, the security review for that convert-to-nia program is also absolutely trivial, because one of the first things that the main function does is to self-impose a `SECCOMP_MODE_STRICT` sandbox.
3. Wuffs uses intrinsics for SIMD, like C/C++, but memory safety is enforced on all loads and stores. And the toolchain also enforces that Wuffs general code can't call into Wuffs AVX2-using code unless your cpuid is AVX2-capable. The memory safety story of all of that is a bit better than just "we do SIMD via assembly in unsafe blocks". Wuffs (the language) doesn't even have an "unsafe" keyword.
These are all valid points, and I think the wuffs project makes a lot of sense for certain use cases. But, I think when characterizing wpd, "we do SIMD via assembly in unsafe blocks" is a bit of a mischaracterization. wpd doesn't implement its decoder using Rust SIMD inside broad unsafe regions. Unsafe code is denied throughout the crate and allowed only in the dedicated assembly binding module; without assembly, the crate uses forbid(unsafe_code).
The handwritten SIMD kernels sit behind safe Rust wrappers that calculate or validate the exact slice bounds before producing raw pointers, and CPU-specific routines are only installed after runtime feature detection. wpd even has guard-page tests specifically checking that the assembly kernels don't read or write beyond those wrapper-established windows. Handwritten SIMD was valuable enough that we felt this was all worth it (see some of the dav1d talks for justifications, we share these views).
Wuffs does still have a stronger formal property in the fact that its compiler checks the loads & stores inside the SIMD implementation itself, but most of the interesting memory safety attack surface (parsing and decoding attacker-controlled structure) remains entirely in safe Rust in wpd anyway, so I think wpd is still safe and incontrovertibly a significant improvement over libwebp anyway.
P.S. you can disable wpd's assembly if you feel so inclined, and get similar performance to Wuffs with safety.
Related is Signal Messenger's webpsan crate, which validates webp container syntax before it is passed to libwebp. It stops just short of decoding actual pixel data, however, as the way webp works requires the entire canvas to be allocated in order to fully decode pixel data, which would have just made webpsan a full-on decoder anyway..
https://docs.rs/webpsan/latest/webpsan/
Nice, I didn't know about this!
Validators separate from parsers have led to many vulnerabilities in the past. There's always something the validator didn't catch that crashed the parser anyway.
A validator before an unsafe parser is strictly better than just the unsafe parser, but worse than a safe parser.
As part of Google's PR for their upcoming Gemini 4 Argon LLM release they said they'd rewritten a few things in rust replacing hand written simd. But I don't think webp was mentioned.
They said:
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.
> For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
I’ve compiled libwebp to WebAssembly and made a web page to use it: https://qip.dev/webp-to-png
You can also run the same wasm on the command line:
Do you have performance numbers versus libwebp?
Not directly mentioned in the article: The decoder is licensed as free software and/but seems to have been developed with extensive use of Claude.
It’s neat that you can compile it without hand-written assembly for more safety, but what does the performance look like in that case? I assume the presented benchmarks don’t show this?
Why not just just have most of it in wasm? There are wasm runtimes that don’t require js, and I’m sure Google can adapt their engine. Then they could integrate the other side with rust
You can always compile to wasm, or ship a wasm decoder. You can even ship a JS decoder, but speed and resource consumption are important, ruling both out.
webp builds with fil-C just fine if you want actual memory safety.
This is great but it’s got caveats:
- assembly code. They may have been careful but it’s an escape hatch
- if it has basically any dependencies then those are likely to transitively pull in more unsafe code
1. Can be compiled without assembly, and I believe SIMD can be written safely if you think about what it is actually doing 2. The only dependency is zerocopy which is used in a lot of major projects considered to be safe