by lifthrasiir 4 years ago

I am the author of J40; ask me anything.

As you can see it is only mature in the sense that it implements a large subset of the specification, and it has tons of rough edges (I know, I know, proper tests and fuzzing are in progress). I also didn't expect this to weigh more than 5K LoC, so my hope is to produce a parallel Rust version in the future.

greatgib 4 years ago

First, very good job to have created such a simple lib without dependency!

Looking at the code example, I was wondering if the 'oops' error check wouldn't be doable at a frame level, or even have some recoverable errors for within a frame?

When loading that kind of objects (image, vidéos,...) It is always annoying when libs have an all of nothing behavior.

  • lifthrasiir 4 years ago

    J40 is indeed designed with a possibility of partial decoding in mind, but both the API and the internal architecture is currently less prepared for such use cases.

phao 4 years ago

Where could one go to read more about the mathematics behind the format, its compression techniques, etc? I remember reading that jpeg 2000 is based on wavelets. Is this the case for jpeg xl?

  • lifthrasiir 4 years ago

    There are multiple strategies working in tandam.

    At the very bottom the entropy coding uses a hybrid of LZ77, Huffman coding and multi-symbol rANS. This is complemented with context modeling which should be pretty familiar to anyone knows Brotli.

    For the lossless (modular) mode the main strategy involves finding a good decision tree to compute a prediction for each pixel, which can be learned for each instance (thus named "meta-adaptive"). This mode allows for a number of additional transformations, some of which also function as a progressive encoding.

    For the lossy (VarDCT) mode the actual transformation is a (vast) superset of the original JPEG, but a lot of more transformations---mostly related to DCT---are available and there are tons of contexts for each coefficient that can be exploited. Not exactly specific to JPEG XL, but libjxl also features a very good psychovisual model to optimize the resulting visual quality.

    Besides from those main modes, there are additional whole-image transformations, color transformations (images can use an absolute color space named XYB when they are allowed to be lossy) and image features. Lastly, images can have multiple frames, some of which are animated and some of which can be merged together when the frame duration is zero, with fully configurable blending modes.

userbinator 4 years ago

Are you familiar with JPEG2000? I looked at this diagram of JPEG XL and was overwhelmed at the complexity, it seems more complex than J2K baseline:

https://en.wikipedia.org/wiki/File:JPEG_XL_codec_architectur...

...and I say this as someone who has written PNG, GIF, and JPEG (baseline sequential) decoders, and didn't think J2K would fit in 5kLoC either.

  • lifthrasiir 4 years ago

    I think the exact code for features in that diagram does fit in 5K lines of code. More specifically an entropy coder is less than 1K LoC, the modular mode and transform is probably 1.5K LoC, the entirety of VarDCT is probably 2K LoC. The problem is that JPEG XL has a lot of other features to make it viable as a long-term format, and I greatly underestimated that complexity at the beginning.

pornel 4 years ago

Which part was hardest to implement so far?

  • lifthrasiir 4 years ago

    Modular predictors (`j40__modular_channel` and related functions in J40). The JPEG XL specification was written after the reference implementation, and there are lots of bugs especially in that part of specification. There are also a significant number of bugs in DCT-like transformations, but those bugs are relatively isolated while modular predictors can wreck the entire image. Those bugs are being reported to the spec editors, so later implementations will have a much easier time.

    • twotwotwo 4 years ago

      I really appreciate the work that must have gone into narrowing down those bugs. There's a reason for the insistence on multiple independent implementations for new standards; the Internet owes you one.

bjourne 4 years ago

Since the library is distributed as an 8k long header file, doesn't compilation times explode?

  • lifthrasiir 4 years ago

    Most lines of j40.h are only enabled with `#define J40_IMPLEMENTATION`, which should be defined only once.

    • bjourne 4 years ago

      Hm, yes, but it still 8k source and, afaik, it will be needlessly recompiled whenever the object file that includes it is modified.

      • lifthrasiir 4 years ago

        Including j40.h more than once does nothing. So it can only affect the preprocessing phase, but windows.h is multiple times larger than J40 (even after `#define WIN32_LEAN_AND_MEAN`) and if you can `#include <windows.h>` anywhere J40 can't be a real problem.