State of the art tooling for 4k intro for the Web
category: code [glöplog]
Hi, we do have a handful of awesome 4k intros for the Web by HBC/½-bit Cheese from a few years ago, and more recently by Bit Labs. But by looking at this content as well as that produced by the JS1k/JS13k scenes, it seems size coded JS/Web demos suffer from harder size limitations than our Windows native ones. So I've looked into the latest tools and tricks for size coding on the Web (thanks p01, 0b5vr and https://in4k.github.io/wiki/javascript), and tried to quantify how much is actually lost in bytes when going from Windows to the Web.
To do so, I’ve rewritten Elevated in JS for the Web and compared its size to the original Windows version with all I’ve learnt.
Because I’m intimately familiar with this piece, I was able to not simply “ported” it but do an actual apples-to-apples comparison, where I've re-written parts of the intro to better compress for the Web. I pretended to be working towards a JS release for a demo party, trying to make the damn thing fit in 4096 bytes in all ways possible. When keeping every polygon, pixels and sound-sample exactly as in the original, the intro turned out to be a 5.5 kilobyte self-contained HTML. When I relaxed the fidelity constraints a little, I was able to recreate what’s basically the same intro in 5.3 kilobytes. In both cases I sacrificed a bit of code to respect window resizing, first user interaction before sound, etc.
So my conclusion so far is that the penalty for bringing 4k intros to the Web in JS is about 1.33x, i.e., a 4k JS intro for the Web is approx equivalent to a 3 kb Windows intro.
This is of course only data point (one intro, one coder).
Now the more useful information that hopefully others can critizice and improve upon is:
I tried most of the different permutations of the packers and the bundler/minifiers above, but in the end this was the winning worfklow: shader minifacation -> WebGL hashing and literal rewrite -> bundle with Esbuild (not minimize) -> minimize and mangle with Terser -> pre-compress with Roadroller -> compress with Zopfli -> append it to an HTML file with the <svg> stub trick. Many of you might know this workflow already, or maybe you have alternate suggestions? Please share.
Now these are the things I’ve learnt, maybe the align with your experience too?
So in the end, these are my take-aways:
If you’ve done experiments of your own or have better ideas and methods, please let me know. Thanks for reading!
To do so, I’ve rewritten Elevated in JS for the Web and compared its size to the original Windows version with all I’ve learnt.
Because I’m intimately familiar with this piece, I was able to not simply “ported” it but do an actual apples-to-apples comparison, where I've re-written parts of the intro to better compress for the Web. I pretended to be working towards a JS release for a demo party, trying to make the damn thing fit in 4096 bytes in all ways possible. When keeping every polygon, pixels and sound-sample exactly as in the original, the intro turned out to be a 5.5 kilobyte self-contained HTML. When I relaxed the fidelity constraints a little, I was able to recreate what’s basically the same intro in 5.3 kilobytes. In both cases I sacrificed a bit of code to respect window resizing, first user interaction before sound, etc.
So my conclusion so far is that the penalty for bringing 4k intros to the Web in JS is about 1.33x, i.e., a 4k JS intro for the Web is approx equivalent to a 3 kb Windows intro.
This is of course only data point (one intro, one coder).
Now the more useful information that hopefully others can critizice and improve upon is:
- I explored multiple ways to encode the data (script, soundtrack and minified shaders) through multiple layouts and prediction strategies, as well as encodings as plain arrays, Base64 and UTF-16.
- I tried multiple bundlers and code minifiers+manglers+rewriters+AST optimizers, including Esbuild, Terser and Closure, all three with and without Roadroller.
- I tried multiple ways to pack the intro, including RegPack, PNG bootstapping, Sagacity's recent Mashi (intended for 64k intros, not 4k), and a custom thing I’ve built with Zopfli/DEFLATE and a boostrapper (similar to 0b5vr’s Compeko and Sagacity's Mashi).
- I also developed some custom scripts to hash WebGL function calls and replace GL literals by numerical values, which does help.
I tried most of the different permutations of the packers and the bundler/minifiers above, but in the end this was the winning worfklow: shader minifacation -> WebGL hashing and literal rewrite -> bundle with Esbuild (not minimize) -> minimize and mangle with Terser -> pre-compress with Roadroller -> compress with Zopfli -> append it to an HTML file with the <svg> stub trick. Many of you might know this workflow already, or maybe you have alternate suggestions? Please share.
Now these are the things I’ve learnt, maybe the align with your experience too?
- A Web intro needs more boilerplate than a native one. A native intro can open a window, go fullscreen and listen to a key to exit in very few bytes. But a JS intro needs a lot to properly respond to the changing window sizes or going fullscreen, listening to the users first interaction to unlock the sound context creation, and doing the minimally necessary CSS.
- Besides boilerplate, while the core JS code minifies and compresses okay with Zopfli/DEFLATE, it’s hard to beat Crinkler on x86 really.
- WebGL is way too verbose and there's only so much that function hashing and replacing literals can do. WebGPU is even worse, about 300 bytes larger. Now, Elevated touches a large surface area of the WebGL API because it’s a traditional mesh renderer that creates textures, FBOs, depth buffers, etc. Shader-only intros that only touch a handful of API points might be a much better fit for JS/the Web.
- Still, the biggest difficulty for me was storing the large amounts of binary data in this intro. I got some success de-interlazing some of the script data channels and doing more efficient encoding than in the original (funny, Elevated could have been a 3.8k intro probably). But no matter the technique, storing whatever data you have in the end is tricky. Neither Base64 nor UTF-16 actually was able to beat plain text after the DEFLATE pass. Data storage is the main reason my JS version of Elevated is not a 4k intro.
So in the end, these are my take-aways:
- Unless there’s a simple solution to the data problem that I’ve missed, it seems that a 4k intro for the Web can be expected to look more or less like a 3k intro for Windows.
- As of today, intros that are “shader only” (no meshes) and contain less data than Elevated (probably using more procedural yet art-directable scripting and soundtrack) are a better fit for the Web.
- It might be worth exploring storing the bulk of the intro in WASM (appended to the final HTML), since WASM can encode the binary data as actual blobs. I've done early experiments with writing the intro in C with Emscripten that look promising, but I'll report fully when I complete this experiment.
- There are a lot of tools out there available (and I haven’t even touched on shader minimizers), with varying degrees of effectiveness but all claiming being the best for extreme size coding (1kb to 4kb), making learning confusing. I think there's an opportunity for a tool that unifies the current workflow that you otherwise have to ducktaped, that is dedicated for 4k intros like Crinkler.
If you’ve done experiments of your own or have better ideas and methods, please let me know. Thanks for reading!
Thanks for that overview iq, I have had exactly this question.
One comment:
I just want to remind people that wasm code can be written directly e.g. through .wat. I have been experimenting with it. By going through the C/emscripten compiler you lose control with the generated bytes, which may negatively effect compression ratio. Like it does in x86 code.
One comment:
Quote:
I've done early experiments with writing the intro in C with Emscripten that look promising
I just want to remind people that wasm code can be written directly e.g. through .wat. I have been experimenting with it. By going through the C/emscripten compiler you lose control with the generated bytes, which may negatively effect compression ratio. Like it does in x86 code.
Agreed with revival, I would not use C/emscripten but create a .wat file containing a bunch of arrays with your data. Then, expose "getOffset" / "getLength" functions where you pass in the index of the array you want, and you'd simply get back a pointer into the exported WASM memory. You would then use "new Uint8Array(memory.buffer, offset, length)" in JS to access the memory directly without any copying.
In light of the myriad of fancy features available in web APIs, I find it rather sad that we are still stuck with 35 year old compression technology.
brotli is supposed to be the modern replacement for DEFLATE on the web. It is supported as content encoding by all major browsers. It is defined in the standard as a format for DecompressionStream, where it is is supported by both Firefox and Safari, but support is unfortunately still lacking from Chromium-based browsers. Apparently they are concerned about the binary size impact of supporting compression and don't want to ship decompression without compression.
If we are lucky, Google will eventually decide to include it, which would help somewhat. I tried compressing the payload of Elusive by Bypass, and here brotli saved 8% compared to zopfli. Quite an improvement, but it still doesn't make a 5.3k demo 4k.
Stubs for brotli exist, but as far as I gather, they are either specific to Windows or have significantly higher overhead.
brotli is supposed to be the modern replacement for DEFLATE on the web. It is supported as content encoding by all major browsers. It is defined in the standard as a format for DecompressionStream, where it is is supported by both Firefox and Safari, but support is unfortunately still lacking from Chromium-based browsers. Apparently they are concerned about the binary size impact of supporting compression and don't want to ship decompression without compression.
If we are lucky, Google will eventually decide to include it, which would help somewhat. I tried compressing the payload of Elusive by Bypass, and here brotli saved 8% compared to zopfli. Quite an improvement, but it still doesn't make a 5.3k demo 4k.
Stubs for brotli exist, but as far as I gather, they are either specific to Windows or have significantly higher overhead.
A few stats (still on Elusive):
Code:
Original 10700
ZX0 4235
zopfli 3942
XZ (LZMA) 3908
Shrinkler 3805
brotli 3612
Crinkler 3380Correction: the 3380 of Crinkler is without the size, model masks and model weights. These add another 24 bytes.
Yes, I did look into Brotli, but as you said Chrome doesn't support it in the CompressionStream API.
And yes revival and Sagacity, I'll start by exporting only the binary blobs in the WASM (good idea to go straight to it through .wat). I'll report the results on that as soon as I have them.
And yes revival and Sagacity, I'll start by exporting only the binary blobs in the WASM (good idea to go straight to it through .wat). I'll report the results on that as soon as I have them.
Thanks for looking into this!
We did C/WASM/WebGPU without Emscripten in our Luminosity 64k from 2024. I wanted to minimize the need to write code in JS. More or less the JS part only provides a block of memory to the C code and takes care of WebAudio and WebGPU (buffer setup) boilerplate. The C code builds with -nostdlib into WASM, and provides the actual meat of the intro (next to the shaders), like data generation, update functions, filling WebGPU buffers and so on. JS is just "orchestrating" the other building blocks (executing the different passes/shaders).
Packaging the whole thing into a single JS/HTML file (base64 encoding the WASM etc.) was somewhat meh and should be much better with sagacity's tool now.
Likely the whole setup is not suitable for 4k, but since it was quite a learning curve for us combining all the different parts and some of this might be applicable to 4k, I thought I throw it out here. If rendering sticks to SDFs, the C part might lose it's justification (ignoring the audio part).
I would prefer WebGPU being exposed to WASM directly (WASI), but last time I checked this was not available. Hence, SW rendering is the way to go for the time being :)
Source is here.
Packaging the whole thing into a single JS/HTML file (base64 encoding the WASM etc.) was somewhat meh and should be much better with sagacity's tool now.
Likely the whole setup is not suitable for 4k, but since it was quite a learning curve for us combining all the different parts and some of this might be applicable to 4k, I thought I throw it out here. If rendering sticks to SDFs, the C part might lose it's justification (ignoring the audio part).
I would prefer WebGPU being exposed to WASM directly (WASI), but last time I checked this was not available. Hence, SW rendering is the way to go for the time being :)
Source is here.
As an addendum to what warp said: I've been trying to follow the same route but with Rust being compiled to WASM instead of C. There really doesn't seem to be a huge amount of benefit. You still need so much interop with JS in order to do anything in the browser so you end up needing to write this whole command dispatcher.
WASI for WebGPU would solve that part, but this is far from being done and it looks like perhaps it will never become part of browsers at all.
WASI for WebGPU would solve that part, but this is far from being done and it looks like perhaps it will never become part of browsers at all.
For starters, there is quite a difference in tech stack if you go oldschool and the bulk of the intro is in HTML,JS,CSS,webGL++ or in WASM
My experience is primarly in 1kb territory, and more "oldschool", but I have seen time and again that hand minification beats Terser and friends hands down! Every time. The minifiers are doing a great job at making the code smaller, but they are completely unaware of the packer that comes after.
If you go oldschool, then you're gonna get much better results with a hand minification or a simpler AST-ish minifier that respects the patterns you put in place but trim some of the "fat" and enforce some packer-friendly patterns. Then do some DLAS style identifiers remapping to steer towards the global optima.
For EXPI, and later Formas, both oldschool 1kb intros, we did something rudimentatry towards that direction and could get 30-70 bytes improvements overal initial Brotli numbers and land just by 1024 bytes. I've been fiddling in that space again lately and got EXPI down to 987 bytes.
For WASM intros, things will shift towards Crinkler or general purpose packers. But the other day, in the Size Optimization discord, we were floating the idea of DLAS style x86 instruction shuffling (respecting the input/output registers and flags) to find smaller packed results, and the same idea could work for WASM.
About data storage, the ideas already mentioned here apply for Deflate, Brotli, .... as well. You can pad the "code" part with the data and access the unpacked buffers.
Blueberry: How big do you think a wasm implementation Crinkler's unpacker would be ? This would plug very well with sagacity's setup from Mashi.
My experience is primarly in 1kb territory, and more "oldschool", but I have seen time and again that hand minification beats Terser and friends hands down! Every time. The minifiers are doing a great job at making the code smaller, but they are completely unaware of the packer that comes after.
If you go oldschool, then you're gonna get much better results with a hand minification or a simpler AST-ish minifier that respects the patterns you put in place but trim some of the "fat" and enforce some packer-friendly patterns. Then do some DLAS style identifiers remapping to steer towards the global optima.
For EXPI, and later Formas, both oldschool 1kb intros, we did something rudimentatry towards that direction and could get 30-70 bytes improvements overal initial Brotli numbers and land just by 1024 bytes. I've been fiddling in that space again lately and got EXPI down to 987 bytes.
For WASM intros, things will shift towards Crinkler or general purpose packers. But the other day, in the Size Optimization discord, we were floating the idea of DLAS style x86 instruction shuffling (respecting the input/output registers and flags) to find smaller packed results, and the same idea could work for WASM.
About data storage, the ideas already mentioned here apply for Deflate, Brotli, .... as well. You can pad the "code" part with the data and access the unpacked buffers.
Blueberry: How big do you think a wasm implementation Crinkler's unpacker would be ? This would plug very well with sagacity's setup from Mashi.
Quote:
Blueberry: How big do you think a wasm implementation Crinkler's unpacker would be ? This would plug very well with sagacity's setup from Mashi.
Good question. We are talking deflated Wasm, right?
The Crinkler unpacker without the header would be around 180-200 bytes of x86. Wasm is more verbose, but the deflate might make up for that, so I would guess in the same ballpark or slightly larger, 200-250 bytes.
I changed the pure JS versionof Elevated to now have all data in a BINARY blob, but unfortunatelly, the intro then is about 150 bytes larger than with pure JS one.
I guess you really need a Crinkler style thing, Zopfli DEFLATE alone seems terrible with this binary data.
For the record, I tried two approaches. First I did the .way -> .wasm + instantiate trick. It was a pretty minimal .way, I'didn't even exporting the offsets and lengths, just the data bank:
But then I realized it was just simple to append the binary blob directly to the HTML and access the ArrayBuffer directly. That saved some nice bytes and removed one tool (wat2wasm) form the pipeline.
I think that in the end we do need something better than Zopfli+DecompressionStream, something more Crinkler like. I don't think we can escape it.
I guess you really need a Crinkler style thing, Zopfli DEFLATE alone seems terrible with this binary data.
For the record, I tried two approaches. First I did the .way -> .wasm + instantiate trick. It was a pretty minimal .way, I'didn't even exporting the offsets and lengths, just the data bank:
Code:
The WASM was only 40 bytes larger the the actual data, then read the data with WebAssembly.instantiate() and access to data at intance.exports.memory.buffer.(module (memory (export "memory") 1) (data (i32.const 0) "\01\00..."))But then I realized it was just simple to append the binary blob directly to the HTML and access the ArrayBuffer directly. That saved some nice bytes and removed one tool (wat2wasm) form the pipeline.
I think that in the end we do need something better than Zopfli+DecompressionStream, something more Crinkler like. I don't think we can escape it.
To clarify what the experiment was:
intro.js + intro.bin are Zopfli'ed together and appended to the SVG stub, which decompresses intro.js+intro.bin, splices to extract intro.json and intro.bin, and runs intro.json which can ready intro.bin as an ArrayBuffer/Uint8Array.
This happens to be 150 bytes larger than enbedding the bin data into the js source code as plain text. Using WASM is even slightly wrose.
Surely code (intro.js) and data (intro.bin) would need different compression context/models/algorithms? I tried compressing them separedly, I was specially interested in trying Z_HUFFMAN_ONLY instead of Z_DEFAULT for the data. No luck.
intro.js + intro.bin are Zopfli'ed together and appended to the SVG stub, which decompresses intro.js+intro.bin, splices to extract intro.json and intro.bin, and runs intro.json which can ready intro.bin as an ArrayBuffer/Uint8Array.
This happens to be 150 bytes larger than enbedding the bin data into the js source code as plain text. Using WASM is even slightly wrose.
Surely code (intro.js) and data (intro.bin) would need different compression context/models/algorithms? I tried compressing them separedly, I was specially interested in trying Z_HUFFMAN_ONLY instead of Z_DEFAULT for the data. No luck.
Quote:
Could separating opcodes and operands into different streams be feasible? This has proven useful for 6502 compression optimisation. The main question would be how expensive (in terms of extra code) it is to re-interleave the streams after decompression, to reconstruct the original code.But the other day, in the Size Optimization discord, we were floating the idea of DLAS style x86 instruction shuffling (respecting the input/output registers and flags) to find smaller packed results, and the same idea could work for WASM.
WASM is stack based, and an array of local variables that you push in and out of the stack as needed. It's not made of the traditional stream of opcodes+operands that you can deinterlace and compress separately as we usually do. Something like this (for the imaginary part of a Mandelbrot/Julia iteration):
That doesn't mean we cannot device good compression algorithms for it.
Code:
f64.const 2.0 ;; [2.0]
local.get $zr ;; [2.0, zr]
local.get $zi ;; [2.0, zr, zi]
f64.mul ;; [2.0, (zr * zi)]
f64.mul ;; [2.0 * zr * zi]
local.get $ci ;; [2.0 * zr * zi, ci]
f64.add ;; [2.0 * zr * zi + ci]
local.set $zi ;; []
That doesn't mean we cannot device good compression algorithms for it.
BTW, and maybe to conclude all of my explorations with Elevated, I made a C++ version of it, compiled into WASM and run it with a small JS, the compressed with Tensered the JS, bundled it with the WASM, and did the Zopli compression plus stub thing ---> I am getting a 7kb intro. So (without a Crinkler equivalente), the WASM path is even worse than just doing it all in JS.
End of experiments.
End of experiments.
iq & al. Thank you so much for going through these experiments.
How big was the "plain" Javascript version of Elevated, without any WASM ?
The boilerplate below gives a W&B centered "start" that takes a canvas fullscreen and get an AudioContext. Hardcoding the canvas' widht and height rather than using innerWidth and innerHeight could save a few bytes.
It takes ~331 bytes raw. As an indication, it packs down to 177 with Brotli.
This boilerplate takes ~592 using the WOFF2 Brotli approach and works directly from file://
This boilerplate takes >410 using the PNG or Deflate bootstrappers but requires the --allow-file-access-from-files flag :\
How big was the "plain" Javascript version of Elevated, without any WASM ?
The boilerplate below gives a W&B centered "start" that takes a canvas fullscreen and get an AudioContext. Hardcoding the canvas' widht and height rather than using innerWidth and innerHeight could save a few bytes.
It takes ~331 bytes raw. As an indication, it packs down to 177 with Brotli.
This boilerplate takes ~592 using the WOFF2 Brotli approach and works directly from file://
This boilerplate takes >410 using the PNG or Deflate bootstrappers but requires the --allow-file-access-from-files flag :\
Code:
<style>html,body{margin:0;inset:0;width:100%;height:100%;background:#000;color:#fff;display:flex;align-items:center;justify-content:center}#c{position:fixed}</style><body>start<canvas id=c><script>c.onclick=c.requestFullscreen;c.onfullscreenchange=()=>{c.width=innerWidth;c.height=innerHeight;A=new AudioContext();/*...*/}</script>Quote:
The Crinkler unpacker without the header would be around 180-200 bytes of x86. Wasm is more verbose, but the deflate might make up for that, so I would guess in the same ballpark or slightly larger, 200-250 bytes.
A closer investigation points at a Wasm version being somewhat bigger than this, more like 300-350 bytes. It turns out many of the x86 tricks used in the Crinkler unpacker (bit manipulation, direct memory operations, flags, crc32) are hard to replicate in a compact way in Wasm.
So Elusive clocks in at around 3700 bytes for the deflated Wasm module, 250ish bytes smaller than pure zopfli. Subtract from this the size of the stub to instantiate the module and extract the memory.
In Mashi I actually have a full WASM context model that looks at the different opcodes and switches models based on the opcodes used. This means that all opcodes get encoded in the same "stream" and operands get encoded to a bunch of different streams (loosely based on how related they probably are).
See https://github.com/datatrash/mashi/blob/main/src/dis_model.rs for details.
See https://github.com/datatrash/mashi/blob/main/src/dis_model.rs for details.
Quote:
Maybe i'm missing something here, but aren't the 2.0 literal and the "$" references to the local variables equivalent to operands in the traditional stream? (That many operands are implicit via the stack shouldn't break the concept.)It's not made of the traditional stream of opcodes+operands that you can deinterlace and compress separately as we usually do.
@Krill: Correct, the compiled wasm binary is still a bunch of opcodes+operands.
I wonder if there is a better tradeoff to be found than full Crinkler by not doing actual compression, rather just a transformation that takes advantage of the subsequent deflating.
BWT+MTF comes to mind as something that has been tried successfully before and is very simple to detransform. It doesn't capture the sparse contexts that make Crinkler so effective, but it still might be a better tradeoff for 4k due to its simplicity.
BWT+MTF comes to mind as something that has been tried successfully before and is very simple to detransform. It doesn't capture the sparse contexts that make Crinkler so effective, but it still might be a better tradeoff for 4k due to its simplicity.
I did try BWT for the shaders, but that didn't help me. It was for the shaders only though.
Roadroller (https://lifthrasiir.github.io/roadroller/) does a bit of transformation + light compression, it sort of prepares things for the subsequent deflating indeed. And it helps a lot, at least with my experiment.
This is the pure JS version:
* 37,678 programmer friendly / development
* 36,073 replaced WebGL literals, and hashed WebGL functions
* 34,884 bundled (Esbuild)
* 18,242 minified (Terser)
* 7,204 pre-compressed (Roadroller)
* 5,558 final html (Zopli defalted + SVG stub)
If I remove all data from it, it becomes a 4,441 html. So I might have been wrong in my previous assesment that data was the main issue after all.
Roadroller (https://lifthrasiir.github.io/roadroller/) does a bit of transformation + light compression, it sort of prepares things for the subsequent deflating indeed. And it helps a lot, at least with my experiment.
Quote:
How big was the "plain" Javascript version of Elevated, without any WASM ?
This is the pure JS version:
* 37,678 programmer friendly / development
* 36,073 replaced WebGL literals, and hashed WebGL functions
* 34,884 bundled (Esbuild)
* 18,242 minified (Terser)
* 7,204 pre-compressed (Roadroller)
* 5,558 final html (Zopli defalted + SVG stub)
If I remove all data from it, it becomes a 4,441 html. So I might have been wrong in my previous assesment that data was the main issue after all.
