Carakan now supports the most modern JavaScript standard1! As predicted, there is enough of a story here for several episodes.
It was exhausting work, full of interesting decisions taken for the most varied reasons. Some things simply suited the engine, some had to be accepted as a compromise, and in some places no alternative existed at all.
It took exactly a month. I realize that when almost all of the labor is carried by AI agents, schedules mean little and code is worth nothing, so I quote this one for comparison; normally a couple of large features and a handful of small ones would appear in a week. This time it was a month of continuous work by the agent swarm, ending in 56,637 lines changed in the engine itself and 5,518,216 lines counting tests and documentation.
The tests were the most painful part, and I cannot wait to share that pain. AI agents are good when what you need is a verifiable result rather than some abstract thing. With Test262 in hand we can come as close to that result as it is possible to come, but these tests have to be run after every change, and run in six different configurations (win x64 release, win x64 debug, win x86 release, win x86 debug, linux release, linux debug). Platform code can differ, compiler behavior can differ, anything at all can happen, so it has to be done.
The debug builds caused the most trouble. Debug code adds substantial overhead, and while release builds got through all 81,178 cases in about three hours, the first full run of a debug build took more than thirty hours. There was a mistake behind that, of course, and once I had sorted it out the run came down to a measly six hours.
A test finds a problem. You fix it. You wait six hours. At best you move on. At worst your fix did not work, or opened a new problem, or worked but broke something somewhere else.
Even with an optimal workflow, with the work spread across several machines, this is still utterly draining. But you creep toward the goal, fixing one problem after another, and then, it seems, that is it. One last qualification run is left. If it passes, you can write the result down and celebrate!
I left the final run going overnight on my personal laptop, so that the first thing I saw on waking would be the result. Instead of numbers I saw an empty desktop.
That was the night Windows Update decided to restart the machine. I was furious—the results were incomplete, and another six hours had to go into them.
But the nightmare is behind us now, and I can tell you how it all works.
Let me explain how these tests are arranged in the first place.
Test262 is the common suite of JavaScript checks, developed under the auspices of TC39, the committee that standardizes the language. The number in the name points at ECMA-262, the document describing how JavaScript is supposed to work. The standard states the rules; Test262 collects the programs that let you check whether they are followed.
Each such program asks the engine a small question whose answer is known in advance. The at() method, for instance, takes elements from the end of an array if you hand it a negative index. A simplified check might look like this:
const numbers = [10, 20, 30];
assert.sameValue(numbers.at(-1), 30);
assert.sameValue(numbers.at(-4), undefined);
The first line of the check asks for the last element; the second runs off the front of the array, where there is no element any more. The assert.sameValue function compares the value it received against the expected one and reports an error if they differ. Execution time is not judged here: slowly is fine, but the engine has to answer correctly.
That is not the end of checking the method, of course. There are empty arrays, missing elements, fractional indices and arguments of an utterly unsuitable type. JavaScript even lets you pass an object instead of an index, one that runs a function of its own when it is converted to a number. That function can modify the array or throw an exception. The standard describes those cases too, down to the order of operations: what manages to happen before the error depends on it. That is how one seemingly simple operation acquires dozens of separate checks.
Some of the tests deliberately contain a broken program. To pass one of those, the engine has to produce the right error at the right moment. If a check expects an error at run time and Carakan could not even parse the new syntax, that cannot be counted as a success either.
Many tests are run twice: in ordinary mode and with the "use strict" directive, which turns on the stricter JavaScript rules. Behavior in these modes differs in places, and all of it has to be accounted for and tracked in order to understand at which point something went wrong. The current suite for ECMAScript 2026 holds 42,517 files, which across the combinations of modes gives 81,178 case variants.
The same suite is used by the developers of V8 in Chrome, SpiderMonkey in Firefox and JavaScriptCore in Safari. Every engine has its own tooling for running Test262 and its own additional tests. The shared suite is especially useful when language features are being added: the same programs check different implementations, and the expected answer does not depend on who wrote the engine.
Carakan can be built either as part of the engine or as a separate console program, jsshell—that was done precisely for the convenience and the speed of testing. You can simply hand jsshell a JS file and get an answer back; that is very easy to automate with something like Python, which is exactly what we do.
Some checks need capabilities an ordinary JavaScript program does not have. Asking the engine to run the garbage collector, say, or to create a separate environment for executing code. For this Test262 describes a special $262 service interface. In essence it is a programmatic remote control, and browser developers have to add it themselves.
Asynchronous tests have quirks of their own. The main body of a program can finish before the Promise handlers return a result. Test262 has ways of checking such code, but the engine has to be taught not to exit after the main script has run, and to report on this properly. Here we had to hook into the loop of the Promise job handler.
jsshell is enough for runs during development, but the final qualification still has to be launched in the full environment. For that there is a small Python server that serves the files and accepts the results. The browser opens a controlling page which runs the checks one after another in an iframe. Every test gets its own environment, so that tests run in isolation.
Most of this infrastructure was written back when ES6 was added, and has been refined ever since while keeping the general architecture. The JS script that drives the checks is written in ES5, of all things—it has to work in any edition of Carakan.
The Test262 source files are kept in a separate repository at a pinned commit matching the current language standard. Simply taking the latest version is not an option—fixes to old tests get in along with checks for features still under discussion, and because of that our results can start to “drift.”
We, on the other hand, keep the results of every run: which test passed or broke at which commit. That makes it possible to run independent checks in different environments, to track regressions and to find them.
If only all of this ran a bit faster…
A run of all the tests in single-threaded mode (Carakan cannot do otherwise so far) without a JIT can easily take three hours or more. This is not a peculiarity of Carakan: V8, JavaScriptCore and SpiderMonkey under the same conditions would take just as long and perhaps even longer. But they have both a JIT and multithreading, so they get through the tests in a matter of minutes.
A debug build is predictably slower than a release one because of the extra scaffolding around the code. But more than a day for a single pass is absolutely abnormal. I had to run the tests under gdb, collect the stacks and look at what was going on in there.
It was the garbage collector, of course.
You could write a separate book about the concept of GC—and one has, in fact, been written. In short: it is a mechanism that finds and destroys objects the program no longer needs, releasing the memory allocated for them. The plus of this approach is that an efficient GC frees the programmer from tracking the life cycle of every object. The minus is… well, basically everything else. You have to work out how to track the state of an object automatically with a hundred percent certainty; if an object is reachable from the code, it must not be deleted.
There are programming languages with a garbage collector, languages without one and languages where it is optional. JavaScript has one.
After each test we clear away its temporary environment. The controlling page, however, keeps running, and everything it holds also has to be walked by the garbage collector. In our case those data included a huge catalog of the tests themselves. The controlling script received the descriptions in portions, put them together, unfolded them into a list of run variants and kept all of it in memory. Tens of thousands of records: which had already passed, which were still to be run. Before the start of every next test, the browser was spending resources walking that whole household.
In Win32 Debug this came especially expensive. Even a short function reading an object’s type flag is dressed in compiler checks and debug scaffolding in such a build. Add the time spent walking the GC to that, then sum it with the time spent walking the rest of the code—and there are your 30 hours.
After the controlling script was fixed, everything got several times faster—this applied to debug and production builds alike. It is still long, but at least it is no longer drearily long.
Carakan now passes 81,176 tests out of 81,178 in the suite for ECMAScript 2026, in all six configurations, with results matching exactly. The two that remain are a single test in sloppy and strict modes, checking Function.prototype.toString() for the legacy RegExp.$& property; its unusual name turns into get $& when the getter function is converted to a string. This is a known disagreement between the check and the description of a legacy extension; the extension itself is not part of the ECMAScript 2026 standard.
For the Intl extension 2,450 of 2,486 checks pass. In 34 cases the tests ask for locales that are not in our ICU package. Two more come down to the space before AM when a time is printed. The test expects an ordinary space, while the version of the international CLDR database we use prescribes a narrow non-breaking one.
Now you see the level of fussiness all of this is checked with.
That completeness is only reached in the console jsshell, though. In the full build 980 SharedArrayBuffer cases are deliberately disabled: shared memory is already implemented in Carakan, but it is still left unavailable to ordinary web pages. For it to work safely, document isolation and the work of web workers need more attention.
And now the main thing, probably. Passing Test262 is very, unbelievably, impossibly, supernaturally cool! According to test262.fyi, only eight engines pass it at roughly the same level (the real count is most likely a little higher, since not every hobby engine makes it into those statistics).
Sadly, that still does not guarantee that JavaScript will work on every site. Test262 does not cover the whole language specification, which is exactly why the major browsers have suites of their own (which we will, of course, try to adapt). In some places something does not work because the expected APIs are still missing from the browser itself, and in others… well, so far it is not even clear why. But that can already be tracked and picked apart case by case.
I’m happy!
The ES2020 specification introduced a new primitive numeric type into JavaScript: BigInt.
I have already said a little about how Carakan stores values. Its universal cell takes eight bytes in a 32-bit build and sixteen in a 64-bit one, and it can hold the data and the type tag. A number or a boolean can be written straight into the cell; for a string or an object it holds a reference to separately allocated memory.
On 32 bits the packing is especially tight: certain bit patterns of NaN, the special “not a number” value, are used as tags. The technique is called NaN-boxing. We have been through the details of that packing before; what matters here is that the size of the cell is fixed and the room for service flags was handed out long ago. So the job is not only to implement the type itself, but also to work out how to fit it into the current architecture.
Computers, generally speaking, are nowhere near as good with numbers as they seem. Counting they do well; representing numbers, not so much. Here again one could dig down through the asphalt and into the Earth’s core, telling of binary fractions, negative zeros, positive infinities and other joys of life.
For now let us look at just a couple of them.
To store an ordinary number of the required width you can set aside a fixed quantity of bits. The number itself goes into those bits, but that way we quickly run up against a limited range. So for very large values and for fractions, floating-point numbers are often used instead, keeping the sign, the significant digits and the exponent separately (roughly like the notation ± 1.2345 × 10²⁰). There is still not much room for significant digits, so the number is rounded when necessary. The rules for how such numbers are to be represented and computed are standardized by the IEEE specifications.
Before BigInt, JavaScript had a single numeric type, Number, with the calculation rules of the 64-bit IEEE 754 format. It uses that same notation, only in binary. Precision gets 53 binary digits—about sixteen decimal ones.
Every integer from −(2⁵³ − 1) to 2⁵³ − 1 can be told apart and stored without loss of precision. Beyond that, precision no longer allows every neighboring integer to be represented: first a step of two appears between the available values, then one of four, and so on. Hence paradoxes like this one in JS:
9007199254740992 + 1 === 9007199254740992 // true
Another gag: decimal 0.1 has no finite form in binary, much as 1/3 has none in decimal. So 0.1 + 0.2 ends up giving 0.30000000000000004—and you find yourself inventing workarounds out of what looks like thin air.
The new BigInt type is meant to solve the first problem. It has to store an integer with a variable number of digits. If the result grows longer, more memory is allocated for it and precision is never lost:
9007199254740992n + 1n // 9007199254740993n
In source code a BigInt is written with the n suffix or produced through the BigInt() function. You can read about its other properties and limitations on the V8 blog; this much is enough for our purposes.
Back to the cells that hold data types.
The long number itself will have to go into a separate block of memory. But the cell still needs a tag by which Carakan will recognize a BigInt and pick the matching arithmetic. The table of existing tags is tightly bound to type checks and machine code generation; extending it would touch the most heavily used parts of the engine.
The garbage collector posed a similar problem. It too tells the types of allocated objects apart by tags in the header. Six bits are set aside for such a tag—64 possible values. There is simply nowhere to stuff a new type.
There was no need to puzzle over it for long: the same task had already been solved for Symbol. The cell has a tag for a reference to an internal engine object whose details can be read from its header. BigInt got such a reference plus an extra flag in the header that tells it apart from the other objects of that kind. The size of the universal cell and of the GC header stayed as it was.
The number itself is stored in parts, each taking 32 bits. Numbers like that can be added in columns, only instead of the familiar digits from zero to nine each part holds a value from zero to 2³² − 1. On overflow we carry one into the next part. The sign is kept separately, and the parts of the number follow the header one after another, in a single block of memory. When the number becomes unreachable, the GC can free the whole block at once.
(I can picture the question: what if someone suddenly wants to port Carakan to processors narrower than 32 bits? Well, let me tell you, BigInt would be nowhere near the top of that list of problems.)
The arithmetic could have been taken from some library, though its memory allocation would then have to be reconciled with the garbage collector and with Opera’s way of handling running out of memory. So a prototype of our own was written first, taught step by step to add with carry, to multiply in columns, to divide digit by digit and the rest of the grade school curriculum. The results were checked against Python, which has arithmetic like this as well. Little by little the implementation grew more involved, got refined—and stayed with us.
The technical limit on a BigInt value in the current implementation is 2⁶⁵⁵³⁶ − 1, just short of twenty thousand decimal digits, which is considerably more than the number of atoms in the observable Universe. The longer the number grows, the more computation it takes, especially for multiplication and division, so this is a sensible place to stop.
Try printing a message about the number of files found. In English that is the whole job: 1 file, 2 files. Russian has three different forms of the word instead of two, and picks between them by the last digits of the number: 1 takes the first form, 2 the second, 5 the third—and at 21 you are back to the form used for a single file, though the number is plainly larger. If you have ever written that function for Russian, forgive the unpleasant reminder.
Dates and numbers tell a similar story. The string 03/04/2026 will be read as two different dates in the US and in the UK, and in Russia writing a date that way will get you fed to a bear. In 1,234 the comma can separate the fractional part or the thousands. Even ordinary alphabetical sorting depends on the language; in Swedish, for instance, Ö comes after Z, while in German it sits next to O. And then there is text segmentation (not every writing system uses spaces), Unicode normalization (ä ⇔ a+◌̈), transliteration…
A set of language and regional settings is called a locale: en-US is English for the US, en-GB for the UK. A program may need to show one and the same sum by the rules of different countries, regardless of the language of the installed operating system.
JavaScript has Intl for this, described by a separate standard, ECMA-402. It can, for example, pick the plural form:
const rules = new Intl.PluralRules("ru");
rules.select(21); // "one"
rules.select(22); // "few"
rules.select(25); // "many"
Once Intl has told it the category, the program can choose the right form of the word. All that was left was a trifle: teaching Intl itself to do that.
None of the problems described are new, and plenty of people have tried to solve them. Different languages and different operating systems had solutions of their own… but honestly, the problem is so universal that nobody, it seems, objected to a single standardized answer.
That de facto standard today is the ICU library, maintained by the Unicode Consortium. It holds text processing algorithms, number and date formatting, calendars, string comparison rules and everything else localization and internationalization require. It is one of the pillars almost the whole digital infrastructure rests on, and it is through ICU that Intl is implemented in Chrome, Safari, Firefox and, probably, in another 99% of the programs that need to work with locales.
“Standard,” however, does not mean “perfect.” ICU is monolithic and heavy, which is why an alternative is being developed under the auspices of the same consortium: ICU4X, written in Rust. Choosing between them, I settled on the older and better-proven option after all: I already have experience integrating third-party C++ libraries, whereas adding bindings around Rust did not make it past Occam’s razor.
Some of the language data comes from CLDR, a maintained database of rules for languages and regions. It holds the names of months and currencies, date patterns, plural rules and a great deal more.
Opera did work with locales inside its own UI before, but what the operating systems provided was quite enough for that. For building Intl it would not have been, quite apart from the fact that different OSes provide different support.
So underneath Intl we have laid ICU 78.3 (a reduced package of 36 locales) + CLDR 48.2 + the Unicode 17.0 tables.
The Intl object is not exactly trivial, but the bulk of its complexity is offloaded onto ICU. The library first has to be connected to Carakan, and it has a canonical path for that: ICU is built separately and exposes a plain C interface, which Opera then calls into.
The alternative—forking ICU4C and rewriting it in Opera’s own idioms—we shall not consider, because it is too early for us to check into the madhouse.
Here is how that integration works, with Intl.DateTimeFormat as the example:
When JavaScript creates an Intl.DateTimeFormat, Carakan parses the parameters and stores the chosen locale, calendar, time zone and date pattern. Those strings and settings live in memory served by the GC.
When formatting, the adapter creates a temporary ICU object, receives the result into a prepared buffer and closes the object through the library’s own function. After the return Carakan turns the result into a JavaScript string. There is no permanent pointer to an ICU formatter inside the JS object. Errors go through the release of temporary resources as well: the library call finishes first, and then the engine reports the failure.
The ICU allocators are connected to Opera through the standard u_setMemoryFunctions mechanism. Memory allocated that way does not become part of the JavaScript heap, and the garbage collector does not deal with it. This lets us control ICU’s allocations and deliberately induce failures during testing (which, as you are about to see, came in handy).
Under memory pressure the ICU library would sometimes crash before it managed to report an error. In Opera, let me remind you, OOM is not a catastrophe: the function is supposed to report it and the calling code to handle it, and that is the whole point of TRAP/LEAVE.
But jumping out from deep inside ICU straight to a Carakan handler is not allowed: we would skip the code that has to release the library’s intermediate objects. This resembles the switching problem in OpPseudoThread I wrote about last week.
There, though, we were switching out of our own code and back into it, and could simply refine the switching logic. Rewriting ICU’s internal out-of-memory handling does not strike me as a good idea.
The solution turned out to be paradoxical. Before an ICU call an emergency memory reserve is prepared, and if an ordinary allocation fails during the call, the adapter hands out the reserve and remembers that a failure occurred. The library runs to the end on that reserve without crashing, the adapter waits for the result and then makes no use of it at all, simply reporting the lack of memory to Carakan. Naturally, if the reserve itself cannot be allocated, that is an instant OOM and ICU is never even reached.
Another interesting case: date formatting crashed with a stack overflow at six nested calls, but started working when the nesting grew to nine. Not exactly logical.
It all falls into place once you recall how the stack (or rather, the stacks) works in Carakan. The allocator gives an OpPseudoThread 16 KiB (a value we inherited, and one that looked perfectly reasonable in an age of thrift), and when that memory runs out the allocator has to add a new segment. The check is simple: less than 3 KiB left—here, have some more memory.
Predicting in advance how much memory will be needed is impossible (well, or possible, but not worth it). We rely on empirical observation and common sense.
When Intl was added, nobody checked how much stack the new functions would want. The functions themselves fit into the stack with a little over 3 KiB to spare, but the very first formatter call demanded 3,696 bytes and ran off the edge of the stack. With nine calls or more, less than 3 KiB was left in reserve, and the allocator added memory right away.
So a separate check had to be added: when Intl is called, the allocator makes sure the pseudo-thread has at least 64 KiB—by measurement that is a fivefold margin even for the fattest call there is. A temporary solution, but sufficient for now.
Next week I will try to tell you about the new RegExp parser, the WeakRef and FinalizationRegistry implementations, shared memory through SharedArrayBuffer and Atomics…
To be continued.