RegExp is a well-known way to give yourself a new problem. Between ES5.1 and ES2026 the standards piled up plenty of powerful new ways to create problems, and every one of them had to be added to Carakan. The additions began with ES2015, but I put off the story of regexes so as to cover the whole evolution in one go. Even with the technical details cut, it came out long, so I won’t dwell on what regular expressions are; if you are reading this, you certainly know them already.
First, a word on how it was done in the original code.
Opera 12.15 had its own RegExp module, implementing all of the core syntax and API of ES5.1 regexes—gim, captures, backreferences, lookahead, the string methods match/replace/search/split—and adding extensions of its own. Some of them were years ahead of the standard: Python-style named groups and references to them ((?P<name>...) and (?P=name)), class intersections ([a-z&&[^aeiou]]—lowercase Latin consonants), character codes in curly braces, and an “extended” mode in which whitespace between pattern elements is ignored. The module also supported “sticky” matching, which would later arrive in the standard as the y flag.
Named groups and class intersections were available from the JavaScript RegExp object without reservation. The x and y flags, however, are hidden behind the ES_NON_STANDARD_REGEXP_FEATURES macro, which is defined nowhere in the 12.15 sources: the module supported those modes but did not let them out. The macro could, admittedly, have been passed in from outside by the build system—I did not check whether the feature is present in the 12.15 binary. Either way, it was a solid and complete implementation for its time.
Note, too, that regexes in Presto are not part of Carakan at all. They are a standalone module, modules/regexp: Carakan hands it the pattern text with its flags and gets back a compiled object. Everything “JavaScript-y”—lastIndex, the result array, the String methods—Carakan does itself, while the module knows only one thing: where in the string a match was found and where its groups are. That boundary existed in the original code and came in very handy: some of the new features required changes only to the module, some only to Carakan, and none of them required rewriting the core JavaScript interpreter.
The engine does not execute the text of an expression directly: first it compiles it into bytecode.
A regex is, in essence, a small program in a small language. Parsing its text again on every search would be wasteful, so the compiler (RE_Compiler) translates the pattern once into a sequence of simple commands for a tiny virtual machine. It resembles a processor with its own instruction set, except that there are few instructions and all of them are about text: “compare a character,” “compare a string,” “check whether a character belongs to a class,” “remember a fork,” “start a group,” “end a group,” “jump,” “success.”
Each instruction is a 32-bit word: the opcode sits in the low byte, a short argument in the upper bits. If the argument doesn’t fit, it takes up the following words. Long literals and character classes are stored separately: a literal is compared by a single instruction, while a class like [a-z0-9_] becomes a separate RE_Class object. For the first 256 characters that is a membership table (a byte per character), for the rest a list of ranges, so ASCII patterns, the most common kind, are checked with a single table lookup.
Here, for example, is a simplified view of the bytecode for /a(b|bc)d/:
0 MATCH_CHARACTER_CS 'a'
1 CAPTURE_START 1
2 PUSH_CHOICE → 5
3 MATCH_CHARACTER_CS 'b'
4 JUMP → 6
5 MATCH_STRING_CS "bc"
6 CAPTURE_END 1
7 MATCH_CHARACTER_CS 'd'
8 SUCCESS
CS in the names stands for case-sensitive; for the i flag there are matching CI instructions. The sources include a native disassembler, RE_Compiler::Disassemble, which prints exactly these listings, but it is switched off even in the debug build, and the places that call it are wrapped in #if 0 on top of that.
The bytecode interpreter is a loop that takes the next instruction, looks at its opcode and executes the matching branch of a switch. Reliable, but every character check also pays for decoding the instruction itself. A JIT removes that middleman: it translates the bytecode into the processor’s machine code, and the regex turns into an ordinary function that compares characters directly without reading any instructions. This JIT has nothing to do with the JIT for JavaScript itself; it is simply the same technology solving similar problems in different subsystems.
That is why a compiled JS function can call a regex in the interpreter, and an interpreted one can call a compiled regex.
The RegExp JIT generator (RE_Native) walks the bytecode and emits machine checks for characters, classes, boundaries, loops and captures. It has a separate part for each architecture: IA-32 (used by both the x86 and the x64 builds), ARM and MIPS.
When to compile is decided by a simple rule. A literal—/.../ right in the program text—goes through an attempt at machine-code compilation immediately, while the script is still being parsed. An object created with new RegExp() starts in the interpreter and moves to the JIT only if it has been called more than 10 times or fed a string longer than 256 characters. The logic is clear: a literal is almost certain to run many times, while a dynamically assembled expression may turn out to be single-use, and compiling it would just waste time.
The JIT does not cover everything. Complex nested backtracking, backreferences and some position checks are marked as “unsupported” and from then on are always run by the interpreter.
How much does the JIT speed up a search? Less than you might think. On a modern machine (Core Ultra 9 275HX, averaged over 10 runs), searching for a date with /[0-9]{4}-[0-9]{2}-[0-9]{2}/ takes 194 ms in the interpreter and 172 ms with the JIT, and searching for a log entry with /ERROR [A-Z]+ [0-9]+/ takes 187 and 163 ms. A speedup of 1.13–1.15×. Not dramatic, but it comes for free.
Now step by step: what happens when exec is called. Take the same expression and the string "xabcd".
Step one: find where it is worth starting at all. Running a full match from every position in the string is expensive, so at compile time a filter, RE_Searcher, is built: a 256-entry table marking which characters a match can begin with. Our expression must begin with a, so x is skipped without starting the machine, and the first candidate is position 1. A typical old-Opera technique: cheap filtering first, expensive work after, and only for plausible candidates.
Step two: run the bytecode from that position. The machine keeps track of the current instruction number, the position in the string and a stack of “bookmarks.” Each bookmark says: “if things don’t work out further on, return to instruction N and position P.”
| Instruction | Position | What happens |
|---|---|---|
MATCH_CHARACTER_CS 'a' |
1 → 2 | a matches |
CAPTURE_START 1 |
2 | group start recorded |
PUSH_CHOICE → 5 |
2 | bookmark: “if need be, try bc from position 2” |
MATCH_CHARACTER_CS 'b' |
2 → 3 | b matches |
JUMP → 6 |
3 | jump over the second alternative |
CAPTURE_END 1 |
3 | the group is b |
MATCH_CHARACTER_CS 'd' |
3 | the input has c: failure |
| backtrack | 2 | pop the bookmark, undo the group end |
MATCH_STRING_CS "bc" |
2 → 4 | bc matches |
CAPTURE_END 1 |
4 | the group is bc |
MATCH_CHARACTER_CS 'd' |
4 → 5 | d matches |
SUCCESS |
5 | done |
Step three: return the result. The module hands out only (start, length) pairs for the whole match and for each group: (1, 4) and (2, 2). It is Carakan that turns them into the familiar array ["abcd", "bc"] with index: 1. Had the bookmarks run out with no match found, the filter would have produced the next candidate, and everything would have repeated from the new position.
This way of searching is called backtracking. It has an unpleasant flip side, but it reproduces backreferences, and the order of trying alternatives that JavaScript requires, directly and relatively simply. A classic finite automaton, for instance, cannot do backreferences at all.
The most interesting part of Opera’s implementation is how the bookmark stack is built:
Choice) are kept in a list of their own on the heap. Neither matching nor backtracking involves recursion: a long string with thousands of forks does not turn into thousands of nested function calls and does not bring the process down with a stack overflow. (The parser, the compiler and the JIT analyzer are recursive, but they have depth limits of their own.)/a*ab/ and a string of a thousand as with a b at the end. The greedy a* first eats all the as and then has to be able to give them back one at a time—so, strictly speaking, every a needs a bookmark of its own. Instead of a thousand entries Opera keeps one: where it started, how many alternatives are left and how big a step to take.a* or [0-9]+ are run by dedicated loop instructions (LOOP_CHARACTER_*, LOOP_CLASS) directly in C++, without a pass through the bytecode for every character, and they create a bookmark only if the rest of the pattern might need something given back.Memory for bookmarks and captures is allocated in blocks visible to Carakan’s garbage collector and freed all at once when the search ends—no fussing with a scatter of small objects.
With each new iteration of the standard the demands on the mechanism grew, and new modes appeared one after another:
u (ES2015) — full Unicode support, so that you can use the 💩 emoji in your expressions;y (ES2015) — “sticky” matching;s (ES2018) — dotAll mode, which does away with hacks like [\s\S];d (ES2022) — positions of the matched groups, which the engine now has to be able to return;v (ES2024) — a new Unicode mode extending u: it allows set operations on characters (intersection, subtraction) and strings inside classes, and is stricter about syntax on top of that.The improvements were not limited to new modes: what patterns can express grew considerably too. The most technically demanding addition was lookbehind assertions, introduced in ES2018. Before that, the left-hand context could only be included in the match itself—/\$(\d+)/—and the needed group fished out afterwards. Checking what comes before a match without including it in the result was impossible, but now expressions like these are supported:
/(?<=\$)\d+/ // a number, but only if preceded by $
/(?<!\$)\d+/ // a number not preceded by $
Also in ES2018 came named groups—(?<year>\d{4}), with the result in match.groups.year instead of a faceless match[1]—and Unicode properties: \p{Script=Greek} finds Greek characters, and \p{Letter} any letter of any alphabet, with no manual listing of ranges. ES2025 allowed flags to be switched on for only part of an expression (in (?i:hello) world case is ignored only in the first word) and the same group name to be repeated in different branches of an alternation.
JS itself also made working with the object easier. Built-in escaping (RegExp.escape) arrived, and so did matchAll, which lets you iterate over every match along with its groups, and the new RegExp() constructor became less strict: you can pass it an existing regex with new flags—new RegExp(/abc/g, "i")—where that used to end in a TypeError. Evolutionary improvements that make a developer’s life much easier but make the implementation noticeably harder—if not so much the regex engine itself as the built-in functions around it.
Bringing Carakan up to ES2015 demanded substantial rework, and the biggest piece of it was the u flag.
Strings in JavaScript are UTF-16: each Unicode code point takes one or two 16-bit “code units.” Anything that didn’t fit into the first 65,536 values, our 💩 included, is written as a surrogate pair, and "💩".length is 2. The old regex engine knew nothing about pairs: to it 💩 is two independent 16-bit code units, and /^.$/.test("💩") returns false. With the u flag the machine is obliged to read whole code points: assemble the pair while reading and match it as a single unit.
Sounds simple, but forget that word when you work with Unicode, because:
u mode goes by the Unicode tables, while without u it uses the old algorithm, which for compatibility’s sake forbids non-ASCII characters from turning into ASCII: the Kelvin sign \u212A (it looks like K, but it is a different character) matches the Latin k under /iu but not under /i;\b are also computed differently when u and i are combined.Indices in the result, meanwhile, stay in code units: /./u on "💩" gives a match of length 2, because that is how JavaScript measures strings. The old start-position filter and the JIT had to be reliably bypassed.
For u, the existing RegExp JIT was not extended. Unicode patterns were sent to the bytecode interpreter, since the machine-code executor worked with 16-bit elements and could not represent code points correctly. For old patterns the JIT was used as before.
The y flag already existed in Opera as an internal “don’t search further” mode; all that was left was to make it standard, with the whole lastIndex contract.
Next came \p{...} from ES2018, and it needed data the old module did not have at all: which characters are letters, which are digits, which script each character belongs to. Unicode 17.0 tables had to be brought in: property names and their aliases, general categories, scripts, binary properties of characters and of strings. At compile time \p{Script=Greek} turns into an ordinary set of RE_Class ranges, and during the search it is a class just like [a-z], only bigger.
ICU is not used here: unlike Intl from the previous part, regexes and identifier parsing run on tables of their own and work even in a build without Intl. The price is size: the bulk of the code added to the module is not logic but those very tables. Their source is the Unicode Character Database, the consortium’s official set of text files: UnicodeData.txt, lists of properties and their aliases, scripts, emoji data—eleven files in all. Their version and checksums are pinned in a manifest, the import verifies SHA-256 and downloads nothing unless explicitly asked, and a Python generator deterministically rebuilds all the shared Unicode tables at once. Moving to the next Unicode version will take a new manifest and a regeneration of the whole set.
The resulting data file for RegExp is 9,096 lines and about 700 KiB of source text: 883 property names and aliases, 368 sets with 22,127 ranges and 3,953 emoji sequences. In the binary that comes to 287 KiB of packed data as calculated before alignment, while the measured growth of Opera.dll on Windows x64 relative to the baseline build is 284 KiB.
Some of the new features came almost for free—thanks to the way the old module was built.
d flag. The machine always returned (start, length) pairs anyway; Carakan simply kept only the substrings. Now, with d, it additionally builds indices: /b(c)/d.exec("abcd").indices is [[1, 3], [2, 3]]. The search algorithm did not change at all.s flag. The dot instruction used to take no arguments and always rejected a line break. Now it has a dotAll flag that skips that check.(?i:...) needs a stack of flags at run time—but no. The mode is already known at compile time: on entering the group the compiler changes its current settings, on leaving it restores them, and every instruction inside immediately gets the right variant: CI instead of CS, a dot with or without the dotAll flag. The machine knows nothing about modifiers at all.(?P<name>...) could already store names, and to the executor a named group is an ordinary numbered group. What was new was the standard syntax, groups in the result and $<name> in replacements.(?<year>\d{4})-\d{2}|\d{2}-(?<year>\d{4}), and the compiler checks this by tracking the path through nested alternatives. Same-named groups are linked into a chain, and groups.year gets whichever one actually took part in the match. \k<year> needed just one new instruction, MATCH_NAMED_CAPTURE, which looks up the group in the chain that matched.RegExp.escape. Doesn’t touch the machine at all: it is an ordinary function that turns a string into a safe piece of pattern. A fun detail: it always encodes a leading digit or Latin letter. RegExp.escape("2026") returns \x32026—an encoded two (\x takes exactly two digits) and the untouched tail 026. Otherwise, once glued into a pattern, the two could stick to a \1 standing before it and turn it into \12.The whole machine is built on moving forward: every instruction that consumes text moves the position to the right (the machine goes back only by backtracking), the filter looks at the first character, bookmarks remember where to return to if things don’t work out ahead. Lookbehind requires looking to the left.
The simple approach—”step back N characters and check”—works in many engines (Python’s re module, for example), where the contents of a lookbehind must have a fixed length. JavaScript grants no such discount: inside there can be \d+, alternatives of different lengths, groups. The standard describes it this way: the contents of a lookbehind are matched right to left. You can even see it in the result:
/^(\d+)(\d+)/.exec("1053") // ["1053", "105", "3"]
/(?<=(\d+)(\d+))$/.exec("1053") // ["", "1", "053"]
In the first case the machine goes left to right, the first group reaches the digits first, and it is the greedy one. In the second the machine goes right to left, the second group comes first—and grabs everything it can, leaving the first group a single digit.
The regex engine gradually learned to walk in both directions:
SET_DIRECTION, switches the reading direction.(?<=ab) reads as “first b, then a, moving left.”MATCH_STRING_CS compares the literal with what stands immediately before the current position and, on success, moves the position left. For u a matching mechanism was added that recognizes a surrogate pair from its tail end.One thing did have to be switched off: the start-position filter. It guesses candidates only from the characters the pattern consumes going forward, and knows nothing about “look-back” context. So an expression with a lookbehind anywhere in it is checked from every position in the string. The exception is the y flag: with it the search starts only at lastIndex anyway.
Before v, a character class was a set of individual characters: ask about one character, get “yes” or “no.” The v flag turns classes into a small algebra of sets:
/[\p{Letter}&&\p{Script=Greek}]/v // intersection: Greek letters only
/[\p{Letter}--\p{ASCII}]/v // subtraction: all letters except ASCII Latin
/[\q{ab|a|xyz}]/v // a class of strings
/^\p{RGI_Emoji_Flag_Sequence}$/v // a property of strings: country flags
The last two are fundamentally new. The 🇺🇦 flag is not one character but two “regional indicators” in a row, that is, four code units. The class stops answering the question “does this character fit” and starts answering “which of several strings starts here.” And if several strings of different lengths fit, which one should it take?
The standard says: the longest first. For this, RE_Class gained RE_UnicodeSet, a small tree of operations whose nodes can be a character set, a property of strings, an explicit list of strings, a union, an intersection, a subtraction or a complement. When checking, the class finds the longest matching string and puts the shorter options into bookmarks—the very same ones used for ordinary alternatives:
/^[\q{ab|a}]b$/v.exec("ab") // ["ab"]
First the class takes ab, but the pattern then expects b, and the string has ended. Backtrack—and the class tries a, after which b matches. No separate engine for string classes was needed.
Syntax in v mode is stricter: a number of characters inside a class ((, ), {, }, /, -, | and others) must be escaped, && and -- have become operators, and other doubled punctuation like !! or ## is reserved for the future.
All of this is parsed by a separate recursive parser written for v from scratch. The leaves of the tree—ranges and character sets—are still stored in the same old RE_Class. Nesting depth is limited to 128 levels. The standard doesn’t require this; it is protection against stack overflow from recursion, and the number itself is inherited from those very class intersections of Opera 12.15, where the same limit capped nested &&[...].
Everything new that changes matching itself is implemented in the bytecode and its interpreter. Implementing the same features in the machine-code generator was not a goal: it was taught to recognize the new instructions and modes and to decline (for now) to compile them.
| What’s in the expression | JIT |
|---|---|
| Old ES5-era patterns | yes, as before |
The d flag |
yes: indices are built after the search |
| Named groups, including duplicate names | yes: to the machine they are numbers |
u, v and everything that comes with them: \p{...}, string classes |
no |
s and local modifiers |
no |
| Lookbehind | no |
Named backreference \k<name> |
no |
“Yes” here means “can,” provided the expression contains nothing from the lower half of the table or any of the other constructs the generator already declined before.
A conservative refusal is better than a partial machine-code implementation that might occasionally produce a wrong result. Old expressions kept their fast path, new ones got correct semantics in the interpreter. But this is an obvious debt, too: code that makes heavy use of \p{...}, v or lookbehind gets nothing from the RegExp JIT.
s is the most galling: for machine code it would seem to be just one skipped check. And there really is no catch—the generator needs to learn to read the new argument of the dot instruction, not to emit the line-break check, and to pass testing on IA-32, ARM and MIPS. It wasn’t done simply because the goal was correctness, and a safe fallback already existed. Modifiers are harder: the old machine-code analyzer assumes the i, m and s modes are uniform for the whole expression, while here they can change several times within a single pattern. Not impossible, just a separate job on every architecture.
Backtracking has a flip side. The textbook example is /^(a+)+$/ on a string of as with an exclamation mark at the end. There is no match, but to prove it, a naive engine will try every way of dividing the as between the inner and the outer +, and each new character doubles the work.
Carakan won’t take that bait so easily. The inner a+ runs as a dedicated loop instruction, and the same-kind bookmarks of the outer + are packed into one—the very compression mentioned above. As a result, thirty as are checked in under a millisecond. But change the example just slightly and the trap springs:
/^(a|aa)*$/.test("aaaaaaaaaaaaaaaaaaaaaaaaaaaaaa!")
Here the number of paths grows like the Fibonacci numbers. On the same machine a string of 30 as is checked in 59 ms, 32 in 158, 34 in 410, and 36 in over a second: every two extra characters mean roughly 2.6 times the work. This is called ReDoS, regular expression denial of service, and it is quite real: in 2019 a single unfortunate expression took down a large part of Cloudflare for 27 minutes.
Opera’s safeguards help here only in part. Its own bookmark stack on the heap won’t let it crash with a system stack overflow, bookmark compression saves memory, and the time-slice check every 65,535 backtracks won’t let a search hold one execution slice forever. But there is no limit as such: having used up its slice, the context yields control to the scheduler and carries on with the same work in the next slice. No dialog, no exception, no cancellation—the search can go on as long as it likes.
There are various ways to defend against this: an explicit backtracking budget, a deadline, rejecting dangerous patterns, or a second executor with guaranteed linear time (a Thompson automaton, for instance, as in RE2) for expressions without backreferences and other constructs that genuinely need backtracking.
V8, for example, as of September 2026 still keeps such an executor behind flags as an experimental feature. When enabled, it takes over after 50,000 backtracks, and not for every pattern: backreferences, u/v, i and some quantifiers are off-limits to it, and lookaround, supported only in late 2024, requires yet another flag. Carakan has nothing of the kind. Specific nested quantifiers like ^(a+)+$ are defused by the old optimizations, but the vulnerability class itself hasn’t gone anywhere since 2013.
WeakRef and FinalizationRegistry appeared in ES2021, and they are probably the most unusual features on the whole list. Everything else—syntax, methods, flags—is a language feature. These two belong to the garbage collector.
I have already said a little about GC. To recap the main point: an object is alive as long as it can be reached through a chain of references from the “roots”—global variables, the call stack, the engine’s internal lists. Anything that cannot be reached, the collector is free to remove.
Usually that is exactly what you want. But picture a cache: say, images the page once loaded and decoded. Put them in an ordinary Map keyed by address, and the cache will hold every image forever, even long after it was thrown off the page. A classic memory leak dressed up as an optimization.
If the key were the image itself, the WeakMap from ES2015 would do: an entry in it lives as long as its key does. But here the key is a string with an address, and what needs holding weakly is the value. That is what a weak reference is for: “I know where the object is, but I won’t hold on to it.” If no other, strong, references to the object remain, the collector will remove it, and the weak reference simply goes empty. The cache becomes a Map from addresses to WeakRefs, and the emptied entries need clearing out from time to time—which is where FinalizationRegistry comes in handy, albeit with caveats covered below.
| Link | What stays alive |
|---|---|
ordinary reference A → B |
B, as long as A is reachable |
WeakRef → object |
the object, only if there is another strong path to it |
WeakMap: key → value |
the value, while the table and the key are reachable, the key on its own |
FinalizationRegistry: target → value |
the target is not held; the value is held by the registry |
“On its own” in the WeakMap row matters: if the key is reachable only through its own value, that loop saves neither the key nor the value.
You can’t polyfill this. Only the collector decides when an object dies, because only it sees the whole reference graph, and no reference counter or clever object property can stand in for it.
WeakRef Doeslet object = { data: "cached" };
const reference = new WeakRef(object);
object = null;
const value = reference.deref();
if (value !== undefined)
use(value);
deref() returns the object if it is still alive, and undefined if the collector has already removed it. And object = null by itself removes nothing: collection may start in a millisecond, in a minute or not at all. Two identical runs of a program are not obliged to see the reference go empty at the same point. That is the contract: a weak reference tells you “maybe, someday.”
So when does that “someday” arrive? Carakan triggers collection primarily by the amount of memory allocated, including the memory that external objects count as their own. After each collection the threshold is recalculated from the volume of surviving data: in the desktop configuration the next collection starts when the heap grows to about four times that volume, but not before 512 KiB and not after one more gigabyte of growth. On top of that, once a second a separate mechanism looks after idle heaps and collects those where at least ten seconds have passed since the last collection and something has happened in the meantime.
You can refer weakly to an object or to a symbol—but not to one that lives in the global Symbol.for registry. Such a symbol can always be obtained again from its string key; it cannot be “lost,” and a weak reference to it would be a lie. Built-in symbols like Symbol.iterator also live forever, but the standard does allow holding them weakly: the line is drawn precisely at the registry. Symbols, incidentally, were not allowed right away: in ES2021 only objects qualified, and unregistered symbols were added in ES2023, together with the ability to use them as WeakMap keys. Numbers and strings qualify even less: they have no “identity” to lose.
deref() comes with a non-obvious guarantee. Look at this code:
if (reference.deref() !== undefined)
reference.deref().render();
If the collector could run between the two calls, the second deref() would return undefined, and the code would crash out of nowhere. So the standard promises: an object that deref() has returned even once will survive until the end of the current job—the synchronous chunk of code together with the chain of Promises launched right after it. This covers only the Promise reactions that run immediately afterwards: a Promise that settles later—on a timer or a network response—belongs to the next job and does not extend the object’s life. The same rule applies when a WeakRef is created: the object just passed in won’t vanish from under your hands.
In Carakan there is a weak_kept_alive list in ES_Runtime for this. The WeakRef constructor and every successful deref() add the object there, and the list itself is an ordinary root for the collector: everything in it counts as reachable. The scheduler clears it, and only once the Promise queue has fully drained.
Hence a consequence: showing an emptied reference live is not so easy. jsshell has a hidden __collectGarbage__() function (Test262 gets the same one under the name $262.gc()), but the straightforward “create a WeakRef, null the variable, run a collection” will never show undefined: the object is obliged to survive until the end of the job. You have to wait for the next one. In a shell without timers the easiest way to do that is through a FinalizationRegistry cleanup task:
var reference;
var trigger = new FinalizationRegistry(function () {
var pressure = [];
for (var i = 0; i < 10000; ++i)
pressure.push({});
__collectGarbage__();
print(String(reference.deref()));
});
Promise.resolve().then(function () {
var target = { value: 42 };
reference = new WeakRef(target);
print(reference.deref().value);
trigger.register({}, 0);
__collectGarbage__();
});
42
undefined
The first collection removes the anonymous object registered in trigger and queues its callback. Before calling the callback the shell finishes the Promise chain and clears weak_kept_alive, and then, inside the callback, a series of allocations runs and the next collection removes the original target. The extra allocations are there only to make the example reproducible; they have nothing to do with the WeakRef contract.
Under the hood, both WeakRef and every registration in FinalizationRegistry are built the same way: an ES_Weak_Cell object with five fields:
| Field | Reference | What it holds |
|---|---|---|
target |
weak | the target being watched: an object or a symbol |
held_value |
strong | the value for the callback |
unregister_token |
weak | the unregister token |
owner |
strong | the registry’s internal state |
next |
strong | the next registration of the same registry |
WeakRef uses only target; its other fields are empty. The registry itself has two layers: the visible JavaScript object and an internal state referenced both by it and by a hidden service cleanup function. The state holds the callback, the head of a singly linked list of cells, the global object of the right window, that same cleanup function and fields for the queue.
A new object needs a new type tag in the GC header, and that, as you may remember, is a problem: six bits, 64 values, none free. The solution is familiar from BigInt: take an existing tag (GCTAG_ES_Boxed_List) and add a MASK_IS_WEAK_CELL flag to the header. This does not get in the way of walking the heap: each object’s size is recorded in the header itself and is not derived from the tag. And during marking the collector sees the flag and applies a special rule to the cell: it follows held_value, owner and next, but leaves target and unregister_token alone.
Besides BigInt and the weak cell, the same “tag plus flag” trick hosts the internal class element, the module binding and built-in strings. There are no free tags left at all: 0 is a free memory block, 1 through 62 are real kinds of objects and service structures, and 63 is “not yet initialized.”
FinalizationRegistry DoesIf WeakRef answers the question “is the object still alive,” FinalizationRegistry lets you find out that it has died:
const registry = new FinalizationRegistry(function (resourceId) {
removeCachedResource(resourceId);
});
const token = {};
let wrapper = makeWrapper();
registry.register(wrapper, 42, token);
// If the registration is no longer needed:
registry.unregister(token);
register takes three arguments: the target being watched (an object or an unregistered symbol, held weakly by the registry), the value the callback will receive (held strongly) and an optional unregister token (weakly again, and again an object or a symbol). One token can cancel several registrations at once, and unregister reports whether it found anything.
Why the value for the callback is held strongly is clear: by the time of the call there would otherwise be nothing to pass. But that is exactly where the main trap comes from:
registry.register(wrapper, wrapper); // TypeError: the standard catches this
registry.register(wrapper, { owner: wrapper }); // but not this
In the second case the held value refers to the object itself, the registry holds the value strongly—and the object won’t die while the registry itself is alive. No callback will come, no error either, just a quiet leak.
The registry itself also has to be needed by someone. If the program has lost both the object and the registry, the collector is free to remove the whole construction without calling any callbacks.
And most importantly: FinalizationRegistry is not a destructor. Collection may happen much later, the process may exit without one at all, and the browser is not obliged to run callbacks before a page closes. Closing files, sockets or transactions this way is not an option. Tidying a cache, gathering diagnostics, freeing a resource “just in case” when its main release path has somehow failed—that is fine.
Back to the resource cache example. Since the callback is not guaranteed, cleanup through the registry is only “best effort,” and empty entries in the Map can still pile up. There is also a subtler trap: a race between generations. If a new image already sits at the same address, a belated callback from the old one must not delete it. So the registry is given the address together with the weak reference itself, and before deleting, the callback checks that it is exactly that reference sitting in the cache:
const cache = new Map();
const cleanup = new FinalizationRegistry(function ({ url, ref }) {
if (cache.get(url) === ref)
cache.delete(url);
});
function remember(url, image) {
const ref = new WeakRef(image);
cache.set(url, ref);
cleanup.register(image, { url, ref });
}
The held value here refers to the WeakRef, not to the image itself, so the leak from the previous example won’t happen. Another option is to cancel the old registration through the unregister token when replacing an entry.
The callback cannot be called straight from the collector. User code can allocate memory (and trigger a nested collection in the middle of the current one), throw an exception, register new objects or unregister old ones—all of it in the middle of the process that is putting memory in order. So GC does the minimum: it clears target and queues the registry.
Even that queuing allocates no new memory: the link field to the next registry is set aside in advance inside the registry itself. That matters because the queue fills up while GC is running, when its internal structures are in an intermediate state, and an allocation could trigger a recursive GC run inside an unfinished collection cycle. OOM is the best you could hope for, but there is simply nowhere to return an exception to the usual way. In effect, this code is obliged to have no point of failure.
Then the scheduler steps in. Here is what happens to a registered object, step by step:
| Moment | What happens |
|---|---|
| script | registry.register(wrapper, 42), then wrapper = null |
| garbage collection, between mark and sweep | wrapper is not marked: target is cleared, the registry is queued |
| sweep | the memory of wrapper is freed |
| the current job and the Promise chain | carry on as if nothing had happened |
| a separate scheduler task | the service function finds the empty cell and calls callback(42) |
The service function acts carefully. It finds a cell with an empty target, takes the value out of it, pins it, removes the cell from the list—and only then calls the user callback. After it returns, the list is read afresh. That way the callback can register, unregister and even trigger garbage collection as much as it likes: the traversal will never be left holding a pointer to an already removed cell.
If the callback throws, its value is lost—the cell was removed before the call. But the other cells stay in place, and once the task finishes, the registry will be queued again. One failed callback doesn’t take the rest down with it.
The standard does not define the order of calls. Carakan walks the list from the head, but you must not rely on that—not on the order of registration, not on the order in which objects die, not on how many values get processed per task.
Web Workers appeared in browsers long ago, and Opera 12 had them. But they talked to the main page only through messages: data was either copied or transferred wholesale, with the sender losing its buffer in the process. For exchanging a couple of strings that is no problem, but the moment you want to process video frames, compute physics or run a game compiled to JavaScript across several threads, constant copying becomes a bottleneck.
SharedArrayBuffer (ES2017) solves this radically: several agents get one and the same block of bytes and can read and write it simultaneously. An “agent” is what the standard calls an independently executing context with its own heap and job queue—the main page or a worker. And the Atomics object provides the operations without which such shared access turns into a lottery.
The beauty of it is that no specialized structure dictating rigid usage patterns is handed “outward.” Each agent gets its own SharedArrayBuffer object, its own Int32Array, prototypes, global object and garbage collector, while underneath they all look at one block of memory:
agent A, heap A agent B, heap B
SharedArrayBuffer A SharedArrayBuffer B
│ │
Int32Array A Int32Array B
│ │
└──────── ES_ArrayBufferBackingStore ───┘
│
shared process bytes
A and B are different objects to their consumers, with different identities: on transfer the receiver creates its own wrapper. You can’t compare them directly—an object from agent A’s heap is simply inaccessible to agent B—but if you send the buffer from A to B and back, agent A ends up with two wrappers of the same block, and === between them gives false. Meanwhile, a write through Int32Array A is immediately visible through Int32Array B.
In Carakan it couldn’t be any other way. The garbage collector of one heap knows nothing about the roots and collection phases of another, and the moment you pass a pointer to an ordinary engine object to a neighboring thread, the two collectors start stepping on each other’s toes. So nothing crosses an agent boundary—no objects, no strings, no exceptions, no Promise jobs—only a reference to a block of bytes and copied numbers. On the plus side, the whole engine did not have to be made thread-safe: object classes, strings and the garbage collector stayed single-threaded, as they always were.
ArrayBuffer to Backing StoreIn the old Carakan the bytes of an ArrayBuffer simply belonged to the object itself: object alive, bytes alive. That won’t do for shared memory, so the layer is split in two. The JavaScript wrapper still lives in the collector’s heap, while the bytes have moved to a separate control block, the backing store. It holds:
Each wrapper holds one reference to the backing store, and the bytes are freed only when the last wrapper, the last temporary user or the last sleeping agent goes away. The counter is atomic, so different threads can release references at the same time without getting in each other’s way.
There was a subtlety, too: the JIT accessed the fields of the old structure at hard-coded offsets. So the order of the old fields was preserved, and the pointer to the new control block was appended after them. An ordinary ArrayBuffer stayed just as fast, but every one of them now pays for the control block: in the browser build that is 80 bytes in the 64-bit version and 44 bytes in the 32-bit one. In jsshell the block is bigger—the waiter queue fields are added there, and on x64 it comes to 112 bytes.
A shared buffer can’t do everything an ordinary one can. It can’t be “transferred” with detachment like an ordinary ArrayBuffer: the other agents have exactly the same rights to the same bytes. When it is sent to another agent, a new wrapper is created and the reference count goes up. It can’t be saved to a file or sent to another process: the block’s address only means something inside the current process. And slice() creates a new, independent buffer with a copy of the bytes.
Nor does the browser’s own code get a raw pointer to shared bytes. The DOM, codecs and WebGL can “borrow” the memory of an ordinary ArrayBuffer, but for a shared buffer that is forbidden; otherwise outside code would read and write that memory bypassing the atomic operations.
A bit of archaeology: the Presto code contains OpSharedMemory, an interface for named shared memory for exchange between processes, with implementations for POSIX and Windows. The code is utterly dead and used nowhere; whether it was yet another unfinished prototype or a forgotten vestige, we will never know for sure. I lean toward the former. It would have been no use for SharedArrayBuffer anyway: the interface can’t reserve addresses or grow, and it would also have needed a separate lifetime object on the GC side.
Recall OpPseudoThread: inside a single system thread, “virtual” threads are created that run in turn.
This is cooperative multitasking, and in theory it would do for implementing shared memory. In a sense it would even be simpler: no races, no data visibility problems. User programs would see no difference, though they would run slower.
But such an implementation would distort the whole point, and so it won’t do. Nor is it all that simple: the scheduler would need a lot of work before algorithms designed for true parallelism could run on it in “one-at-a-time” mode.
The other option is to use real system threads and do everything properly. But here is the snag: Opera is single-threaded, and its workers live on the main thread under an internal scheduler anyway. Today’s consumers would gain nothing from a truly multithreaded mechanism. For shared memory to work in the browser, workers have to become real agents in their own system threads, with a full shutdown protocol: closing a worker or a document has to wake the sleeping agents, cancel asynchronous waits, wait for the threads and free the memory.
That is a big and difficult task in its own right, and it is better put off until multithreading support is added to the browser as a whole.
But there is another reason not to hurry: a shared buffer plus a worker incrementing a counter in a loop is a homemade, very high-precision timer, and that is exactly the kind of timer side-channel attacks like Spectre need.
In January 2018 all the major browsers disabled SharedArrayBuffer at once, and in the end brought it back only for pages in “cross-origin isolation,” which a site switches on with the COOP and COEP headers. Opera 12 knows nothing about such isolation, and it would have to be built from scratch: policies, the crossOriginIsolated flag, rules for transfer between origins.
But we had come too far to just stop. Another compromise: implement and test a full multithreaded SharedArrayBuffer in jsshell, but not make the feature available in the browser.
In jsshell each agent starts in its own OS thread and creates a full Carakan instance there: its own runtime, global object, heap, job queues and JIT. The global state of jsshell had to be sorted out by owner, making the g_opera root and the stack thread-local. Immutable tables and shell parameters fixed before the agents start remain shared, and the executable memory heap for the JIT on Windows got synchronization.
So: the feature works for tests only. The Test262 interface $262.agent can start agents, hand them a shared buffer, collect reports from them, and at shutdown wake the sleeping agents, wait for all the threads and free the memory. Full Web Workers will need work in the browser itself.
Atomics IsTake the simplest counter in shared memory and two agents simultaneously executing:
counter[0] = counter[0] + 1;
That is three separate actions: read, add and write. Both agents can read the same old value, both add one, and both write the same result. One increment is lost. This is a classic race: the result depends on how the operations of different threads happened to interleave in real time.
The usual remedy for such races is a lock, also known as a mutex. In Python a critical section is guarded by threading.Lock, in C++ by std::mutex, and at the operating system level mutexes and other waiting primitives are used. While one thread holds the lock, the others do not enter the protected section. It is a universal mechanism: under one lock you can change several fields consistently and check arbitrary conditions. The price is the cost of waiting, of putting threads to sleep and waking them, and, when something goes wrong, deadlock, with two threads waiting for each other forever.
For a single integer such protection is overkill. The processor can read a cell, change it and write it back as one indivisible operation, and exactly those operations are collected in the Atomics object:
Atomics.add(counter, 0, 1);
Now nobody can wedge in between the read and the write, and on x86 this call becomes a single lock xadd instruction. For counters, flags and other state that fits in one cell, this is simpler and cheaper than a mutex. For anything more complex—read two fields and write a third—a lock is still needed, just one built on top of atomics.
Carakan implements the whole set: arithmetic and logic (add, sub, and, or, xor), exchange, compareExchange, load, store, isLockFree, plus wait, waitAsync and notify. The key one is compareExchange: it writes the new value only if the cell still holds the expected old one, and both locks and lock-free data structures are built on it. Atomics work only with integer typed arrays, including the 64-bit BigInt ones.
The visibility rules are taken from the C++11 memory model. All Atomics operations run in the strictest mode, sequentially consistent: all agents see them in one and the same order. That is precisely what makes it possible to pass data between threads:
// producer
data[0] = 42;
Atomics.store(flag, 0, 1);
Atomics.notify(flag, 0);
// consumer
Atomics.wait(flag, 0, 0); // sleep while the flag is 0
print(data[0]); // guaranteed to be 42
wait and notify are modeled on Linux futex: an agent goes to sleep only if the cell still holds the expected value, and another agent, having changed it, wakes those sleeping on that address. Carakan doesn’t call futex itself—the engine has to work the same on Windows and POSIX—so it has a waiter queue of its own.
Underneath, all of this sits on a small layer, op_atomics.h, which is essentially a little <atomic> of its own on top of compiler intrinsics: __atomic_* in GCC and Clang, _Interlocked* in MSVC.
It is not only Atomics that goes through this layer, but also the most ordinary view[0] = 1. If one thread writes an address with a plain C++ assignment while another reads it at the same moment, the race happens in C++ itself, and the behavior of the whole engine becomes undefined. So for a shared buffer even ordinary accesses are atomic, only in the weakest mode, relaxed: a value that is read will always be entirely either the old one or the new one, never a mix of bytes from both, but ordering between neighboring cells is not guaranteed. For the same reason set, copyWithin, slice and the typed array constructors cannot simply call memcpy on shared memory.
A separate trap is user code in the middle of an operation:
Atomics.store(view, index, { valueOf() { buffer.grow(8192); return 1; } });
While the engine converts the argument to a number, valueOf() manages to grow the buffer, trigger garbage collection or throw an exception. So after all conversions the bounds are checked again: a pointer and a length obtained before the call may already be stale.
The JIT does not yet speed up either accesses to shared typed arrays or Atomics calls—everything goes down the slow but safe interpreter path. To fix that, the generator has to be taught to tell shared memory from ordinary memory, pick the width and ordering of an operation and keep track of the length of a growable buffer. Like other JIT improvements, this one has been put on the back burner.
The results of the final stage of work were visible right away. Many sites that simply hadn’t worked before started working—slowly, creakily, with glitches in the layout, but working. Not all of them, and not all at once: the very first live tests turned up plenty of small bugs. The Vue and React frameworks, before doing anything at all, checked for certain APIs missing from Opera, and those had to be added along the way. A lot still doesn’t work, but Carakan has nevertheless reached the stage where each new fix opens up a huge part of the previously broken web to the browser.
Before leaving this chapter, it is worth looking back: what else had to be done, what the engine has turned into and what lies ahead for it.
Here are a few more significant changes, each of which deserves a chapter of its own. For now I only have time for a brief overview.
Classes gained fields, private fields and methods (#name), static initialization blocks and the #name in object check. A private name is an unforgeable key that exists only inside the class body. Its value can be kept in the same property storage as ordinary fields, but neither property enumeration, nor Proxy, nor reflection may see it. And field initialization is tied to a precise moment in construction: in a derived class this doesn’t exist until super() is called, and the fields appear right after it.
In Carakan a private name is a separate kind of key in the same property storage, not a parallel table of private fields. Every evaluation of a class creates a fresh key: internally it resembles a symbol, but no user Symbol can obtain it. Ordinary lookup, enumeration, descriptors, cloning and Proxy simply filter such keys out. There is no separate “brand” either: an instance’s brand is the mere presence of its own element with that key. That is why #name in object looks only at own properties, without prototypes or Proxy traps, and the “freshness” of the key shows nicely in a class factory:
function make() {
return class {
#x = 1;
static has(o) { return #x in o; }
};
}
const A = make(), B = make();
A.has(new A()); // true
A.has(new B()); // false: each evaluation of the class has its own #x
Asynchrony grew from Promise to async iterators and generators, for await...of, Array.fromAsync, iterator helpers and new Promise combinators. The main difficulty here is not the number of methods but the order: one extra or one missing microtask changes the sequence of calls, and Test262 checks for that. The standard itself changed in this area too: in ES2019, for example, await stopped spending extra microtasks on wrapping something that is already a Promise. So the job queue, its owner and the microtask checkpoint in the browser became part of the language semantics rather than mere scheduling infrastructure.
Modules gained dynamic import(), import.meta, import attributes, JSON modules and top-level await. The last one turns the module graph into an asynchronous state machine: a parent can wait for a dependency, a cyclic group of modules has to finish consistently, and the failure of one module has to reach everyone who imports it. Carakan’s resumable execution of module bodies proved a suitable foundation, but the browser loader, CSP, the network and the document lifecycle still had to be tied to the state of the language.
Binary data gained more than SharedArrayBuffer: resizable ArrayBuffers, transfer with detachment (transfer), typed arrays with automatically tracked length, BigInt64Array, Float16Array, new DataView methods, base64 and hex encoding right on Uint8Array. The backing store discussed above is a single shared rework that made it possible not to implement each of these features as a special case of its own.
The rest is numerous but architecturally simpler: ?. and ??, ... in objects, new array and Set methods, stable sorting, Object.groupBy and Map.groupBy, Promise.any, Promise.withResolvers and Promise.try, Error.cause, exact preservation of function source text for toString(), Math.sumPrecise, well-formed Unicode strings. Some of this syntax the compiler reduces to already existing bytecode, while the library methods are written as built-in functions in C++ and use the old object model.
Then again, “simpler” doesn’t mean “simple.” Take Math.sumPrecise: it looks like a sum += value loop. But the standard requires first obtaining the exact mathematical sum and rounding it only once—otherwise Math.sumPrecise([1e30, 0.1, -1e30]) returns zero instead of 0.1, and the result starts depending on the order of the addends. The implementation relies on the fact that any finite double is a multiple of 2⁻¹⁰⁷⁴: positive and negative addends are accumulated separately, in two huge integers of 34 64-bit parts each (2,176 bits apiece), then subtracted and rounded once. Plus NaN, both infinities, the sign of zero, a length limit and properly closing the iterator on error. On the outside, one Math method; on the inside, long arithmetic of its own, almost like the one in BigInt.
So “ES2026 support” touches the lexer, scopes, the bytecode, the interpreter, the JIT boundaries, the garbage collector, strings, arrays, typed arrays, the scheduler, the module loader, structured cloning and the browser host—nearly everything the engine has. It is, without exaggeration, the biggest change since the “revival” of Opera began.
Work on this scale often ends unpredictably, even when everything is planned in advance. On top of that, I wanted to keep to the spirit of the old school, designing everything according to the principles built into the architecture.
Luckily, those principles worked in the project’s favor, suggesting solutions more than once. Here and there a seam would show up that was just right for grafting a new feature onto. This is entirely to the credit of the original designers—an experienced engineer always sees the rough paths a system will develop along and leaves room for extending it.
Neither the compiler, nor the garbage collector, nor the JIT had to be rewritten wholesale. The closest thing to a rewrite was ArrayBuffer memory ownership: the object used to own its bytes itself, and now it is a wrapper over a separate backing store with a reference count, storage kinds, address reservation and a waiter queue. The old lifetime model was effectively replaced by a new one, but the external typed array classes and the field offsets baked into the JIT were preserved—so even this was a rewrite of one internal mechanism, not of a whole subsystem.
Still, the upgrade exposed a few places that are asking to be reworked. The type tags in GC headers are exhausted, and new types have to be disguised as existing ones. The state of FinalizationRegistry is an array of fixed slots rather than a self-documenting type. And then there is single-threadedness. The room for extension has run out here: in some places the boundaries need pushing, in others the principles themselves need rethinking. But even after those changes, I’m sure, the engine will remain recognizably Carakan.
And the biggest current problem is the monstrous gap between completeness and speed: semantically the engine has caught up with the standard, but the JIT has not. I am definitely not going to rush here, though: first, native execution needs to be properly tested and optimized. That is a job for more than one week and more than one person, but I intend to come up with something.