About a month ago Carakan got ECMAScript 2015 support. A swarm of AI agents upgraded the engine in a week, which cannot fail to impress, because the difference between ES5 and ES6 is revolutionary! New syntax, an object model, standardized behavior, non-linear execution…
The ES6 → ES2026 upgrade (I am folding every intermediate specification into that) promised to be colossal for other reasons. Without reworking the language itself, the standards added an enormous number of new APIs, changed the behavior of RegExp, introduced new types, described a concurrency model, and the devil knows what else. After all, six years passed between ES5 and ES6, and nearly twice that between ES6 and ES2026.
Work on ECMAScript 2026 began right after ES2015. About a week ago the process moved on from intermediate tests of individual pieces of code to testing the whole engine as an assembled build. That takes a hell of a lot of time and effort, and I would rather not be distracted in the middle of it, so today I will share just one story. It shows a single one of the hundreds of problems that came up; the fix was not the longest, but it was probably among the most technically intricate.
I have mentioned many times that Opera works with the DOM and JavaScript on one thread. There are reasons for that, both historical and practical; I will not dwell on them now. Yet the JS engine still has to handle several tasks that at least look simultaneous—running a different program for each open tab, say.
That is what the OpPseudoThread class in Carakan is for. The meaning matches the name exactly: it is an implementation of a pseudo-thread that can stop and resume without being an operating system thread. By modern standards it is a relative of the coroutine, written back when mechanisms like that were rare in ordinary applied C++.
Every pseudo-thread is given its own region of memory to use as a stack. On it Carakan can descend deep into JavaScript execution, suspend it, and later return to exactly the same place. If the engine needs to call browser code on the ordinary system stack, the pseudo-thread parks its own stack for a moment, switches the processor to the original one, and performs the call there. If the memory it has is not enough, the pseudo-thread can allocate another segment and carry on there.
Architecturally this is elegant and efficient. The switch itself is performed by tiny fragments of machine code. They save the registers, substitute the stack pointer, and return into a different call chain. From the point of view of the calling code the work looks entirely single-threaded, and that code has no idea any switching is going on.
From the point of view of Windows, though, tricks like that look suspicious.
When your code crashes, you can usually work out what went wrong. Dereferenced a null pointer, ran past the end of an array, tried to execute garbage—the debugger gives you an address, a call stack, and at least a rough direction to dig in. There are exceptions: when the problem refuses to reproduce under a debugger, or when you have to debug multithreaded code. Worst of all is when those conditions come together.
One day the Carakan test suite on Windows broke off in the middle of an out-of-memory check. No error, no result—just the program dying with status STATUS_BAD_STACK. Experiments with other build variants produced STATUS_BAD_FUNCTION_TABLE. Either a bad stack or a bad function table in memory; nowhere near enough to understand the cause (though by now, with hindsight, you may be guessing at something). Under a debugger all of it looked… ambiguous—and by reading on, you will see why.
The MSDN descriptions of the two codes have something in common: the words during an unwind. Windows defines STATUS_BAD_STACK as an invalid or unaligned stack encountered during an unwind, and STATUS_BAD_FUNCTION_TABLE as an invalid function table encountered in the same place. So the “out-of-memory check” sets the Windows unwind machinery in motion.
Still not clear!
Once again, recall how exceptions are implemented in Opera: through the TRAP/LEAVE mechanism. It is convenient to think of it as “try to perform an operation” and “return to the nearest handler with an error immediately.” Along the way, Opera keeps a list of temporary objects that have to be deleted on such a return: memory must be released both on success and on an abnormal exit.
At the core of this mechanism are the op_setjmp/op_longjmp macros. On the Windows path they map onto calls to setjmp/longjmp, on the UNIX path onto _setjmp/_longjmp. The first pair is in the ISO C standard, the second only in POSIX. The POSIX version does not unwind frames, simply restoring the saved registers and transferring control to the required point. Linux does not much care what adventures the program went through to get from there to here. Skipping the unwind makes these functions faster, which is apparently what settled their use.
On 64-bit Windows things are different. Before the jump, the runtime unwinds the call chain: it walks the frames, checks them, and restores state from service tables. Ordinary compiled code carries a description of how to do that. The OpPseudoThread switchers, copied into executable memory as byte arrays, have no such tables. Worse still: the chain suddenly jumps from one physical stack to another. That is not at all what Windows expects!
So when an internal LEAVE tried to reach a handler on the other side of a switch, the operating system started unwinding frames, found the break, and stopped the process. STATUS_BAD_STACK meant that the next frame did not belong to the expected stack. STATUS_BAD_FUNCTION_TABLE appeared when the unwind reached generated code for which no suitable description existed.
The hypothesis was confirmed by a minimal test. Gotcha, you little biter!
Although OpPseudoThread had been there all along, the problem had never shown itself in quite this way. Why?
First, we had made fixes to the stack switching subsystem. Before that, Windows considered the system stack to be the active one at all times and returned more cryptic results on a crash. Once the boundaries became correct, the operating system could crash with confidence and a precise error. Which is to say Opera could well have been dying of the same cause long ago, only with different symptoms.
Second, an error like this slipped under the radar of the tests. As already said, the problem does not reproduce on Linux, and on Windows it only showed up in an x64 environment, which is to say from Opera 12.00 onward.
In our case the trigger was the async generators test, which deliberately creates an out-of-memory condition. At that moment Carakan is doing work on its own stack, moves across to the system stack for a while, discovers that it cannot continue, and throws a LEAVE. In single-threaded logic an ordinary internal error should have reached the topmost handler and turned into a tidy test result.
In pseudo-thread logic the LEAVE tries to travel back across the break between stacks, which makes the operating system close the process on the spot.
Technically the problem is not confined to OOM. The same boundary is crossed by Carakan’s calls into the browser environment, by operations that reserve additional stack, by deep recursion, and by a multitude of internal calls. Running out of memory simply turned out to be a convenient and reproducible way to run into an architectural sore spot.
I am not inclined to call this a Windows problem. Both call variants violate the architectural contract. That Linux jumped quietly across stack boundaries for years merely allowed the violation to go unnoticed.
So you have an error on your hands, but you cannot report it “upward.” Something to think about, is it not?
The hint turned up in the same place as the source of the problem. The designers of the mechanism had a premonition—I cannot find a better word for it—of the problem, and left a warning in the OpPseudoThread.h header file:
A pseudo thread's implementation should take be care about using certain
mechanisms. TRAP/LEAVE "inside" a pseudo thread probably works, whereas
leaving out of the thread (across a call to Start() or Resume()) might not
be a good idea. But really, the safest is probably to use Yield() in
exceptional situations, and translate it to LEAVE afterwards. Also, relying
on stack unwinding obviously has the same issues as with TRAP/LEAVE; if
there is any chance that a yielded thread is destroyed rather than resumed
(which seems inevitable) all data referenced from the stack must be possible
to clean up externally. But since available stack space is also fixed (at
thread creation time) conservative usage of the stack is wise disregarding
the unwind issues.
A perfectly accurate prediction of a problem they never ran into, and an equally accurate solution: do not leave a pseudo-thread through LEAVE, do a Yield() instead and convert the error after the return.
That is exactly how the whole path works now. Code running on the other side of a switch is given a local handler on its own physical stack. If a LEAVE occurs, it is caught in the same place where it was raised. What crosses the boundary is not control, but an ordinary numeric error value. Then the regular switcher restores the registers and the stack pointer the way it was designed to. Once execution is back on a continuous system stack, Opera raises the saved error to the outer handler again.
The result is a kind of trampoline. An exception runs up to the boundary, turns into data, bounces across it, and becomes an exception again only on the safe side.
For some transitions even that is not enough. Inside the thread’s stack, between the current point and the handler, there can be machine code generated by Carakan itself, which likewise carries no unwind data for Windows. Determining in advance whether any particular path is safe is impossible: an error in such a check leads to the immediate death of the process. So on a doubtful path the current Carakan activation is abandoned entirely, the pseudo-thread returns to the root system stack in the regular way, and the error is raised from there.
This is a conservative decision. We do not try to continue an internal call after a serious error, but in exchange we do not depend on which generated frames happened to lie between the two points. For OOM and other terminal errors that is exactly what is needed.
Returning control safely is only half the job. On a LEAVE, Opera has to walk the chain of registered objects and release the resources they own. Historically this chain is shared across the logical thread: it can hold, all mixed together, items from the system stack, from Carakan’s main stack, and from the additional segments allocated for deep recursion.
As long as longjmp jumped wherever it was told without asking questions, the construction looked as though it worked. Once the boundaries became explicit, it became necessary to know which physical stack each section of the chain belonged to—so a separate cleanup list had to be added for every segment.
An ordinary LEAVE, however, cleans up list items only as far as the first CleanupCatcher: that one corresponds to the nearest TRAP, removes itself from the chain, and does a longjmp into a frame that is still alive. But when an activation is abandoned entirely, no TRAP on its segments is a destination any more—all those frames will be discarded by the regular Yield() right down to the root. Stopping at the first catcher would leave the objects registered before it, further toward the tail of the list, with no cleanup at all.
Hence a separate CleanupSegment() mode. It runs to the end of the chain for the physical segment being left. For an ordinary item it invokes the virtual cleanup, while a CleanupCatcher it recognizes separately and merely detaches with the base CleanupItem::Cleanup(), without calling the override that does the longjmp. Once the lists are split by segment, the end of such a chain really does mean the boundary of a single stack, so the walk never touches live objects on the system stack or a neighboring one. The same operation is then repeated for every parked outer block.
Phew! It is rarely enough to find a problem and a solution for it. The fix itself can uncover problems that are incredibly hard to foresee. I am sure I will run into this sort of thing more than once.
OpPseudoThread is a magnificent mechanism for its task (one that now also works the way it should). But could it be replaced with real system threads?
Yes! Firmly and unequivocally!
I am so sure of it because the code already holds a ready-made solution, OPPSEUDOTHREAD_THREADED. It is an implementation of OpPseudoThread running on top of pthreads and Win32. Start() creates a worker on a separate thread, after which the calling side and the worker hand the right to execute back and forth through a mutex and a condition variable. Yield() wakes the original side and puts the worker to sleep; Resume() does the reverse; Suspend() asks the original thread to run a callback; if Reserve() finds itself short of space, it creates another nested system thread with a new stack. A finished thread is even kept in a global cache for reuse. It bears some resemblance to the OpWorkerThread we built earlier, and I will certainly come back to that.
This variant does not, of course, execute JavaScript in parallel. It merely provides an alternative mechanism for executing it. The comment on TWEAK_OPPSEUDOTHREAD_THREADED describes it as a development and debugging aid, convenient for profilers but too heavy for production: a system stack, a kernel object, and a scheduler switch cost noticeably more than saving a few registers.
But could a mechanism like that (after some work) be used for true parallelism?
No, and the choice of mechanism is not the reason.
Carakan, the DOM, the garbage collector, and the surrounding browser code are all built around cooperative ownership of state. Any genuinely parallel execution would mean revisiting heap ownership, DOM access, thread-local state, stopping the GC, passing results between schedulers, and solving problems next to which the one described above would look like a light warm-up.
That is all for today.
Next comes debugging the new Carakan, which by my reckoning will take an infinite amount of time. Telling the story of what was implemented and how may take roughly as long. I expect the next part will take longer than usual to write, and it will most likely be a sub-series of its own: “We Took Carakan; Losses Were Heavy.”