The Engine We Lost: Digital Archaeology of Opera Presto

The Engine That Can Google Again

Contents

XXXII. Fetch for Opera Presto

Take the last official release of Opera Presto and try to google something in it. As basic an operation as they come—you would think it could work in browsers far older than this one. Yet even if you fight your way past the security warnings and spoof the user agent, all you get is an error message. Behind that message sits a requirement: Fetch.

Some history first.
Before Fetch, web developers used XMLHttpRequest, or XHR. The interface was invented at Microsoft, the other browsers picked it up, and the whole AJAX era grew up around it: pages learned to talk to the server and refresh their own contents without a full reload. The XHR specification carries an Opera fingerprint—its editor was Anne van Kesteren of Opera Software, who went on to write the Fetch Standard.

XHR turned out to be remarkably hard to kill, but I would not call it beautiful. A single object holds the request parameters, receives the response, changes its internal state, and dispatches events, all at the same time. You create it, call open(), hang handlers on it, call send(), watch readyState, and work out separately which state counts as an actual answer. The interface came together gradually, went through several iterations, and captures the era when asynchronous JavaScript meant callbacks.

Fetch did not arrive merely as a nicer XHR. Its job is broader: to give the web platform one shared idea of a request, a response, headers, a body, redirects, and access to somebody else’s domain. JavaScript is not the only thing that loads resources—HTML, CSS, fonts, workers, the media player, and dozens of other APIs do it too, and historically each of them grew its own set of rules. The Fetch Standard set out to describe a single loading machine, and the fetch() function became the part of it a developer can see.

A familiar situation, once again: when Opera Presto was frozen, work on Fetch was underway but nowhere near even a rough draft. The first sketch of the Fetch API appeared on May 27, 2014; Chrome and Firefox shipped it in 2015, and Safari only caught up in 2017.

Today the technology is table stakes, and a pile of libraries and services hang their work on that one call. Without fetch() you cannot even google. Which is why, once Promise and async/await landed in Carakan, Fetch became one of the most obvious and most critical things to add.

XXXII.I The False Simplicity of Fetch

Let’s surface back into the present.

Show a browser an <img>, a <script>, or a link to a stylesheet and it knows exactly what to do: assemble an HTTP request, find the resource in the cache or go out to the network for it, sort out redirects, cookies, and certificates, then hand the bytes to whichever subsystem asked for them. Which is why the modern fetch() function looks almost insultingly simple:

const response = await fetch("/api/news");
const news = await response.json();

Give the browser an address, wait for the response, read JSON out of it. In an engine that already does all of this, you would think such an API would take half an hour to implement. You could even take the XMLHttpRequest that is already there, wrap it in a Promise, and call the job done.

You could—but you can’t.

Yes, XHR and Fetch use the same browser machinery. DNS, connections, proxies, TLS, cookies, HTTP authentication, and the disk cache are shared between them. What differs is how the browser describes the operation, drives it, and presents the result to JavaScript.

  XHR Fetch
Core abstraction One mutable request object Request and Response values plus a single-use operation controller
Setup open(), then headers, responseType, timeout, then send() Every parameter is captured in a Request snapshot before the load begins
Result Written back into the same XHR object A separate Response is created
Asynchrony States and events Promise
Reading data responseText, response, responseXML text(), json(), blob(), arrayBuffer(), formData()
Reading twice Usually possible The body is single-use: reading it sets bodyUsed; a second consumer needs clone()
Progress Has upload/download progress events No progress events
Cancellation xhr.abort() An external AbortSignal that can be handed to several operations
Timeout Built-in timeout property Usually a timer plus AbortController
Synchronous mode Exists, historically Absent by design

XHR is built as a small state machine. A call to open() moves it from UNSENT to OPENED, arriving headers move it to HEADERS_RECEIVED, an incoming body to LOADING, and completion to DONE. At every step the browser mutates fields on that same object and fires readystatechange, progress, load, error, timeout, or abort. The object can be reset with a fresh open() and used all over again.

Fetch, by contrast, treats a load as a transformation:

Request → network operation → Response

Request carries the method, the URL, the headers, the body, and the security policies. Once the operation starts, none of that is supposed to change behind your back. The internal DOM_FetchController lives for exactly one operation; when that operation ends, the controller resolves or rejects the promise and destroys itself. The resulting Response exists independently of it.

The security model is where the two diverge most. XHR has same-origin/CORS and withCredentials, but it has no notion of a response that “arrived, yet whose contents you are not allowed to see.” Fetch introduces explicit same-origin, cors, and no-cors modes, credentials and redirect policies, and the response types basic, cors, opaque, and opaqueredirect. no-cors, for instance, may let the request go out and then return an opaque response whose status, headers, and body are all invisible to the script. For XHR, a failed CORS check simply looks like a network error.

Errors are represented differently too. On a network failure XHR gets status == 0 and an error event, while Fetch rejects the promise. But an HTTP 404 or 500 is not a network error for either API: Fetch returns an ordinary Response with ok == false, and XHR ends with a load event and the matching status.

So Fetch cannot be implemented as a JavaScript wrapper over XHR. The two share a low-level transport, but their lifecycles, their rules for reaching the response, their redirect handling, their cancellation model, and their JavaScript execution order are all incompatible. In Presto we reused XHR’s networking bricks—DOM_HTTPRequest, the URL loader, the CORS manager, cookies, and the redirect machinery—but built a separate DOM_FetchController on top of them.

XMLHttpRequest → DOM_XMLHttpRequest → DOM_HTTPRequest → URL/HTTP/TLS/cache
fetch()        → DOM_FetchController → DOM_HTTPRequest → URL/HTTP/TLS/cache

XXXII.II Compromises I Had to Make

Implementing it took two compromises.

Modern Fetch is bound up with Readable Streams. In a proper implementation the Promise from fetch() resolves as soon as the headers arrive, and the body can be read in chunks as it comes off the network. That matters for large files, for streaming protocols, and for processing data without an enormous intermediate buffer.

Presto has no Readable Streams. The Streams API sits in the backlog and will get its turn one day. For now the Fetch implementation is deliberately buffered. Opera receives the response body in full, stores the bytes, and only then hands over a Response. All the methods are asynchronous, but Request.body and Response.body return null: there really is no stream behind them.

The cost is a longer wait for the response and more memory per request, which makes this Fetch unfit for streaming video, a giant download, or parsing an endless response line by line. What it does do is work, and the internal representation is split so that a real Stream can be attached later without throwing Request and Response away.


Picture a page that hits /account twice at once. The first request is the ordinary kind: the browser sends the stored cookies and the server recognizes the user. The second one has to be anonymous—credentials: "omit" explicitly forbids sending cookies or HTTP authentication data. One address, two opposite requirements.

Opera is not ready for that case. The URL_Rep class inside the engine is a super-object for an address. Loads, caches, history, and a good deal more hang off it. If two images on a page point at the same file, it makes sense to find it once and not download it twice. So requests to a single address converge, in many cases, on a single internal super-object.

For images that is an excellent optimization. For Fetch it is a potential problem. Write the “send cookies” rule into the shared address object and the second request may overwrite the setting made by the first. The anonymous request then leaves with cookies attached, or the ordinary one loses its authorization. Which of the two wins depends on which was sent last—and that is not easy to work out even in single-threaded Presto, now that Carakan has schedulers for asynchronous work.

Cases like this were accounted for at design time and checked in the earliest stages of the implementation. In this particular one the bug I expected refused to reproduce, and that in itself was strange.
I had to unwind the whole call chain from DOM_HTTPRequest::GetURL(). At first it looked as though adding any custom header to the call changes its identity in Opera’s internal representation—and Fetch happened to be adding an Accept header. The answer looks like an accident until you remember that the engine was designed by people who knew exactly what they were doing.
One level down. Every visited address lives in the Url_Store registry, and opening a Url always checks what the browser already has for that address. And here is where it stops being an accident: if an incoming request carries non-standard parameters of any kind, that URL_Rep is marked unique, and later requests for its address go around the registry, the caches, and history.
The potential problem from the example above is still there, though. Adding that header on a fetch() call was technically unnecessary—DOM_HTTPRequest::GetURL() adds Accept: */* on its own—it simply came out that way in the first implementation. That code could later be removed, and addresses in the registry would start converging on one URL_Rep again.
As a guard, every Url requested through Fetch now carries the uniqueness flag from the outset. It costs all such requests their caching, unfortunately, but how to get around that is a question for next time.

XXXII.III The Twenty-First Redirect

Remember how we have run into cases where new code exposed old bugs?

One of the WPT tests sends the browser down a chain of redirects. The standard allows no more than twenty of them; on the twenty-first, fetch() is supposed to refuse to go on and return an error. Opera crashed instead.

The cause was in the old networking code. As it tore a load down, it destroyed the resource object at the very moment a network function was still working with it. The handler itself was meant to be deleted slightly later, once that function returned, but before that it managed to receive one more message and reach for a resource that no longer existed.
Fetch was the first thing to reproduce the exact sequence of events. The bug could have been lying there for years, because the older ways of loading never reached that state.
The fix was simple: do not take a load apart while its own handler is still running.

XXXII.IV What Can Opera Do Now?

Well, it can google now.

One level down, that means fetch(), Headers, Request, Response, AbortController, AbortSignal, and URLSearchParams are now available in the window and in dedicated and shared workers. You can make HTTP and HTTPS requests, load data: and blob: URLs, send the common kinds of body, read responses as text, JSON, a form, a Blob, or an array of bytes, clone unused requests and responses, cancel operations, and control credentials and redirect modes. same-origin, CORS, and a safely restricted no-cors all work; an opaque response stays genuinely opaque.

That does not mean the whole modern Fetch Standard is implemented. There are no Readable Streams, no Service Workers, and no Cache API; no request interception or background execution, no streaming uploads or downloads, no keepalive, no Subresource Integrity, and none of several newer policies. Some Request parameters are recognized and then rejected, because the engine cannot deliver the behavior they promise. The work is far from finished, but what is done already makes Opera far more capable on today’s web.

XXXIII. One Engine, Two Databases

I have mentioned before that Opera built its own full-text search across the entire history of visited pages. A very cool piece of engineering that somehow never got a cool name. It is simply Search engine—a micro-DBMS built on B-trees, a compressed index, and purpose-built data structures.

And yet Presto also carries SQLite (version 3.7.9 at the moment of the freeze, bumped by us to 3.53.3), which supports full-text search of its own. We have seen more than once that Opera wrote its own solutions when it saw reason to, and moved to somebody else’s without much fuss when that brought an advantage. Two databases with overlapping capabilities raise a question worth digging into.

XXXIII.I The Invention of Search engine

The search_engine module makes this easy: its author left a full product presentation right there in the source tree. The module is not a relational DBMS and has no SQL, no schemas, no query planner, and no general-purpose secondary index system. Its strength is a handful of structures chosen in advance for the specific operations Opera performs:

This is a very specific tool used in very specific places: searching page history and the contents of mailboxes in the M2 client. It is designed for a small memory footprint and tight time budgets—fast startup, short but frequent inserts and deletes. Every mechanism is fitted to the expected flow of operations and to the browser’s event loop, and written, of course, in the house Opera style.

Every serious article needs a diagram. I stole this one from the original 2006 presentation.

And this is 2006.
Other options were surely at least considered. Looking back from here at what existed in full-text search at the time, the field comes down to this:

Full-text search in the “big” databases is not worth considering here. In small, embeddable SQLite, FTS1 and FTS2 had only just appeared, and they had obvious performance problems.

So why not Sphinx or Xapian? Was this NIH syndrome after all?
Let’s stop clinging to that date first. 2006 is the year Search engine already existed and worked, which means the decision itself was made earlier—in 2005, quite possibly 2004. Sphinx was embryonic back then, and it was designed as a client-server system besides. That closes the case, though other reasons to pass on it could be found.
Xapian is a far more interesting matter. Its ancestor Muscat already had experience searching hundreds of millions of web pages. It is a library, not necessarily a separate server; if I were drawing up a shortlist of candidates, Xapian would be on it. But one open license is not the same as another. Xapian is licensed under GPLv2+, which effectively rules out integrating it into a closed application. That settles this case too, without even bringing up that a stable Xapian 1.0 only arrived in 2007 and would still have needed a substantial adaptation layer.

Which leads to an entirely obvious conclusion: nothing suitable exists, so build your own.

So they did.

XXXIII.II Why SQLite?

So you have a gorgeous database of your own. What did you need SQLite for?

For Web SQL.
There is an instructive story attached to it. The idea of handing web developers an SQL-like interface was in the air, and in 2008 Apple added the API to Safari. Chrome and Opera added it in 2010.
What is the simplest and fastest way to get SQL support into your browser? Take SQLite, obviously. Apple did that, Chrome did that, and Opera, last of all, did the same.

Mozilla and Microsoft did not. They refused to implement the standard at all, though for different reasons.
The Mozilla Foundation insisted that the “obvious path” breaks a W3C rule: “for a draft to become a web standard, at least two independent implementations of the technology must exist in the world.” When every browser uses the same solution, that solution turns into the standard, quirks and bugs included. Anyone who later tries to write an alternative has to reproduce those bugs and quirks exactly.

Microsoft’s position was that a single SQL standard does not exist in practice. Yes, there is nominally ANSI SQL, but every database supports something of its own—T-SQL, PL/SQL, MySQL, PostgreSQL. Microsoft had its own SQL Server, and shipping somebody else’s SQLite dialect in Internet Explorer would have stung its pride.

Both Mozilla and Microsoft held that the web needed not a relational SQL database but a low-level object store: NoSQL. Such a store could be specified strictly—which methods exist, how indexes behave—without tying anything to the parsing of complicated SQL text. Web developers gave them plenty of grief for it: as they saw it, corporate bureaucracy and stubbornness were taking away a tool that was pleasant to work with and pushing them toward the complicated, asynchronous, unfamiliar IndexedDB.

One way or another, by the formal W3C rules Mozilla and Microsoft were right. On November 18, 2010, the Web SQL specification was declared deprecated and no longer under development.

Did the other browsers remove Web SQL then? Of course not.
The standard had taken root so firmly, especially in iOS web apps, that ripping it out in one motion was simply dangerous, and it stayed switched on for years afterward. Safari dropped support only in September 2019, and Chrome not until October 2023. Opera, for its part, died with the standard still enabled.

Web SQL has not gone anywhere even now. With WASM you can compile SQLite straight into the browser, using IndexedDB or OPFS as the storage. It sounds like magic and like an unheard-of pile of hacks at the same time, but this is the case where the “magic hacks” became the standard solution, supported from the SQLite side and the browser side alike.

XXXIII.III Yes, But Why Keep Both?

I still have not answered the question. You have one excellent database; what do you need a second one for? Their capabilities clearly overlap, and the logical move would be to delete one of them and extend the other where needed.

And indeed both have pages or blocks, caching, B-trees, space allocation and reuse, transactions, journals, and corruption handling. Could Opera have added SQL on top of Search engine?
Here I step onto the shaky ground of guesswork. Picture an Opera engineer assigned to implement a new standard. New, well—Safari has had it for a couple of years, but Safari does not count, because there Web SQL exists for the sake of web apps. Now, though, Chrome has announced it too, which means it is our turn.
Search engine could be extended, but the tool is awfully specialized. Porting everything SQLite has into it is years of work. Simply adding SQLite is weeks.
Now run it the other way. You now have SQLite, which at that point implements FTS3. Why keep Search engine? Because Search engine already works beautifully on every platform. Porting search over to SQLite is low priority, and putting it off is the rational call.
Six months later Web SQL is declared deprecated, and you have every reason to believe SQLite will be gone in a couple of years. A couple of years later, the thing that is gone is Opera itself.

I cannot prove this is how it went, but it holds together.

Is there any point in holding on to Search engine now, when SQLite can be built with an excellent FTS5? In theory you could measure. Write a compatibility layer, generate piles of random data. For history search, measure cold and warm startup, adding a document while a page loads, deleting a range of history, prefix and phrase search, ranking the top N results, snippet generation, memory, and on-disk size. For M2: import, daily inserts and deletes, search, opening large folders, recovery after a forced shutdown, and size. Every run has to be done on Windows and Linux, on SSD and HDD. And then you decide.

Happily, my job is to restore the browser as it was, so I get to skip all of that nonsense.

XXXIII.IV The Forgotten Feature

None of which means that Opera’s engineers, holding a tool this advanced if this narrow, never thought about reusing it. While digging through the code I found a related feature: TWEAK_URL_SEARCH_ENGINE_CACHE.

The idea is simple. The browser cache is files. There are a lot of them, and they are small. Directory listings on an HDD are slow. So put Search engine to work storing the cache index: the resources themselves have no business being in there, of course, but the metadata fits perfectly once you build the right structures for it.

A fair amount was built, but the feature was never finished. Its flag is off in every configuration, and the code contains outright errors and plain unfinished bits—it was clearly abandoned. Why?
Because a simpler answer turned up. The dcache4.url index file that already existed was taught to store the metadata of every resource and, when a resource is smaller than 2 KB, the resource itself. Resources between 2 KB and 16 KB went into container files tied to domains, and anything larger stayed a file on disk, with only the directory structure optimized.

As a result, a small favicon, a CSS fragment, or a server response no longer requires:

The documentation preserves the measurements: at the price of a few megabytes added to the index and slightly higher memory use, overall cache read speed went up by 50% and write speed by a factor of ten.

The TWEAK_URL_SEARCH_ENGINE_CACHE code turned out to be unnecessary, and it was forgotten. Finishing that system now makes even less sense, so I deleted it.

XXXIV. The IndexedDB Iceberg and the Fate of Web SQL

All of this has been the preamble, carrying you toward the real task as inexorably as an iceberg toward the Titanic.

The object store standard that Mozilla and Microsoft argued for was adopted: IndexedDB. It is a transactional object-oriented NoSQL key-value database, asynchronous, with a set of features aimed squarely at web app scenarios. And this one everybody implemented their own way. Well, almost.
Microsoft used its proprietary Extensible Storage Engine—the same technology that powers Windows Search indexing and stores the Active Directory user catalog. Their IndexedDB came out crooked and not quite complete, adding one more item to Internet Explorer’s long list of sins.
Apple and, unexpectedly, Mozilla used SQLite, translating NoSQL into SQL transparently to the developer.
Google, meanwhile, wrote an engine of its own, LevelDB, sharpened strictly for NoSQL operations and faster for it. It now lives in every Chromium-based browser.

It would be entertaining to watch what would have happened after the death of Internet Explorer had Google taken the same path as Apple and Mozilla. Spoiler: the old circus would not have repeated itself, the situations are not alike.

So what do we choose?
In 2026 there is no shortage of NoSQL stores. Let’s pretend the obvious answer is not staring us in the face and walk the whole list.

  1. Memcached. Solves a different problem: a distributed in-memory cache that tolerates losing and evicting data, running as client-server. Reworking it would cost a great deal.
  2. Redis. Another class of solution again. It can act as persistent storage, but its lifecycle does not match the IndexedDB profile at all.
  3. LevelDB. Google’s own, straight out of Chrome—fast and reliable. There is a catch: transactions, indexes, uniqueness constraints, migrations, and blobs are not part of the store. Chromium implements that layer itself, which means we would have to as well.
  4. RocksDB. Technically suitable, but its profile is server workloads and large volumes of data. It brings complicated configuration, background threads, caches, and a heap of files in exchange for capabilities that would never be used. Overkill.
  5. Berkeley DB. A fit, if a heavy one. But the licensing is awkward: either a commercial Oracle license or the GNU AGPL.
  6. LMDB. A good, compact embedded B-tree store with ACID and MVCC. But the whole database is mapped into the address space, which runs against Opera’s philosophy. On top of that it still needs its own implementation of schemas, foreign keys, migrations, and partial unique indexes.
  7. Something of our own, on top of that same Search engine, for instance. Technically possible, and it would be interesting, were there not a queue of a dozen other subsystems waiting to be written.
  8. SQLite. Already integrated into the code, and it brings transactions, savepoints, indexes, integrity constraints, and migrations. It makes the best use of Opera’s existing machinery and moves more of the dangerous invariants onto a proven engine.

SQLite fits Opera’s approach, so IndexedDB will be built on top of it.

And Web SQL? It is there, in fully working order, and nobody needs it. I left it in the engine, hidden behind a user setting. The question of removing it is worth revisiting only once Presto fully supports IndexedDB, WASM, and OPFS.

XXXV. What’s Next?

It hit me all at once that I have spent my entire month of vacation on this project. Not a shred of regret—it was fascinating. But I will most likely have to slow down now.
Even so, I expect to add IndexedDB to Opera and the tools for working with it to Dragonfly. That work has already begun, and the finish line is in sight.

After that, with no dates promised, comes ECMAScript 2026, including Unicode 17, the current RegExp features, the new data types, and all of Intl. It sounds wildly ambitious, but without it everything already done somehow loses its point.

The rest, we’ll see.