← All posts Engineering · 12 min read

Ten ways to turn HTML into a PDF, ranked by what they cost you

Our export service launched a fresh Chromium for every request. It survived nine concurrent users and then took the host with it. The ten options we wrote on the whiteboard, the four we actually measured, and the unexciting one we shipped.

RBRita Ben-Ari · Maintainer, core
25 August 2026

An answer with citations is only useful inside Quire. The moment somebody has to send it to a colleague who does not run Quire, to an auditor, or to a regulator, it has to become a file, and the file is always a PDF. So quire export renders the evidence sheet — the answer, every cited span, the page images those spans came from — as HTML, and then something has to convert that HTML into a paginated document.

The first version of that something was four lines long. Launch Chromium, open a page, page.pdf(), close Chromium. It is the code every tutorial shows you, it worked on every laptop it was written on, and it is a memory bomb. Each launch is a fresh browser process tree: 90–140 MB resident before it has rendered anything, several hundred milliseconds of startup, and no upper bound on how many of them exist at once. A user on a 4 GB VM exported a batch of thirty evidence sheets from a shell loop, and the OOM killer took the whole quire serve process with it, index and all.

That was issue #2841. It stayed open five weeks. This is what reading the alternatives taught us.

First: does it need a browser at all

This is the fork that decides everything downstream, and it is worth being honest about which side of it you are on.

A browser engine is enormous, and you pull it in for exactly one reason: it is the only renderer that agrees with the browser your users previewed the document in. Evidence sheets are not simple. They carry a heading path, a two-column citation layout, page-image crops with highlight overlays positioned by bounding box, footnotes that must not split, and a repeated table header when a line-item table runs across a page break. All of that is print CSS — break-inside, @page margin boxes, running headers — and print CSS support is precisely where the lightweight renderers diverge.

We tested the fork instead of arguing about it. Our export template rendered through a browser engine and through wkhtmltopdf produced visually identical output on 31 of 40 fixture documents. The nine failures were the same two problems every time: highlight overlays offset by a few points, because the older engine positions absolutely positioned children of a transformed parent differently, and table headers not repeating across page breaks. Both are cosmetic. Neither is acceptable on a document whose entire purpose is showing someone which region of which page a number came from.

A renderer that is right 78% of the time is fine for a report and useless for evidence, because the 22% is invisible until somebody relies on it.

So: browser. That decision costs roughly 300 MB of container image and a real operational burden, and if your documents are invoices with a simple flow layout, do not copy it. Options 7 and 8 below will serve you better and cost almost nothing to run.

The ten options

Everything we considered, in the order we considered it. Memory is the column that matters, because in self-hosted software memory is the resource that fails hardest — CPU contention makes you slow, memory exhaustion makes you dead.

#OptionHow it worksMemoryBest for
1Playwright + shared browserLaunch one Chromium and reuse it. Create and close a page per job.MediumMost applications
2Playwright + browser poolA fixed pool of browsers or contexts with a hard cap on concurrent renders. No unbounded process count.Medium, boundedSimultaneous requests
3Playwright + queue (BullMQ/Redis)The HTTP handler enqueues a job; a fixed number of workers drain the queue.Low in the API, medium in workersTraffic spikes
4Dedicated PDF microserviceThe browser lives in its own container. The main service only posts HTML to it.Very low in the APIProduction systems
5Remote Chrome / BrowserlessConnect over CDP to a Chromium you do not run in-process.Very low locallyServerless, elastic apps
6Puppeteer + clusterA managed cluster of browser workers, strict concurrency, automatic worker recycling.High but controlledHigh PDF volume
7wkhtmltopdfAn external one-shot binary. Old WebKit, effectively no JavaScript, very cheap.Low to mediumInvoices, simple documents
8PrinceXML or similarA commercial print engine with the best print-CSS support available. Runs as its own process.LowComplex print layouts
9HTML-to-PDF APIPost HTML to a vendor, receive a PDF. No browsers on your infrastructure at all.Very lowWhen simplicity beats cost
10Serverless PDF workersIsolated functions or containers with per-invocation memory and concurrency limits.Very low in the APIBursty load

Three we could not use, for a reason that is not technical

Options 5, 9 and 10 are the cheapest to operate, and we would probably have picked one of them if Quire were a hosted product. All three are disqualified by the same sentence on the front page: your documents do not leave your machine.

An evidence sheet is the most concentrated possible extract of a private corpus. It holds the answer, the cited spans, and cropped images of the source pages. Posting that to a rendering vendor moves the confidential part of the document to a third party at the exact moment the user believes they are producing a local file. It does not matter that the vendor is reputable or that the payload is transient. The guarantee is does not leave, and a guarantee with an exception is a marketing claim. Same reasoning as the local-first post, and this is the least convenient place we have had to apply it.

Option 5 survives in a narrower form. Connecting over CDP to a browser you run yourself — a sidecar container on the same host — is fine, and it is how the Kubernetes configuration in Recipes works. It is the remote vendor that is out, not the remote protocol.

What we measured

Four candidates, our 40-document fixture set, a 2 vCPU / 4 GB container, because that is the smallest machine we tell people to run quire serve on. Load is 20 export requests arriving over 10 seconds, which is not much traffic and was enough to separate them decisively. RSS is peak resident set of the whole container.

ApproachPeak RSSp50 renderp95 renderFailures of 20
Browser per request (the bug)3.9 GB, then OOM1.9 s—11
Shared browser, unbounded pages2.1 GB0.7 s4.8 s2
Shared browser, 4 concurrent780 MB0.7 s2.6 s0
Sidecar container, 4 concurrent190 MB API + 810 MB sidecar0.8 s2.7 s0

Two rows in that table were not obvious to us beforehand.

The first is that a shared browser with no concurrency limit is not a fix. It removes the process-per-request cost and leaves the actual problem untouched, because twenty pages rendering at once in one browser is still twenty live renderer processes — Chromium isolates per page regardless. It looks like a fix in single-user testing, which is exactly why it is the version people ship. The limit is the fix. The shared browser is only what makes the limit cheap.

The second is that the sidecar bought no performance at all: 190 MB against 780 MB in one process, identical latency. What it buys is a blast radius. When the renderer wedges, and it does, the container restarts and the index does not. That is worth a lot in software other people operate, and nothing whatsoever on a laptop.

What we shipped

Option 1 with option 2 bolted on, in-process by default, with option 4 available as configuration. Concretely:

  • One browser instance, launched lazily on the first export and never at startup. Most installs never export anything, and they should not pay 120 MB for the possibility.
  • A semaphore of four concurrent renders, configurable, defaulting to min(4, cpu_count). Everything past it waits. Waiting is a feature: a queued export finishes, a rejected one does not.
  • A page per job, closed in a finally, with a per-render timeout that closes the page whether or not the render finished. The most important line in the file, for reasons in the next section.
  • Recycling every 200 renders or 30 minutes idle, whichever comes first. The browser is torn down and the next request relaunches it.
  • An optional sidecar, one environment variable, for anyone running quire serve for a team. Same code path, connected over CDP instead of in-process.

No Redis, no BullMQ, no cluster manager. Option 3 is the right answer for real spike traffic and the wrong answer for a self-hosted tool, because it adds a service the operator has to run, monitor and back up in order to solve a queueing problem that a semaphore and an in-memory wait queue already solve at our volumes. If somebody turns up rendering thousands of documents an hour, the queue is where they should go — and the sidecar boundary is deliberately the place they would plug it in.

The failures that were not about memory

Every one of these cost more time than choosing between the ten options did.

  • Leaked pages. A render that threw between opening a page and closing it left the page alive. Memory grew about 8 MB per failed export and nothing in the metrics pointed at it, because the browser count stayed at one. It presented as a slow leak in the API process. It was a missing finally.
  • Zombie Chromium on shutdown. Killing the parent does not reliably kill the browser tree. On a host where quire serve was restarted nightly by a supervisor, orphaned Chromium processes accumulated for eleven days before anyone noticed the machine had no memory left. Fixed with an explicit close on SIGTERM and SIGINT, and by running an init process in the container so that reaping is somebody's job.
  • The 64 MB /dev/shm default. Chromium in a default Docker container crashes on large pages with an error mentioning neither shared memory nor the page. Two days. The fix is one flag on the container.
  • Fonts. The image had none. Every exported PDF rendered in a fallback face, and no reviewer caught it because they all tested locally, where fonts exist. Digits in a tabular column stopped lining up. The fixture comparison is pixel-based now, and it fails the build.
  • --single-process. Somebody found it, it halved memory in a test, and it produced silently truncated PDFs on documents over roughly 40 pages. Do not.
  • Waiting on the wrong signal. Network idle is not the same as fonts loaded and layout settled. We wait on document.fonts.ready plus an explicit flag the template sets once it has positioned the highlight overlays. Before that, about one export in fifty had overlays a few pixels off — the exact defect the browser engine was chosen to avoid.

What we would do first

If you are about to build this: cap concurrency before anything else, because the cap is what stops the crash and the shared browser is only an optimisation on top of it. Close the page in a finally. Recycle the browser on a counter. Then, when it is other people operating your software, move the whole thing into its own container so its failures stay its own.

And check whether you need the browser at all. We do, and it is the most expensive dependency in the project by an order of magnitude. If your document is an invoice rather than a piece of evidence, options 7 and 8 will render it correctly for a fraction of the cost, and this entire post describes a problem you can decline to have.

The export service lives in github.com/getquire/quire under services/export; the pixel-comparison fixtures are in tests/export/fixtures. Container flags are documented in Docs under Deployment.