Changelog

New features, improvements, and fixes — everything we ship, as we ship it.

Subscribe via RSS
Clear Pick a year to narrow by month; a date range overrides both pickers.
Improved Fixed

Markdown conversions are smaller and faster #

markdown-to-html no longer runs Pygments highlighting. The generated page never included a Pygments stylesheet, so all of those token spans rendered with no color at all while making the HTML about 3.6x larger and peak memory about 6x higher. A 763 KB file with 128 code blocks took 29 seconds and 1578 MB, and now takes 5 seconds and 252 MB. Fenced blocks come out as

, which client side highlighters like Prism or highlight.js can pick up, and code heavy documents that used to die with a 504 convert. PDF to Markdown releases each page's cache as it goes and pulls table content out immediately instead of holding live page objects for the whole document, so peak memory tracks the largest page rather than the file. Spreadsheet and CSV to Markdown stop parsing at the 5000 row cap. One edge case changed there: blank rows are dropped after the cap, so a sheet with blank rows scattered through the first 5001 rows renders fewer body rows than it used to.

Fixed

A failed batch creation no longer strands queued pages #

The batch envelope is written before the per-URL rows now. With the old order, a failure between the two steps left operation rows that nothing could resume or sweep, because the resume path only walks envelopes. Those rows consumed the caller's quota and sat in queued forever.

Security Fixed

Batch credentials are no longer stored #

Basic auth, cookies and headers sent with a /v2/perceive batch were written into the job record as JSON and outlived the request, readable by anyone with database or backup access. They are stripped before the record is saved, and only the names of the dropped keys are kept. If the gateway restarts mid-batch and that batch used credentials, its unfinished URLs come back as failed with batch used credentials that are not stored; resubmit the batch instead of quietly rendering login walls for pages you paid to have rendered authenticated.

Fixed Security

Reusing a job_id returns 409 instead of clobbering a row #

job_id is client supplied, so re-sending your own after a timeout is normal, and it now resets the poll row idempotently instead of logging a duplicate key violation. A job_id that belongs to a different project returns 409 job_id already in use and leaves that project's row untouched. The success and failure writers are scoped by project as well, so a guessed ID can no longer flip someone else's job to failed with a caller supplied error message, or point it at a download URL signed for the wrong project.

New Fixed

Runaway renders and post-processing have hard ceilings #

Full page screenshots are capped at 25000 px tall and single page PDFs at 60000 px, tunable with SCREENSHOT_MAX_HEIGHT_PX and PDF_MAX_SINGLE_PAGE_HEIGHT_PX. Infinite scroll pages report heights of 50000 px and up, and a full page raster costs width times height times 4 bytes in one allocation, which took the whole worker down instead of returning anything. Ghostscript grayscale post-processing now times out after 120 seconds, and the child process is killed and reaped on timeout or cancellation, so a pathological PDF fails with an explicit error instead of pinning CPU and outliving the request that started it.

Improved Fixed

Large files stream instead of being buffered #

The download proxy at /v1/convert/download streams from storage in 64 KB chunks instead of reading the whole object into memory first, so time to first byte drops on large files and several clients slowly pulling 150 MB files no longer risk an out of memory. Batch URL jobs, /v2/perceive zip bundles and /v2/ingest assembly write to a temp file as each result finishes and upload it as a stream, so at most one result is resident at a time. Multi-page PDF to JPEG archives are stored rather than deflated, which makes them marginally larger on the wire and quicker to produce, since JPEG does not deflate anyway. Image to PDF hands the image to the renderer as a file instead of a base64 data URI, which used to materialise about five copies of it. The S3 client is built once per process instead of once per call, which takes a TLS handshake off every upload, download, delete and health check.

Improved Security Fixed

Upload size limits measure the real file #

Fifty upload endpoints were reading the Content-Length header, which measures the whole multipart envelope, and a client that sends no Content-Length at all, which includes chunked and HTTP/2 uploads, was not size checked at all. They now measure the uploaded file's own bytes, so plan ceilings (5 MB Free, 15 MB Starter, 50 MB Pro, 150 MB Business) apply on every upload endpoint and the 413 boundary is exact. Admission control weighs bytes as well as request count: with more than 256 MB of conversions in flight the API answers 503 with Retry-After: 10 instead of running the box out of memory. That budget is tunable with MAX_PENDING_CONVERSION_BYTES.

Security Fixed

Images over 40 megapixels are rejected before decoding #

Fourteen image conversion entry points now check the declared canvas size right after reading the file header and return HTTP 400 before allocating a single pixel, with the message Image is too large to process: WxH (N pixels) exceeds the 40000000 pixel limit. A kilobyte upload can legally declare a canvas that decodes to several gigabytes, and Pillow's own bomb detector only fires around 358 MP. This is a real narrowing: a legitimate 50 MP camera photo that used to convert is now rejected. SVGs behave differently. An intrinsic canvas over 40 MP is scaled down with the aspect ratio preserved rather than refused, and an SVG whose intrinsic size cannot be read renders at 2048 px wide. The ceiling is tunable with IMAGE_MAX_PIXELS everywhere except compress-image, which keeps its own 40 MP limit and its own wording.

Fixed New

Watchers stop reporting changes that never happened #

A watch check is a diff against a stored baseline, so it now always captures with Chromium, the same engine that captured that baseline. Before this a check could fall through to the no-browser fetch path, see raw un-hydrated HTML, and fire a change webhook and email for a page that had not changed. The trade-off is that every watch check consumes a Chromium slot. Render quality scoring also recognises an un-hydrated single page app shell: a framework mount node holding fewer than 20 words on a page under 500 visible words scores 0.30 instead of 1.00, which is below the quality floor, so /v2/perceive escalates to a real browser render and a bad watch check is recorded for audit without being diffed or notified on.

Fixed Security

The no-browser fetch path works again #

The TLS fetch engine passed a resolve= argument that curl_cffi does not accept, so every request on that path raised a TypeError and fell through to a full Chromium render. It had been failing 100% of the time for about two weeks, and the raw Python error text was being appended to the warnings array of /v2/perceive responses. Eligible HTML-only pages now skip the browser, which leaves the single Chromium slot for pages that actually need it. The DNS pin that stops a redirect from rebinding to an internal address between validation and connect also takes effect for the first time. It is reapplied on every hop, cleared when a hop does not resolve, skipped for bare IP hosts, and IPv6 addresses are bracketed so the entry parses.