Changelog

New features, improvements, and fixes — everything we ship, as we ship it.

Subscribe via RSS
Clear Pick a year to narrow by month; a date range overrides both pickers.
Improved Fixed

Browser rendering upgraded to Chromium 149, automation fingerprints removed #

Every read that needs a real browser now runs on Chromium 149 (up from 143) through a patched automation layer. That covers /v2/perceive, the pages it renders for /v2/distill, /v2/ingest and /v2/watch, and the V1 URL-to-PDF, URL-to-image and URL-to-Markdown converters. The browser no longer exposes the signals bot-detection scripts most often use to spot automation: the DevTools Runtime.enable side effects and the --enable-automation switch are gone, so the old JavaScript-level stealth patches are no longer needed.

Nothing changes in the API. Requests, responses and options are the same, and render_quality counts JavaScript errors exactly as before, so scores stay comparable with earlier reads. js_code still runs in the page's own context and can read its globals, and a wait_for that times out still degrades to a warning. PDFs and screenshots now come from the newer Chromium, so layout can differ slightly from earlier captures of the same page.

  • Private addresses blocked throughout a render. Iframes and redirects that point to private or internal network addresses are now blocked during rendering, just as the requested URL itself always was. Every redirect hop is checked, and a chain longer than 8 redirects now fails the read instead of being followed.
  • Faster recovery after a browser crash. If the process that drives the browser dies, the next request relaunches it, where requests used to fail until the browser's scheduled restart.
New

The render_quality scorer is open source #

The scorer that produces render_quality, deductions and is_blocked on every web read is now a standalone MIT library: render-quality on PyPI and npm. It ports the seven HTML-only deductions (anti_bot_challenge, bot_detection, login_wall, empty_body, unhydrated_shell, http_error, soft_404) with the same weights and thresholds the API uses, so you can reproduce an API verdict offline, score pages from crawl4ai, Firecrawl or Jina Reader with the bundled adapters, and gate your own pipeline on the same 0.40 floor. The four runtime deductions (js_errors, resource_failures, domain_redirect, slow_render) need a live browser and remain API-only, so the hosted score can be lower than the library's for the same HTML, never higher.

Improved

Playground renders are capped #

Anonymous renders from the public playground (a JWT minted from the playground key) now run the TLS rung and plain Chromium only, inside a 60 second ladder budget, with no stealth escalation. cache_mode: bypass|refresh and wait_for are rejected with 422, /v2/perceive/batch is rejected with 403, and each visitor IP is capped at 10 playground requests per minute and 60 per hour. Requests made with your own API key are unchanged: full ladder, full budget, every option.

Improved

Reads that hit 401, 407 or 451 return immediately #

The render engine ladder (TLS -> Chromium -> stealth Chromium) used to re-render pages that answered 401 Unauthorized, 407 Proxy Authentication Required or 451 Unavailable For Legal Reasons on every rung, burning up to three renders to fetch the same error page. Like 404/410, these are now treated as the origin's definitive answer: the first render returns with status_code set and the http_error deduction applied. 403, 429 and 5xx still escalate, since those are frequently bot-gate pages a stealthier browser gets past.

Fixed

A browser crash mid-render is now a 503 with Retry-After, not a 502 or 500 #

When our headless Chromium died during a render, url-to-pdf, url-to-screenshot and url-to-markdown answered 502 Bad Gateway / empty_render, which blamed the target site for our fault, and /v2/perceive answered a generic 500. Neither gave clients anything to retry on.

All of them now return:

HTTP/1.1 503 Service Unavailable
Retry-After: 5
{"error": "Service Unavailable", "code": "browser_restarting", "detail": "The rendering browser restarted while rendering <artifact> for <url>. Retry in a few seconds."}

The browser is relaunched in the background as soon as the failed request releases its slot, so a retry after 5 seconds lands on a warm browser. Genuine target-side failures (upstream_timeout 504, upstream_unreachable 502, empty_render 502) are unchanged, and /v2/perceive now uses the same typed envelope for them instead of a 500, with the operation_id in an X-Operation-Id header so support can still find the request.

Improved

Request rate limits are on; free plan raised to 30 requests per minute #

Short-window fairness limits are now enforced on all billable POST endpoints (/v1/convert/*, /v1/extension/*, /v2/*). These are separate from your monthly ops quota (402).

Plan Private key, per minute per hour per day
Free 30 (was 5) 100 1,000
Starter 50 1,000 10,000
Pro 150 3,000 30,000
Business 300 6,000 60,000

When a window is exhausted you get 429 Too Many Requests with Retry-After (seconds until the window resets) plus RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset headers. GET polling, downloads and /health are never limited.

Public (pk_) keys used in browsers also get a per-visitor limit of half your plan's public-key per-minute and per-hour windows (at least 3 per minute and 30 per hour). It is keyed on each visitor's own IP address, with IPv6 visitors grouped by their /64 network, so one heavy visitor of your widget cannot use up your project's window for everyone else.

Fixed

Ingest skips pages it could not actually read #

/v2/ingest now applies the same render-quality verdict every /v2/perceive read carries. A page whose render is blocked or scores below 0.40 (a login wall, a soft-404 with body text, an anti-bot challenge page) is recorded as skipped with the reason, for example the page could not be read (render_quality 0.30; http_error), contributes no chunks to the JSONL and is not billed. Previously such pages were chunked, billed and shipped into your corpus.

Every chunk of a web page now carries the verdict in its metadata so you can filter at load time:

{"metadata": {"source_url": "...", "render_quality": 0.9, "deductions": {"slow_render": 0.1}, ...}}

Uploaded files are not rendered and their records are unchanged. A job whose every page was skipped fails with the dominant skip reason in error_message.

Fixed Improved

Blocked, error-page and login-wall reads are free #

Every perceive response now carries billed: boolean. A read is not charged against your monthly operations when is_blocked is true or when its deductions include http_error (the origin answered 4xx/5xx) or login_wall. You paid for content, not for a challenge page or a sign-in form.

An unbilled read is a verdict, not a delivery: it returns render_quality, deductions, status_code and warnings with outputs: {} and no structured data, and with direct_download: true it returns that JSON instead of a file. Unbilled verdicts are never served from the 1 h cache, so retrying the URL always re-renders. Cache hits of normal reads remain billed as before.

Improved Fixed Security

Distill fixes and clearer conversion errors #

If a distill call came back with everything null, there was often no way to tell whether the data wasn't on the page or the page never reached the model. You were charged either way. Pages with a </body> inside a comment or a <style> block no longer get cut short, <title> and <meta description> are readable again, and when a page is trimmed to fit or the model runs out of room mid-answer, the response says so instead of returning a silent null.

You can trust data against your schema now, so there's no need for defensive parsing. A scalar target_field gets the value rather than the whole record, missing values are real null instead of the string "null", arrays stay arrays, and failed URLs carry "data": null instead of dropping the key and breaking typed clients. Prompt-only mode keeps its types, so an integer comes back as 22.

Mistakes cost nothing now. A target_field that isn't in your schema, a regex field with no capture group, or a schema with no fields return 422 with the reason, instead of quietly falling through to a paid extraction. A single bad URL used to fail the whole batch with a 500 and lose the results for URLs that had already finished; now it's one failed row. Validation errors no longer echo back headers and cookies, which is where API keys and session cookies tend to live.

File conversion errors now tell you what's actually wrong. PNG to SVG checks the file signature and says the file isn't a PNG, rather than passing through a decoder message with a server path in it. JSON to CSV and XML to CSV also accept single objects and rows with different keys, which used to fail outright.