How our web perception API scores pages, and where it's wrong
In August 2026 a test read of OpenAI's platform docs got redirected to a 75-word login page that says "Sign up or login", and our scorer gave it 0.75. The login check looked for "log in" and "sign in" as whole words. login isn't log in, and Sign up isn't sign in, so it matched nothing.
That was the smaller of two holes the testers found that week. A documentation URL that answered HTTP 404 had come back as content with render_quality: 1.0. So we pulled five documentation 404s and scored them: 0.45, 0.85, 0.85, 1.0 and 1.0. Every one cleared the 0.40 floor we tell customers to gate on. We sell a web perception API, and it was handing an agent a 404 and letting it believe it had read the docs.
Nothing in the scorer looked at the HTTP status. The status code reached three places in our own code and we threw it away in all three. The browser hook, for one, received the response object and never read its status. The root-cause note from that week calls render_quality "an anti-bot detector wearing the name of a quality score."
Both fixes were committed on 6 August and shipped the next day. We added two deductions, http_error and soft_404, rewrote the login_wall check, and put status_code and the named deductions on every response so the reason for a low score is on the wire. Re-scored, that docs 404 lands at 0.30, and at 0.35 even when the server lies and answers 200. The OpenAI login page went from 0.75 to 0.20.
EnConvert is a web perception API for AI agents. Whatever you ask /v2/perceive for (markdown, a screenshot, structured data and so on) comes back from one render with a verdict on that render: render_quality from 0 to 1, the named deductions that lowered it, is_blocked, the origin's status_code and billed. Perceive is generally available inside an API we still label beta. The web reading API post covers the request side.
How the web perception API scores a read
The arithmetic is deliberately dumb. Each deduction that fires removes a fixed amount, and the score is 1.0 minus the sum, clamped to 0..1. Every weight is a multiple of 0.05, so every possible score is one too.
The floor is 0.40 and the check is a strict less-than, so 0.40 passes and 0.35 fails. We sized the weights so one detected failure, with nothing else wrong, lands at or near the floor:
- login_wall or soft_404 alone: 0.35
- empty_body (hard tier), unhydrated_shell or http_error alone: 0.30
- anti_bot_challenge alone: exactly 0.40
- bot_detection alone: 0.50
The last two are why the score on its own isn't a gate. A page that trips only anti_bot_challenge (challenge markers among 100 or more visible words, or an article quoting challenge text) lands on 0.40 and passes a < 0.40 check. A block page that trips only bot_detection sits at 0.50. Both set is_blocked, which is true exactly when anti_bot_challenge or bot_detection fired. So the rule I'd put in agent code checks both. Treat the page as usable only if is_blocked is false and render_quality is at least 0.40.
The 11 deductions
Deduction | Weight | Fires When |
|---|---|---|
| 0.60 | 2 or more of 5 challenge markers ("just a moment", "checking your browser", "ray id", "cf-browser-verification", "challenge-platform") appear anywhere in the raw HTML, markup included |
| 0.50 | 2 or more of 11 block phrases ("access denied", "captcha", "are you a robot" and so on) in the visible text, or 1 strong signature such as "why have i been blocked" on a page under 200 words |
| 0.65 | auth words ("sign in", "login", "create account" and others) counted on word boundaries: 3 or more under 200 words, 2 or more under 120 words, or any password input |
| 0.70 / 0.15 | under 20 visible words / under 100 |
| 0.70 | a framework mount node |
| 0.20 / 0.10 | more than 10 / more than 3 console errors |
| 0.15 / 0.05 | more than 10 / more than 3 failed images, scripts, stylesheets, fonts or media |
| 0.10 | the registrable domain changed between the URL you asked for and the one you landed on |
| 0.10 / 0.05 | page load over 30 s / over 15 s |
| 0.70 | the main document answered 400 or above |
| 0.65 | the title or first h1 is a 404 phrase, the page has under 250 distinct words, and |
The two bot rules read different text. Cloudflare's strongest fingerprints are element ids and script paths, so anti_bot_challenge matches raw markup. bot_detection counts only visible text, because a normal page can ship "captcha" and "unusual traffic" inside its JavaScript bundle without showing them (our regression test is a crypto dashboard that does that).
Soft 404 detection is the fiddliest rule. Google's definition of a soft 404 is a page that says it doesn't exist while answering 200, and single-page-app docs platforms do it a lot. The title or h1 has to be the error phrase (the patterns are anchored), so an article titled "Understanding the 404 not found error" doesn't fire. It counts distinct words rather than raw words because the August docs 404 padded itself with hundreds of words of font-probe filler ("word word word") but had only about 150 distinct ones.
unhydrated_shell exists because an empty <div id="root"> wrapped in 120 words of nav and footer once scored a clean 1.0 on the fast no-browser path (that story is in the browser post ).
The four runtime deductions js_errors, resource_failures, domain_redirect, slow_render) read signals from the fetch itself rather than the HTML, and together they can remove at most 0.55. Two of them, js_errors and resource_failures, can't fire on the fast fetch at all, since there's no browser there to count console errors or failed images.
How the verdict changes routing and billing
The score also picks the renderer. HTML-only reads try a fast fetch with a real Chrome TLS fingerprint first, and escalate to a real headless Chrome render when the page looks blocked or scores under 0.40. Screenshot and PDF requests go straight to Chrome.
A 401, 404, 407, 410 or 451 is the origin's final answer, so the first render comes back with status_code set and http_error applied. A 403, 429 or 5xx still escalates, because those are often bot gates that the Chrome render can get past. Until 17 September, 401, 407 and 451 escalated too, and we re-rendered the same error page for nothing.
You pay one op per URL read, cache hits included. Three things make a read free: is_blocked, or an http_error or login_wall deduction. Those reads come back as HTTP 200 with billed: false, outputs: {} and structured: null, and they're never cached, so a retry renders again (the perceive docs have the full contract). We hold the artifacts back on those reads because if they still shipped, anyone who controls a site could get free PDFs and screenshots by answering with a 404 and a full body, or by parking a password field on a short page.
Every other read under 0.40 is billed and delivered. A soft 404 at 0.35 or an empty shell at 0.30 comes back with its outputs and costs the op, and only the score and its deductions say it's wrong. Code that checks billed and nothing else will pass soft 404s to the agent.
Two reads I made on 8 October 2026 show the difference, and both score 0.35. The first is a soft 404. deno.com answers a path that doesn't exist with HTTP 200 and a page whose heading is "404". Captured at 08:59:06 UTC, with the signed URL's query string cut:
{
"operation_id": "per_42578404a2a645ab9a7d7ba6b7a48de4",
"status": "completed",
"url": "https://deno.com/this-page-does-not-exist",
"url_final": "https://deno.com/this-page-does-not-exist",
"content_hash": "e7caeeaa8df9d6da5a98d4ee75614d4b31d2d63daad6fa4b139a059ee11cb1a9",
"status_code": 200,
"render_quality": 0.35,
"deductions": {
"soft_404": 0.65
},
"is_blocked": false,
"billed": true,
"options_echo": {
"outputs": [
"markdown"
],
"only_main_content": true,
"truncate_data_arrays": true,
"allow_degraded": false,
"extract": [],
"cache_mode": "bypass",
"mobile": false,
"respect_robots": false,
"direct_download": false,
"wait_for": null,
"wait_timeout_ms": 30000,
"viewport": null,
"block_resources": [],
"js_code_provided": false,
"schema_provided": false,
"pdf_options_provided": false,
"auth_provided": false,
"cookies_provided": false,
"headers_provided": false
},
"cache_hit": false,
"outputs": {
"markdown": {
"url": "https://nyc3.digitaloceanspaces.com/econverter/live/files/1/v2-perceive/per_42578404a2a645ab9a7d7ba6b7a48de4_markdown.md?[signed query string cut]",
"object_key": "live/files/1/v2-perceive/per_42578404a2a645ab9a7d7ba6b7a48de4_markdown.md",
"size_bytes": 2640,
"content_type": "text/markdown; charset=utf-8",
"expires_in": 900
}
},
"structured": null,
"extraction_tier": "heuristic",
"tokens": {
"input": 0,
"output": 0
},
"cost_cents": 0.0,
"duration_ms": 5447,
"error": null,
"warnings": [
"engine tls_http verdict=thin (quality 0.35); escalating",
"engine chromium verdict=thin (quality 0.35); escalating",
"engine ladder returning best-effort tls_http render (quality 0.35); all engines blocked/thin",
"only_main_content: no extraction retained enough of the page's content; returned the full page instead (fidelity guard)."
]
}Billed, and delivered as a 2,640-byte markdown file of Deno's not-found page. The fidelity guard also fell back to the full page. If your code reads billed: true as "got the page", this is what it hands the agent.
The second is a login wall. Logged out, GitHub's profile settings URL redirects to its sign-in page. Captured at 08:59:43 UTC, with options_echo cut because it matches the one above:
{
"operation_id": "per_278ab8f677ee408cb81d0a9d6796ba6a",
"status": "completed",
"url": "https://github.com/settings/profile",
"url_final": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fsettings%2Fprofile",
"content_hash": "78a20844042a6f8fb3e0ee75c4db0a282e8177f8a5b360ec47f289120f875c4b",
"status_code": 200,
"render_quality": 0.35,
"deductions": {
"login_wall": 0.65
},
"is_blocked": false,
"billed": false,
"options_echo": { ... },
"cache_hit": false,
"outputs": {},
"structured": null,
"extraction_tier": "heuristic",
"tokens": {
"input": 0,
"output": 0
},
"cost_cents": 0.0,
"duration_ms": 9418,
"error": null,
"warnings": [
"tls_http followed 1 redirect(s) to https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fsettings%2Fprofile",
"engine tls_http verdict=thin (quality 0.35); escalating",
"engine chromium verdict=thin (quality 0.35); escalating",
"engine ladder returning best-effort tls_http render (quality 0.35); all engines blocked/thin",
"render quality: the page was scored as not delivered (deductions: login_wall); no artifacts were produced and this read is not billed.",
"render quality: page appears to be a login wall (score 0.35)."
]
}Same score, same 200 from the origin. This time there are no outputs and no charge, because login_wall is one of the three things that make a read free. In both reads Chrome saw exactly what the fast fetch saw, so the API returned the fast fetch's version.
The same 0.40 floor gates other V2 endpoints too. /v2/ingest skips pages that are blocked or under it and doesn't bill them. It writes render_quality and deductions into every chunk's metadata (the RAG ingestion post covers that gate). A watcher check that is blocked or under the floor is never diffed, so a detected challenge page can't fire a change alert (Watch is in private beta). Paid LLM extraction is skipped on reads that are blocked or under the floor.
Where the scorer is wrong
We built the scorer on 7 June 2026 and checked it against a labelled corpus of 25 pages that should fail and 25 that shouldn't. The pages are synthetic, HTML snapshots modelled on real vendor pages with made-up browser signals, because we had no production reads to label yet.
We re-ran it in October with the current scorer and the production gate is_blocked, or a score under 0.40). It catches 22 of the 25 bad pages and false-alarms on 1 of the 25 good ones.
The three misses:
- A Cloudflare error 1020 "Access denied" page scores 0.85. It carries one challenge marker ("Ray ID") and one generic block phrase ("access denied"). Each rule wants two, and "access denied" isn't one of the strong single-phrase signatures. Only the thin-page tier of empty_body fires, on 62 words.
- An EU cookie-consent interstitial on a cross-domain TrustArc portal scores 0.90. The article never rendered, and domain_redirect is the only deduction that fires.
- A skeleton page, with the nav and sidebar rendered and the main area all shimmer placeholders saying "Loading your projects...", scores 0.40. That's 0.15 for a thin body, 0.20 for console errors, 0.15 for failed resources and 0.10 for a slow load. The floor is a strict less-than, so it passes.
The false alarm is a 1,416-word security article explaining Cloudflare challenges. It quotes "Just a moment...", "Checking your browser" and example Ray IDs, trips anti_bot_challenge and gets is_blocked. In the hosted API that read comes back unbilled with no outputs, so you'd get the article's score and no markdown.
Fifty synthetic pages is a small sample, so treat 22 of 25 as a check on the rules and nothing more. The corpus records carry no status code, so http_error never fires in this test. If a 1020 page in the wild arrives with a 403, http_error would catch it, but I haven't checked that against a live capture, so it stays in the miss column.
Our public benchmark, perceive-benchmark, runs six readers, a naive fetch included, against 200 bot-gated, paywalled, JS-heavy and control URLs. Its v2 runs go out from GitHub-hosted runners, and on 28 September 83 of the 200 calls to our own hosted API ended in HTTP 524, Cloudflare's origin-timeout error. Until that's fixed, we aren't quoting a hosted-API number from it.
Crawl4AI, Jina Reader, Firecrawl and Zyte
We didn't invent block detection. Our browser render runs on Crawl4AI with a patchright-driven Chromium, and Crawl4AI ships its own detector. What I found in each project's docs or source, as of October 2026:
- Crawl4AI's antibot_detector.py exports is_blocked(status_code, html, error_message), with tiered checks for Cloudflare, DataDome, PerimeterX and other vendors. Its stated philosophy is that false positives are cheap, because a fallback rescues them, and false negatives are catastrophic. We run Crawl4AI 0.8.9, so its flag also lands in our warnings on browser renders.
- Jina Reader's open-source code adds Warning: lines to the output: "Target URL returned error" with the code for any non-2xx, and a CAPTCHA hint on short pages whose HTML contains captcha or cf-turnstile-response. I confirmed that in the open-source branch, not on the hosted r.jina.ai.
- Firecrawl returns metadata.statusCode and metadata.error. Its pricing page says, effective 4 September 2026: "A scrape that returns no result is not charged. A page that responds with an error status such as 403 or 404 is still returned to you and costs 1 credit."
- Zyte API answers HTTP 520 when it couldn't avoid a ban in reasonable time, and doesn't charge for it.
All four are defensible designs, and Crawl4AI leaning towards flagging makes sense inside a crawler that retries. Most of what we built isn't about bots. Only two of our 11 deductions are. Billing reads the deduction names, so an HTTP error or a login wall is unbilled, like a block. And the same floor applies in perceive, ingest, watch and LLM extraction. I didn't find a numeric per-read score in the docs or source of those four this month. I may have missed one.
Run the scorer on any reader's output
On 17 September we published the scorer as an MIT library, render-quality, with the same weights and thresholds the API uses.
pip install render-quality
npm install @enconvert/render-qualityIt ports the seven deductions that only need HTML. The four runtime ones need signals that HTML doesn't carry (console errors, failed subresources, the final URL, load time), so they stay in the API. For the same HTML and status code, the hosted score can be lower than the library's, never higher. The Python package needs only BeautifulSoup, and the npm one has no dependencies.
The example fetches with httpx (pip install httpx). Any fetcher works.
import httpx
from render_quality import QUALITY_FLOOR, score
resp = httpx.get("https://example.com/docs/some-page", follow_redirects=True)
verdict = score(resp.text, status_code=resp.status_code)
if verdict.is_blocked or verdict.render_quality < QUALITY_FLOOR:
print("not the page:", verdict.deductions)
for reason in verdict.reasons:
print(" ", reason)Pass the status code. It defaults to None, which means http_error never runs, and a hard 404 gets caught only if something else gives it away, such as a 404 title or h1 on a page with under 250 distinct words, or a near-empty body. That's the hole we fell into in August.
If you already use another reader, the adapters do the field mapping. from_crawl4ai(result) reads a CrawlResult's html and status_code. from_firecrawl(scrape_json) takes the status from metadata.statusCode and prefers rawHtml, because Firecrawl's html is a cleaned copy that drops the markup the challenge check reads. from_jina(text, status_code) scores Jina's markdown with score_text, which has no DOM to inspect, so it skips unhydrated_shell and the password-field check and says so in reasons.
The 1020 page and the challenge-text article were known hard cases in the June corpus, and they still score wrong today. If you hit a page the scorer gets wrong, the HTML and the status code are enough to reproduce it with the library. Send me both as an issue on the render-quality repo.