What a web reading API should return besides the markdown
Late on 16 September (UTC), while we were finishing the release that changed how blocked reads are billed, we sent g2.com through our own web reading API. DataDome answered with a 403 and an empty page. The read came back with render_quality 0.0, status_code 403 and two deductions, empty_body and http_error, at 0.7 each. Two deductions worth 0.7 each, so the score clamps at zero.
The warnings on that read also carried Crawl4AI's own message, "Blocked by anti-bot protection: DataDome captcha". Our render runs on Crawl4AI, and its detector recognised the DataDome page. Our scorer only saw an empty page behind a 403. Neither of our two anti-bot deductions fired, so is_blocked was false.
Under the rules we were running that night, an empty 403 like g2's came back as a delivered read with a low score attached, and it cost an op. You paid us for a 403. If the scorer did see a bot block with nothing behind it, /v2/perceive failed with a 502 instead (unless you set allow_degraded, which handed back the challenge page itself), and the agent got a gateway error with the deduction names buried in an error string and no score.
The release we published on 17 September changed both. Blocked and error-status reads now come back as an HTTP 200 with billed: false and no artifacts, so a read like g2's costs nothing and still tells the agent what happened (the docs have the exact rule).
g2.com is also row 1 of the benchmark we ran on 7 July: 50 deliberately hostile URLs (18 bot-gated, 15 paywalled, 17 JavaScript-heavy), fetched from a residential IP in India with httpx, a Chrome user agent and no JavaScript. 26 of the 50 came back as a block or gate page that a naive consumer would have stored as content, g2 among them (an empty shell behind a 403). That count included 401, 403, 429, 451 and 503 bodies on purpose, because code that reads resp.text without checking the status ingests those too. The file's other arm, labelled Enconvert, was Crawl4AI run locally with our browser config, not the hosted API, so I'm leaving its numbers out.
EnConvert is a web reading API for AI agents. You send a URL to /v2/perceive, and from one render you get back markdown, HTML, a screenshot, a PDF, links, images or structured data, plus a verdict on the read in four fields render_quality, named deductions, is_blocked and billed).
What a web reading API should hand back
A plain HTTP fetch gives you the content and leaves the question of whether it's the page you asked for to a status code your code may never check. You'd see the captcha and move on. Your agent gets a string, and it will summarise a DataDome challenge and cite it back to the user as g2's homepage. A RAG pipeline embeds it and retrieves it weeks later as documentation (the RAG ingestion post is about that failure).
outputs takes any of markdown, html_cleaned, html_raw, screenshot, screenshot_full_page, pdf, links, images and structured, and defaults to ["markdown", "structured"]. Files come back as signed URLs that expire after 15 minutes. structured is inline. If one expires, GET /v2/perceive/{operation_id} re-signs everything without a new render or an op, as long as you're inside your plan's retention window (an hour on the free plan). The URL to Markdown post covers what we strip from the markdown and how we check we didn't strip the article with it.
The verdict is those four fields plus the origin's status code:
- render_quality: 0.0 to 1.0. It starts at 1 and every deduction that fires subtracts its weight. Under 0.40 is a failed render.
- deductions: the names and weights that fired. There are eleven. Two are about bots anti_bot_challenge, bot_detection). The others cover things like http_error, soft_404 (a "page not found" served with a 200), login_wall, empty_body and unhydrated_shell (a JavaScript app whose mount node never filled in).
- is_blocked: true when one of the two bot deductions fired, meaning the scorer took the page for a block or challenge page.
- billed: whether this read cost you an op.
- status_code: what the origin's main document answered.
Crawl4AI already ships its own anti-bot detector (the one that flagged g2). I wanted one graded number with named reasons on top of that, covering the failures that have nothing to do with bots, and the same 0.40 floor reused in ingest, distill and watch. Ingest skips pages under it before it chunks anything, and the LLM extraction step doesn't run on a page under it. On /v2/perceive itself the floor doesn't decide delivery or billing, so that gate is yours.
The gate to put in agent code
The rule is is_blocked or render_quality < 0.40, and you need both halves.
bot_detection on its own takes 0.5 off, which leaves the page at 0.50, above the floor. Only is_blocked catches it. anti_bot_challenge alone leaves exactly 0.40, and 0.40 isn't less than 0.40, so again is_blocked is what stops it. And the g2 read scored 0.0 with is_blocked false. A gate on is_blocked alone would have passed it as a good read, and your agent would have gone looking for markdown that isn't there.
import os
from enconvert import Enconvert
FLOOR = 0.40
def usable(op) -> bool:
if op.is_blocked:
return False
return op.render_quality is not None and op.render_quality >= FLOOR
client = Enconvert(api_key=os.environ["ENCONVERT_API_KEY"])
op = client.v2.perceive("https://example.com", outputs=["markdown", "screenshot"])
print(op.render_quality, op.is_blocked, op.billed, op.deductions, op.status_code)
md = op.outputs.get("markdown") # absent when the read was not billed
if usable(op) and md is not None:
print(md.url) # 15-minute signed URL
else:
print("not content:", op.status_code, sorted(op.deductions))Don't index into outputs. An unbilled read comes back with outputs: {} and structured: null, so use .get(). A read can also be billed and still fail the gate. A soft 404 on its own scores 0.35, delivers its markdown and costs an op, because the origin said 200 and handed over a page. billed: true only means we delivered a page, so keep the gate on billed reads too.
When the gate says no, hand the agent the status code and the deduction names instead of the page. A model can read "403, http_error, empty_body" and tell the user it couldn't get the site.
The scorer is a heuristic, and it's wrong sometimes. It misses some block pages, and it can flag a genuine article that quotes Cloudflare's challenge text (that read then comes back empty and unbilled). All eleven weights are in the web perception post, along with the misses and the small synthetic test set we measure it on.
The same read from curl
Asking for markdown and a viewport screenshot:
curl -X POST https://api.enconvert.com/v2/perceive \
-H "X-API-Key: $ENCONVERT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com", "outputs": ["markdown", "screenshot"]}' \
| jq '{render_quality, is_blocked, billed, deductions, md: .outputs.markdown.url, png: .outputs.screenshot.url}'A blocked or not-delivered read is still an HTTP 200, so your code has to check billed, is_blocked and the score in the body. Here are two full responses, one of each kind, both asking for outputs: ["markdown"] and captured on 8 October 2026 at about 06:06 UTC. First g2.com, complete and unedited:
{
"operation_id": "per_c2dbcd0357fe4b0387ff7376951aee70",
"status": "completed",
"url": "https://www.g2.com",
"url_final": "https://www.g2.com",
"content_hash": "d308fb540ff4329b651792b8c582571a7c407b513cdfad1b071c69593f1462a3",
"status_code": 403,
"render_quality": 0.0,
"deductions": {
"empty_body": 0.7,
"http_error": 0.7
},
"is_blocked": false,
"billed": false,
"options_echo": {
"outputs": [
"markdown"
],
"only_main_content": true,
"truncate_data_arrays": true,
"allow_degraded": false,
"extract": [],
"cache_mode": "enabled",
"mobile": false,
"respect_robots": false,
"direct_download": false,
"wait_for": null,
"wait_timeout_ms": 30000,
"viewport": null,
"block_resources": [],
"js_code_provided": false,
"schema_provided": false,
"pdf_options_provided": false,
"auth_provided": false,
"cookies_provided": false,
"headers_provided": false
},
"cache_hit": false,
"outputs": {},
"structured": null,
"extraction_tier": "heuristic",
"tokens": {
"input": 0,
"output": 0
},
"cost_cents": 0.0,
"duration_ms": 8909,
"error": null,
"warnings": [
"engine tls_http verdict=thin (quality 0.00); escalating",
"crawl4ai flagged this page (Blocked by anti-bot protection: DataDome captcha); outputs were captured by our hooks and are returned anyway.",
"engine chromium verdict=thin (quality 0.00); escalating",
"engine ladder returning best-effort tls_http render (quality 0.00); all engines blocked/thin",
"render quality: the page was scored as not delivered (deductions: empty_body, http_error); no artifacts were produced and this read is not billed."
]
}Same answer as in September: a 403, a score of 0.0, empty_body and http_error, is_blocked false, billed false and outputs: {}. The warnings trace the read up the ladder. The fast fetch came back thin, so did Chrome, and once every engine had come back blocked or thin the API kept the best-effort render and scored it as not delivered. All of that took 8.9 seconds. The second warning still says outputs "are returned anyway". That text is older than the billing rule, and outputs is the field to trust.
Then python.org, straight after. I've cut options_echo (identical to the one above) and the query string of the signed URL:
{
"operation_id": "per_9fe9e26650084c82bd913cfe219772cb",
"status": "completed",
"url": "https://www.python.org",
"url_final": "https://www.python.org",
"content_hash": "d354d8cf6dd6257aa05be2e6845db6747404793f1f98abd671491341a27c7a2a",
"status_code": 200,
"render_quality": 1.0,
"deductions": {},
"is_blocked": false,
"billed": true,
"options_echo": { ... },
"cache_hit": false,
"outputs": {
"markdown": {
"url": "https://nyc3.digitaloceanspaces.com/econverter/live/files/1/v2-perceive/per_9fe9e26650084c82bd913cfe219772cb_markdown.md?[signed query string cut]",
"object_key": "live/files/1/v2-perceive/per_9fe9e26650084c82bd913cfe219772cb_markdown.md",
"size_bytes": 3019,
"content_type": "text/markdown; charset=utf-8",
"expires_in": 900
}
},
"structured": null,
"extraction_tier": "heuristic",
"tokens": {
"input": 0,
"output": 0
},
"cost_cents": 0.0,
"duration_ms": 1321,
"error": null,
"warnings": []
}A score of 1.0 with nothing deducted, billed, and a 3,019-byte markdown file behind a URL that expires in 900 seconds. No warnings, and it came back in 1.3 seconds
The Python SDK is pip install enconvert. The perceive docs cover the options I skipped here, like wait_for, js_code, and cookies and headers for pages behind a session.
What a read costs
One URL read is one op, whatever outputs you ask for (a schema that triggers LLM extraction on a paid plan also draws AI credits). A read isn't billed when is_blocked is true or the deductions include http_error or login_wall. Those come back as a 200 with billed: false and no artifacts. Every other read that completes bills one op, including cache hits. An identical request within an hour is served from cache, which saves you render time, not ops.
Unbilled verdicts are never cached, so retrying a blocked URL renders it again, and the retry is free if it fails again. Requests that fail outright (a 504 timeout, a 502 because the site couldn't be reached) aren't counted either, since ops are counted on completion.
The free plan is 500 ops a month with no card. Paid plans are on the pricing page.
How the page gets rendered
HTML-only reads try a fast fetch with a real Chrome TLS fingerprint first, and escalate to a real headless Chrome render when the page looks blocked or scores under 0.40. Screenshot and PDF requests go straight to Chrome. The Chrome side runs on Crawl4AI with patchright, and it dismisses cookie banners and scrolls for lazy-loaded content before capturing.
That fallback once hid a bug of ours for about two weeks: the fast path was dead, and reads kept arriving because every request fell through to Chrome (the browser API post has the story).
Other readers, and when I'd use them
All of this is as of October 2026, from each vendor's own docs or code.
The shortest path from a URL to text is Jina Reader. Prefix the URL with https://r.jina.ai/ and LLM-ready text comes back (in its own Title: / URL Source: format by default, or plain markdown if you send x-respond-with: markdown), without a key at 20 requests a minute. It's Apache-2.0. Its open-source code adds Warning: lines for non-2xx statuses and for short pages that look like a CAPTCHA. If you're prototyping and you'll parse those lines, it's the quicker start.
If you need to crawl and map whole sites from one vendor, with managed proxies included, Firecrawl is the better pick. The core is AGPL-3.0. Its scrape response carries metadata.statusCode and metadata.error, and I didn't find a block flag for HTML pages in the documented schema.
If you want to run the whole thing on your own machines, use Crawl4AI. Our rendering runs on it, and it's Apache-2.0 with the anti-bot detector built in.
For static pages at volume, Browserbase Fetch returns raw page content with no browser session for about $1 per 1,000 pages (per their March 2026 changelog). Markdown and JSON output, which they call Fetch Extract, are priced separately. Fetch doesn't execute JavaScript and has a 5 MB content limit. When your agent needs to log in or click through something, use their browser sessions.
If you're keeping the reader you have, the scorer behind render_quality is an MIT library pip install render-quality or npm install @enconvert/render-quality) with adapters for Crawl4AI, Firecrawl and Jina Reader output. It runs the seven checks that only need HTML. The other four need a live browser, so they stay in the API.
Limits I'd want to know about
The API is labelled beta. Perceive is generally available inside it, but the docs still warn that parameters and response formats may change.
We don't solve CAPTCHAs or get around paywalls, and we don't run residential proxies, so a DataDome wall like g2's blocks us too. A blocked read carries no content, not even the challenge HTML, so perceive can't show you what the site served.
Right now renders run one at a time on a single node, and we're in one region. Our trust page calls that "a real ceiling". Send a burst and browser renders queue for that one slot. If none frees up in time you get a 503 with Retry-After, and the rate limit (30 requests a minute on a free-plan private key) applies on top.
For a week, log the deductions (or your reader's equivalent) and the status code next to every page your agent reads, then count how many of the pages it cited were error pages, challenge pages, login walls or empty shells. Whichever reader you use, I'd like to hear that number ([email protected]).
The playground's Perceive tab takes a URL with no key.