---
title: "Web content extraction API: what each perceive output gives you"
description: "Markdown, links, images, structured data, screenshots and PDFs from one URL read and one op: what each output contains and when to use which."
canonical: "https://www.enconvert.com/blog/web-content-extraction-api"
locale: "en"
author: "Het Dave"
published: 2026-10-08
---

# Web content extraction API: what each perceive output gives you

- Published: 2026-10-08
- Author: Het Dave

Markdown, links, images, structured data, screenshots and PDFs from one URL read and one op: what each output contains and when to use which.

When we measured the 13 August markdown release, one notebook page dropped 94% once its embedding vectors were truncated. An output cell had printed them as plain text, and every number landed in our markdown.

Nothing was broken. The converter kept what was on the page, and no model needs a raw embedding vector in its context, so a web content extraction API has to decide what to leave out.

That truncation is the `truncate_data_arrays` option ([changelog](https://www.enconvert.com/changelog/entry/cleaner-perceive-markdown)). A run of 64 or more comma-separated numbers collapses to its first 16 values plus a count `... [truncated <dropped> of <total> values]`), inside code fences too, since notebook output cells become fenced blocks. Unset, it follows `only_main_content`, which defaults to true. On the default main-content path, if the tidy pass cuts the markdown by more than half, `warnings` says so. `"truncate_data_arrays": false` brings the numbers back.

The page was LlamaIndex's Pinecone metadata-filter example, which has since moved to developers.llamaindex.ai and still prints those vectors. I read it twice on 8 October 2026, at about 12:02 UTC. With `truncate_data_arrays` unset, the markdown came back at 10,971 bytes with one warning: "long numeric data arrays were truncated in the markdown (set truncate_data_arrays=false to keep them in full)." With `"truncate_data_arrays": false` it came back at 167,735 bytes with no warnings, so the default file is 93.5% smaller. Both files have 283 lines, because truncation shortens lines and never drops them. Lines 196 to 205 of the default file, unedited:

```
INFO:httpx:HTTP Request: POST https://api.openai.com/v1/embeddings "HTTP/1.1 200 OK"
HTTP Request: POST https://api.openai.com/v1/embeddings "HTTP/1.1 200 OK"

[NodeWithScore(node=TextNode(id_='7fed3d0b-e2d7-432a-9231-3ac02130c00f', embedding=[0.00310940156, -0.0246712118, -0.0222742166, -0.0364649445, -0.00715911388, 0.0117236068, -0.0400604382, -0.0275654588, -0.0157589763, 0.00740136392, 0.0278459582, 0.029962454, 0.0187679715, -0.00440511806, 0.00412780605, 0.00442743069, ... [truncated 1520 of 1536 values]], metadata={'director': 'Christopher Nolan', 'theme': 'Fiction', 'year': 2010}, excluded_embed_metadata_keys=[], excluded_llm_metadata_keys=[], relationships={}, hash='7937eb153ccc78a3329560f37d90466ba748874df6b0303b3b8dd3c732aa7688', text='Inception', start_char_idx=None, end_char_idx=None, text_template='{metadata_str}\n\n{content}', metadata_template='{key}: {value}', metadata_seperator='\n'), score=0.317687511),
 NodeWithScore(node=TextNode(id_='6b7bdb44-9dc2-48f6-b9cb-05dcaea4889b', embedding=[0.0131883333, -0.0137667693, -0.0396421254, -0.00729471678, -0.0107653309, 0.0171731133, -0.0282276608, -0.0507480912, -0.0210422054, -0.01610622, 0.0254768785, 0.0143452045, 0.0170060098, 0.00128059229, -0.0103797065, -0.000858012936, ... [truncated 1520 of 1536 values]], metadata={'author': 'J.K. Rowling', 'theme': 'Fiction', 'year': 1997}, excluded_embed_metadata_keys=[], excluded_llm_metadata_keys=[], relationships={}, hash='1b24f5e9fb6f18cc893e833af8d5f28ff805a6361fc0838a3015c287510d29a3', text="Harry Potter and the Sorcerer's Stone", start_char_idx=None, end_char_idx=None, text_template='{metadata_str}\n\n{content}', metadata_template='{key}: {value}', metadata_seperator='\n'), score=0.467440248)]
```

The same ten lines with truncation off. I've cut lines 202 and 203 after their first 12 values, and the marker at each cut gives the full length of the line:

```
INFO:httpx:HTTP Request: POST https://api.openai.com/v1/embeddings "HTTP/1.1 200 OK"
HTTP Request: POST https://api.openai.com/v1/embeddings "HTTP/1.1 200 OK"

[NodeWithScore(node=TextNode(id_='7fed3d0b-e2d7-432a-9231-3ac02130c00f', embedding=[0.00310940156, -0.0246712118, -0.0222742166, -0.0364649445, -0.00715911388, 0.0117236068, -0.0400604382, -0.0275654588, -0.0157589763, 0.00740136392, 0.0278459582, 0.029962454, [cut here: this line is 23,188 characters long]
 NodeWithScore(node=TextNode(id_='6b7bdb44-9dc2-48f6-b9cb-05dcaea4889b', embedding=[0.0131883333, -0.0137667693, -0.0396421254, -0.00729471678, -0.0107653309, 0.0171731133, -0.0282276608, -0.0507480912, -0.0210422054, -0.01610622, 0.0254768785, 0.0143452045, [cut here: this line is 23,173 characters long]
```

EnConvert's `/v2/perceive` is a web content extraction API for agents. You send a URL, and from one capture of the page you can take markdown, cleaned or raw HTML, a link map, an image list, structured metadata, a screenshot and a PDF, plus a `render_quality` score that says how far to trust the capture. It's one op per URL however many outputs you ask for. Perceive is generally available, and the beta label sits on the API as a whole.

The markdown side has [its own post](https://www.enconvert.com/blog/url-to-markdown-api.md), and Krystin covered raw versus clean text in [How to extract clean web content for RAG pipelines and LLM agents](https://www.enconvert.com/blog/how-to-extract-clean-web-content-for-rag-pipelines-and-llm-agents.md) . I'll stick to the other eight outputs.

## Nine outputs from one web content extraction API call

```
curl -X POST https://api.enconvert.com/v2/perceive \
  -H "X-API-Key: $ENCONVERT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "outputs": ["markdown", "links", "images", "structured", "screenshot_full_page"],
    "extract": ["metadata", "structured_data", "headings", "tables"]
  }'
```

File outputs come back as signed URLs under `outputs.<name>.url`. `structured` is inline in the JSON. To try it before writing code, use the keyless Perceive tab in the [playground](https://www.enconvert.com/playground.md?mode=perceive) .

Here is that request against Wikipedia's Mohs scale article, which has JSON-LD and three real data tables, captured on 8 October 2026 at about 12:05 UTC. I added `"cache_mode": "bypass"` so the read was fresh. I've cut the signed URLs' query strings, the `options_echo` block, and all but the first three rows of each table. Nothing else is changed:

```
{
  "operation_id": "per_b7adc935f535448fa4c51b890ef00432",
  "status": "completed",
  "url": "https://en.wikipedia.org/wiki/Mohs_scale",
  "url_final": "https://en.wikipedia.org/wiki/Mohs_scale",
  "content_hash": "376e1f880669317c7f5183636f23f34a8a9aebd9bbf8f5e289c546459b3ad0c9",
  "status_code": 200,
  "render_quality": 1.0,
  "deductions": {},
  "is_blocked": false,
  "billed": true,
  "options_echo": { ... },
  "cache_hit": false,
  "outputs": {
    "markdown": {
      "url": "https://nyc3.digitaloceanspaces.com/econverter/live/files/1/v2-perceive/per_b7adc935f535448fa4c51b890ef00432_markdown.md?[signed query string cut]",
      "object_key": "live/files/1/v2-perceive/per_b7adc935f535448fa4c51b890ef00432_markdown.md",
      "size_bytes": 28520,
      "content_type": "text/markdown; charset=utf-8",
      "expires_in": 900
    },
    "links": {
      "url": "https://nyc3.digitaloceanspaces.com/econverter/live/files/1/v2-perceive/per_b7adc935f535448fa4c51b890ef00432_links.json?[signed query string cut]",
      "object_key": "live/files/1/v2-perceive/per_b7adc935f535448fa4c51b890ef00432_links.json",
      "size_bytes": 49345,
      "content_type": "application/json",
      "expires_in": 900
    },
    "images": {
      "url": "https://nyc3.digitaloceanspaces.com/econverter/live/files/1/v2-perceive/per_b7adc935f535448fa4c51b890ef00432_images.json?[signed query string cut]",
      "object_key": "live/files/1/v2-perceive/per_b7adc935f535448fa4c51b890ef00432_images.json",
      "size_bytes": 71784,
      "content_type": "application/json",
      "expires_in": 900
    },
    "screenshot_full_page": {
      "url": "https://nyc3.digitaloceanspaces.com/econverter/live/files/1/v2-perceive/per_b7adc935f535448fa4c51b890ef00432_screenshot_full_page.png?[signed query string cut]",
      "object_key": "live/files/1/v2-perceive/per_b7adc935f535448fa4c51b890ef00432_screenshot_full_page.png",
      "size_bytes": 1460161,
      "content_type": "image/png",
      "expires_in": 900
    }
  },
  "structured": {
    "metadata": {
      "title": "Mohs scale - Wikipedia",
      "description": null,
      "keywords": null,
      "author": null,
      "og:image": "https://thumb.wikimedia.org/wikipedia/commons/thumb/f/fd/Mohssche-haerteskala_hg.jpg/960px-Mohssche-haerteskala_hg.jpg?utm_source=en.wikipedia.org&utm_campaign=index&utm_content=thumbnail",
      "og:image:width": "908",
      "og:image:height": "1200",
      "og:title": "Mohs scale - Wikipedia",
      "og:type": "website"
    },
    "structured_data": [
      {
        "@context": "https://schema.org",
        "@type": "Article",
        "name": "Mohs scale",
        "url": "https://en.wikipedia.org/wiki/Mohs_scale",
        "sameAs": "http://www.wikidata.org/entity/Q41472",
        "mainEntity": "http://www.wikidata.org/entity/Q41472",
        "author": {
          "@type": "Organization",
          "name": "Contributors to Wikimedia projects"
        },
        "publisher": {
          "@type": "Organization",
          "name": "Wikimedia Foundation, Inc.",
          "logo": {
            "@type": "ImageObject",
            "url": "https://www.wikimedia.org/static/images/wmf-hor-googpub.png"
          }
        },
        "datePublished": "2001-11-06T03:42:32Z",
        "dateModified": "2026-08-28T13:41:17Z",
        "image": "https://upload.wikimedia.org/wikipedia/commons/f/fd/Mohssche-haerteskala_hg.jpg?utm_source=en.wikipedia.org&utm_campaign=index&utm_content=original",
        "headline": "qualitative ordinal scale characterizing scratch resistance of various minerals"
      }
    ],
    "headings": [
      {
        "level": 2,
        "text": "Contents"
      },
      {
        "level": 1,
        "text": "Mohs scale"
      },
      {
        "level": 2,
        "text": "Reference minerals"
      },
      {
        "level": 2,
        "text": "Examples"
      },
      {
        "level": 2,
        "text": "Use"
      },
      {
        "level": 2,
        "text": "Comparison with Vickers scale"
      },
      {
        "level": 2,
        "text": "Footnotes"
      },
      {
        "level": 2,
        "text": "See also"
      },
      {
        "level": 2,
        "text": "References"
      },
      {
        "level": 2,
        "text": "Further reading"
      }
    ],
    "tables": [
      {
        "headers": [
          "Mohshardness",
          "Referencemineral",
          "Chemical formula",
          "Absolutehardness[14]",
          "Example image"
        ],
        "rows": [
          [
            "1",
            "Talc",
            "Mg3Si4O10(OH)2",
            "1",
            ""
          ],
          [
            "2",
            "Gypsum",
            "CaSO4·2H2O",
            "2",
            ""
          ],
          [
            "3",
            "Calcite",
            "CaCO3",
            "14",
            ""
          ],
          ... 7 more rows cut
        ],
        "caption": "",
        "summary": "",
        "metadata": {
          "row_count": 10,
          "column_count": 5,
          "has_headers": true,
          "has_caption": false,
          "has_summary": false,
          "id": "mwfg",
          "class": "wikitable sortable jquery-tablesorter"
        }
      },
      {
        "headers": [
          "Hardness",
          "Substance"
        ],
        "rows": [
          [
            "0.2–0.4",
            "Potassium[15]"
          ],
          [
            "0.5–0.6",
            "Lithium[15]"
          ],
          [
            "1",
            "Talc"
          ],
          ... 29 more rows cut
        ],
        "caption": "",
        "summary": "",
        "metadata": {
          "row_count": 32,
          "column_count": 2,
          "has_headers": true,
          "has_caption": false,
          "has_summary": false,
          "id": "mwARw",
          "class": "wikitable"
        }
      },
      {
        "headers": [
          "Mineralname",
          "Hardness (Mohs)",
          "Hardness (Vickers)(kg/mm2)"
        ],
        "rows": [
          [
            "Tin",
            "1.5",
            "VHN10 = 7–9"
          ],
          [
            "Bismuth",
            "2–2.5",
            "VHN100 = 16–18"
          ],
          [
            "Gold",
            "2.5",
            "VHN10 = 30–34"
          ],
          ... 14 more rows cut
        ],
        "caption": "",
        "summary": "",
        "metadata": {
          "row_count": 17,
          "column_count": 3,
          "has_headers": true,
          "has_caption": false,
          "has_summary": false,
          "id": "mwAeo",
          "class": "wikitable"
        }
      }
    ]
  },
  "extraction_tier": "heuristic",
  "tokens": {
    "input": 0,
    "output": 0
  },
  "cost_cents": 0.0,
  "duration_ms": 36254,
  "error": null,
  "warnings": []
}
```

| Output | What You Get |
| --- | --- |
| `markdown` | Main content by default. Image URLs become alt text, except when the Readability candidate wins |
| `html_cleaned` | Crawl4AI's cleaned HTML, with nav, header and footer still in it |
| `html_raw` | The HTML as captured (rendered DOM, or the fast fetch's HTML) |
| `links` | Object with `internal` and `external` array |
| `images` | List of `{src, alt, desc}` entries, plus `width` and `format` when known |
| `structured` | Inline JSON: metadata, JSON-LD, headings, tables, main content |
| `screenshot` | Viewport PNG, 1920 x 1080 by default |
| `screenshot_full_page` | Full-height PNG |
| `pdf` | One continuous page, or paginated with `pdf_options` |

`links`, `images`, `html_cleaned`, `metadata` and `tables` come from the scraper in [Crawl4AI](https://github.com/unclecode/crawl4ai), the open-source crawler our render runs on, and the filtering in the next two sections is theirs. We add the `render_quality` score and choose which of their fields reach you.

### The links file

Each entry in `internal` and `external` has `href`, `text`, `title` and `base_domain`. `base_domain` is the domain with its subdomains dropped, the page's for internal links and the target's for external ones. Read from docs.python.org, a link to blog.python.org counts as internal to python.org. A crawler can walk `internal` and just log `external`, filtering by host first if it should stay on one subdomain.

The list covers the whole page whatever `only_main_content` says, nav and footer included. Entries are de-duplicated by resolved URL and the earliest anchor on the page wins, so a link in the nav before the article keeps its nav text. On the Mohs scale page, the first five `internal` entries are all page chrome (a skip link and four sidebar links), and the first two `external` ones are a donation link and a photography-contest banner. Unedited:

```
{"href": "https://en.wikipedia.org/wiki/Mohs_scale", "text": "Jump to content", "title": "", "base_domain": "wikipedia.org"}
{"href": "https://en.wikipedia.org/wiki/Main_Page", "text": "Main page", "title": "Visit the main page [alt-z]", "base_domain": "wikipedia.org"}
{"href": "https://en.wikipedia.org/wiki/Wikipedia:Contents", "text": "Contents", "title": "Guides to browsing Wikipedia", "base_domain": "wikipedia.org"}
{"href": "https://en.wikipedia.org/wiki/Portal:Current_events", "text": "Current events", "title": "Articles related to current events", "base_domain": "wikipedia.org"}
{"href": "https://en.wikipedia.org/wiki/Special:Random", "text": "Random article", "title": "Visit a randomly selected article [alt-x]", "base_domain": "wikipedia.org"}

{"href": "https://donate.wikimedia.org/?uselang=en&wmf_campaign=en.wikipedia.org&wmf_medium=sidebar&wmf_source=donate", "text": "Donate", "title": "", "base_domain": "wikimedia.org"}
{"href": "https://commons.wikimedia.org/wiki/Commons:Wiki_Loves_Monuments_2026_in_the_United_States", "text": "Photograph a historic site, help Wikipedia, and win a prize. Participate in the world's largest photography competition this month!\n                Learn more", "title": "", "base_domain": "wikimedia.org"}
{"href": "https://www.wikidata.org/wiki/Special:EntityPage/Q41472", "text": "Edit links", "title": "Edit interlanguage links", "base_domain": "wikidata.org"}
{"href": "https://commons.wikimedia.org/wiki/Category:Mohs_scale", "text": "Wikimedia Commons", "title": "", "base_domain": "wikimedia.org"}
{"href": "https://upload.wikimedia.org/wikipedia/commons/transcoded/1/16/En-us-Mohs.ogg/En-us-Mohs.ogg.mp3", "text": "", "title": "Play audio", "base_domain": "wikimedia.org"}
```

If I only wanted the article's links I'd read them from the markdown, which keeps them inline. And if JavaScript injects links after load, pass a `wait_for` selector that only appears once they exist (that also forces the Chrome render).

### Images and what Crawl4AI drops

Crawl4AI filters `images` before you see it. It drops anything that looks like an icon, logo or button (judged by the parent's class, the `src` or the `alt`), images sitting directly in a `<button>`, inline `display:none` and `data:` URIs. It scores the rest, a point each for a width attribute over 150, a height attribute over 150, alt text, sitting in the earlier half of the page's images, a recognisable image format in the URL, a `srcset`, and a `<picture>` parent. Three points gets an image in. An image low on the page with no size attributes, no alt and no `srcset` doesn't make it.

A responsive image appears once per candidate URL (the `src` plus each `srcset` entry, with `width` from its `w` descriptor). We don't pass Crawl4AI's group id through, so grouping the variants is on you. They share `alt` and `desc`.

On the Mohs scale page, the 14 entries in `images.json` are 7 pictures at two sizes each. Every `src` starts with `//` rather than `https:`, because that's how Wikipedia writes them, so resolve each one against the page's scheme before you fetch it. `desc` is Crawl4AI's guess at the text around the image. For the talc photo inside the minerals table it's that table row, and for 6 of the 14 entries it's 10,804 characters of the article itself. The first five entries, with the four long `desc` values cut:

```
{"src": "//thumb.wikimedia.org/wikipedia/commons/thumb/f/fd/Mohssche-haerteskala_hg.jpg/250px-Mohssche-haerteskala_hg.jpg?utm_source=en.wikipedia.org&utm_campaign=parser&utm_content=thumbnail", "alt": "Open wooden box with ten compartments, each containing a numbered mineral specimen.", "desc": "From Wikipedia, the free encyclopedia [cut here: this desc is 10,804 characters long]", "format": "jpg"}
{"src": "//thumb.wikimedia.org/wikipedia/commons/thumb/f/fd/Mohssche-haerteskala_hg.jpg/500px-Mohssche-haerteskala_hg.jpg?utm_source=en.wikipedia.org&utm_campaign=parser&utm_content=thumbnail", "alt": "Open wooden box with ten compartments, each containing a numbered mineral specimen.", "desc": "From Wikipedia, the free encyclopedia [cut here: this desc is 10,804 characters long]", "format": "jpg"}
{"src": "//thumb.wikimedia.org/wikipedia/commons/thumb/5/53/Mohs_scale_vs_absolute_hardness.png/250px-Mohs_scale_vs_absolute_hardness.png?utm_source=en.wikipedia.org&utm_campaign=parser&utm_content=thumbnail", "alt": "", "desc": "From Wikipedia, the free encyclopedia [cut here: this desc is 10,804 characters long]", "format": "png"}
{"src": "//upload.wikimedia.org/wikipedia/commons/5/53/Mohs_scale_vs_absolute_hardness.png?utm_source=en.wikipedia.org&utm_campaign=parser&utm_content=thumbnail_unscaled", "alt": "", "desc": "From Wikipedia, the free encyclopedia [cut here: this desc is 10,804 characters long]", "format": "png"}
{"src": "//thumb.wikimedia.org/wikipedia/commons/thumb/f/fd/Talc_block.jpg/120px-Talc_block.jpg?utm_source=en.wikipedia.org&utm_campaign=parser&utm_content=thumbnail", "alt": "", "desc": "1\nTalc\nMg3Si4O10(OH)2\n1", "format": "jpg"}
```

None of this reads pixels. EnConvert has no OCR. Text inside an image (chart labels, a price table saved as a PNG) never reaches the markdown.

### Structured data without a model

With `structured` in `outputs` and no `extract` list you get `metadata` and `structured_data`. `metadata` is the page's `<head>` as Crawl4AI reads it: title, description, keywords, author, and the `og:`, `twitter:` and `article:` meta tags it finds there. `structured_data` is every JSON-LD block on the page in a single list (broken blocks are skipped and arrays are flattened). JSON-LD is what the site published for machines to read. Look there before you scrape a visible price.

Add `headings` for `[{level, text}]` in DOM order, and `tables` for `{headers, rows, caption, summary, metadata}` on each `<table>` Crawl4AI scores as data rather than layout. `main_content` puts the main-content markdown inline, cut at 50,000 characters with a warning.

`prices`, `contacts` and `technologies` are reserved names that aren't live. Ask for one and you get a line in `warnings` and nothing in `structured`. Named fields from pages you've never seen are the `schema` option. It's an LLM pass, only on paid plans, and [the LLM web scraping post](https://www.enconvert.com/blog/llm-web-scraping-api.md) covers it.

## Screenshots and PDFs

Use a screenshot when the layout carries the meaning. A pricing grid built from `<div>` cards can lose its columns in markdown, and `tables` won't see it because it isn't a `<table>`. Charts and forms have the same problem.

`screenshot` is the viewport, 1920 x 1080 unless you pass `viewport`. `mobile: true` makes it 390 x 844 if you haven't set one. `screenshot_full_page` measures how tall the page is and sets the viewport to that height before it captures. That measurement has been clamped at 25,000 px since 3 August. Infinite-scroll pages report heights of 50,000 px and up, and a full-page raster is width times height times 4 bytes in one allocation (1920 x 50,000 x 4 is 384 MB).

![Full-page screenshot of Wikipedia's Mohs scale article, captured by EnConvert perceive at 1920 x 6,675 pixels](https://econverter.nyc3.digitaloceanspaces.com/live/blog/01e084ea6312.png "screenshot_full_page of the Mohs scale article")

The `screenshot_full_page` from the same request, scaled down: the real file is 1920 x 6,675 px and 1,460,161 bytes.

`pdf` without `pdf_options` is one continuous page, byte-identical to our v1 url-to-pdf, and that page tops out at 60,000 px. Send `pdf_options` and it paginates, A4 portrait by default. Take the field names from the [url-to-pdf reference](https://www.enconvert.com/docs/endpoints/convert/web-pages/url-to-pdf.md): `page_size` (or `page_width` and `page_height`), `orientation`, `margins` (millimetres), `scale`, `grayscale`, `header`, `footer`. An unknown top-level key gets a 422 naming the field, but an unknown key inside `pdf_options` is dropped without a word and you get defaults. That inconsistency is ours, so check the PDF before you assume your options took.

Asking for pixels also changes how the page is fetched. Screenshot and PDF requests go straight to Chrome, while HTML-only reads try a fast fetch with a real Chrome TLS fingerprint first and escalate to a real headless Chrome render only when the page looks blocked or scores under 0.40. Every output in a request comes from the same capture, so adding a screenshot puts the rest of your outputs on the Chrome-rendered page too.

## Signed URLs, retention and billing

Every file URL is signed for 900 seconds. `GET /v2/perceive/{operation_id}` hands back fresh ones, and it neither re-renders nor bills. The files themselves last as long as your plan keeps them (an hour on the free Founding plan, 30 days on Production). `direct_download: true` skips the JSON and streams the file as the response body. It only works with exactly one artifact output `structured` doesn't count).

The 1-hour cache is keyed on the URL plus the render options, `outputs` included. Markdown now and links in a second call is two reads and two ops. Both in one call is one. Cache hits bill too.

A challenge page, a login wall or an HTTP error gets you a verdict and no files. `billed` is `false`, and `outputs` and `structured` come back empty `{}` and `null`). A soft 404 still bills and still hands you files, so gate on `render_quality` before you trust a tidy link map of a missing page ([the web reading API post](https://www.enconvert.com/blog/web-reading-api.md)has the gate).

## When another tool fits better

As of October 2026, [DIffbot's Extract](https://www.diffbot.com/docs/extract/) renders the page, classifies it and returns typed JSON (Article, Product, Event, Job and more). If you want a product's price from pages you've never seen, it's built for that, and our `prices` extract isn't live.

Cloudflare's [Browser Run](https://developers.cloudflare.com/browser-run/quick-actions/crawl-endpoint/) (also checked in October 2026) splits the same jobs into an endpoint each `/markdown`, `/links`, `/json`, `/scrape` and more).

[Mozilla's Readability](https://github.com/mozilla/readability) and [Trafilatura](https://github.com/adbar/trafilatura) are Apache-2.0 libraries that pull the article out of HTML you already have. We run readability-lxml as a candidate in our main-content step, and when it wins, image URLs stay in the markdown. For JavaScript pages or screenshots you'd still have to run a browser alongside either library.

The [perceive reference](https://www.enconvert.com/docs/endpoints/perceive.md) lists every parameter, and the url-to-pdf one has the `pdf_options` fields.
