Python File Conversion SDK#

enconvert is the official Python client for the EnConvert API: one pip install, one API key, and typed methods for converting files and for reading the live web. It runs on Python 3.9 or newer with a single runtime dependency, requests, and ships inline type hints plus a py.typed marker so mypy and Pyright see everything. The client has two surfaces. Nine conversion methods cover 43 {input}-to-{output} pairs plus URL to PDF, screenshot, and Markdown, backed by three status-polling helpers. The client.v2 namespace adds web intelligence: perceive, discover, lookup, distill, ingest, and watch.

PyPI: enconvert · Source: conversionapi/python-sdk · Python: 3.9+ · Runtime dependency: requests>=2.28 · License: MIT

Install#

pip install enconvert

uv add enconvert and poetry add enconvert work the same way. requests is the only runtime dependency, and the type stubs ship inside the wheel, so there is no types- package to chase.


Quick start#

Read a page the way your agent should, with a quality score attached, then convert a local file through the same client.

import os

from enconvert import Enconvert

client = Enconvert(api_key=os.environ["ENCONVERT_API_KEY"])

op = client.v2.perceive("https://example.com", outputs=["markdown", "structured"])
print(op.outputs["markdown"].url, op.render_quality)

print(client.convert_to_pdf("report.docx", save_to="report.pdf").presigned_url)

Every method is synchronous and blocking. The SDK is server-side only: it authenticates with a private API key, so never ship it inside a desktop, mobile, or browser client. Get a key from your dashboard, and see Authentication for how keys are scoped.


What the client exposes#

Enconvert is the whole public surface. Conversion methods hang off the client directly; everything web-facing lives under client.v2.

Conversion method Endpoint Returns
convert_url_to_pdf(url, ...) POST /v1/convert/url-to-pdf ConversionResult
convert_url_to_screenshot(url, ...) POST /v1/convert/url-to-screenshot ConversionResult
convert_url_to_markdown(url, ...) POST /v1/convert/url-to-markdown ConversionResult
convert_image(file, output_format=...) POST /v1/convert/{input}-to-{output} ConversionResult
convert_document(file, ...) POST /v1/convert/{input}-to-{output} ConversionResult
convert_to_markdown(file, ...) POST /v1/convert/anything-to-markdown ConversionResult
convert_to_pdf(file, ...) POST /v1/convert/anything-to-pdf ConversionResult
convert_website_to_pdf(url, ...) POST /v1/convert/website-to-pdf BatchSubmission
convert_website_to_screenshot(url, ...) POST /v1/convert/website-to-screenshot BatchSubmission
get_job_status(job_id) GET /v1/convert/status/{job_id} JobStatus
get_batch_status(batch_id) GET /v1/convert/batch/{batch_id} BatchStatus
wait_for_batch(batch_id, ...) GET /v1/convert/batch/{batch_id} (polled) BatchStatus
client.v2 capability Methods Returns
Perceive perceive, perceive_direct, get_perceive_operation, download_perceive_artifact, perceive_batch, get_perceive_batch PerceiveResult, PerceiveDirectResult, PerceiveBatchResult
Discover, Lookup, Distill discover, lookup, distill DiscoverResult, LookupResult, DistillResult
Ingest ingest, ingest_files, list_ingest_jobs, get_ingest_job, cancel_ingest_job, retry_ingest_webhook, get_webhook_secret, rotate_webhook_secret IngestJob, IngestJobList, WebhookSecret, WebhookRetryResult
Watch create_watcher, list_watchers, get_watcher, get_watcher_snapshots, update_watcher, delete_watcher Watcher, WatcherList, WatcherSnapshotList

Every argument after the first positional one is keyword-only, snake_case, and optional unless a table says otherwise. Results are frozen dataclasses, so build a new instance rather than mutating one. REST shapes are in Endpoints overview.


File conversion#

Every conversion method takes save_to (a str or os.PathLike; the result is streamed there and parent directories are created) and output_filename (override the generated name). Both are left out of the tables below.

convert_url_to_pdf#

Render any reachable URL to PDF.

result = client.convert_url_to_pdf(
    "https://example.com", single_page=False, viewport_width=1440, save_to="report.pdf"
)
print(result.presigned_url, result.file_size)
Option Type Default Description
single_page bool True True gives one continuous page. False paginates using pdf_options.page_size.
pdf_options PdfOptions -- Page size, orientation, margins, scale, grayscale, header, footer. See PDF options.
viewport_width / viewport_height int 1920 / 1080 Browser viewport size in pixels.

load_media and enable_scroll are both True by default: the first waits for images and video, the second scrolls top to bottom so lazy loaders fire. Three more options reach pages behind a gate: auth (HttpBasicAuth), cookies (list[BrowserCookie]), and headers (dict[str, str]).

Do not combine auth with an Authorization header. The API rejects the conflict rather than guessing which credential wins. Pick one.

convert_url_to_screenshot#

Capture a PNG of any URL.

client.convert_url_to_screenshot("https://example.com", viewport_width=1440, save_to="shot.png")

Takes the same viewport, media, scroll, filename, and browser-access options as convert_url_to_pdf, minus single_page and pdf_options.


convert_url_to_markdown#

Pull clean GitHub-Flavored Markdown out of a URL. Navigation, footers, ads, and scripts are stripped, the main article body is kept, and YAML frontmatter with title, description, url, links, and images is prepended.

client.convert_url_to_markdown("https://example.com/article", save_to="article.md")

Same option set as convert_url_to_screenshot. When you also want a quality score, page metadata, or structured extraction on the same read, use v2.perceive instead.


convert_image#

Convert between jpeg, png, svg, heic, and webp, or rasterize a PDF to JPEG. The input format comes from the filename extension.

client.convert_image("photo.heic", output_format="webp", save_to="photo.webp")
client.convert_image("scan.pdf", output_format="jpeg", save_to="scan.jpeg")

output_format is required and must be one of jpeg, png, svg, heic, or webp; the aliases jpg, yml, htm, and md are normalized for you. file accepts a path string, an os.PathLike, raw bytes, or a FileData(data, filename) wrapper. Raw bytes carry no filename and upload as upload.bin, so prefer FileData whenever the extension matters.

from pathlib import Path

from enconvert import FileData

blob = FileData(data=Path("photo.heic").read_bytes(), filename="photo.heic")
client.convert_image(blob, output_format="webp", save_to="photo.webp")

convert_document#

Convert documents and data formats. output_format defaults to "pdf", and pdf_options is honored when the output is PDF.

from enconvert import PdfMargins, PdfOptions

client.convert_document("report.docx", save_to="report.pdf")
client.convert_document("data.json", output_format="yaml", save_to="data.yaml")
client.convert_document(
    "README.md", pdf_options=PdfOptions(page_size="A4", margins=PdfMargins(top=20)), save_to="r.pdf"
)

Supported inputs: .doc, .docx, .xls, .xlsx, .ppt, .pptx, .html, .htm, .odt, .ods, .odp, .ots, .pages, .numbers, .md, .markdown, .csv, .json, .xml, .yaml, .yml, .toml.

EPUB has no dedicated document pair. Send .epub through convert_to_pdf or convert_to_markdown.


convert_to_markdown#

Auto-detect an uploaded document server-side and return clean Markdown. This is the RAG-ingestion building block for a single file.

client.convert_to_markdown("handbook.docx", save_to="handbook.md")

Accepted inputs: PDF, DOCX, PPTX, XLSX, CSV, HTML, EPUB, TXT and MD, plus legacy and ODF office formats. Images are not supported here, and the endpoint takes no PDF options. For a whole site rather than one file, use v2.ingest.


convert_to_pdf#

Auto-detect an uploaded file server-side and return a PDF.

from enconvert import PdfOptions

client.convert_to_pdf("slides.pptx", save_to="slides.pdf")
client.convert_to_pdf("scan.pdf", pdf_options=PdfOptions(grayscale=True), save_to="gray.pdf")

Accepted inputs: office, ODF, Pages, Numbers, RTF, CSV, HTML, Markdown, plain text, raster images, SVG, EPUB, and an existing PDF as passthrough. Because a .pdf input is passed through, this doubles as a grayscale normalizer.

Only pdf_options.grayscale is honored on this endpoint. The other page-geometry fields are ignored here. When you need real page setup, route the file through convert_document or convert_url_to_pdf instead.

convert_website_to_pdf and convert_website_to_screenshot#

Discover every page of a site, convert each one in the background, and collect a single ZIP. Both are asynchronous by design and return a BatchSubmission immediately.

batch = client.convert_website_to_pdf(
    "https://example.com", crawl_mode="sitemap", exclude_patterns=["/blog/tag/"]
)
print(batch.batch_id, batch.url_count, batch.discovery_method)

status = client.wait_for_batch(batch.batch_id, save_to="site.zip")
print(status.completed, "of", status.total, "pages converted")

for item in client.get_batch_status(batch.batch_id).items:
    print(item.source_url, item.status, item.download_url)
Option Type Default Description
crawl_mode "auto" \| "sitemap" \| "full" server default sitemap reads sitemap.xml only. full adds a breadth-first crawl. auto picks the highest mode available to the key.
include_patterns / exclude_patterns list[str] -- Keep or drop URLs by path fragment. Excludes apply in full crawl mode only.
notification_email / callback_url str -- Address to email, and webhook to call, when the batch finishes.
single_page bool server default PDF only.
pdf_options PdfOptions -- PDF only. See PDF options.

Per-page render options apply to every discovered URL as well: viewport_width, viewport_height, load_media, enable_scroll, auth, cookies, and headers. Anything left unset keeps the gateway's own default.

wait_for_batch accepts interval_seconds (default 5.0), timeout_seconds (default 1800.0), and save_to. It raises APIError(504, ...) if the batch is still processing when the deadline passes.


Supported conversion pairs#

The SDK carries a copy of the gateway's conversion map and rejects an unimplemented pair locally, before any network round-trip, naming the valid outputs for that input in the message.

Input Outputs
json csv, toml, xml, yaml
xml csv, json
yaml, toml json
csv json, xml
markdown html, pdf
html pdf
doc, excel, ppt, odt, ods, odp, ots, pages, numbers pdf
jpeg, png, svg, heic, webp each other, all 20 pairs
pdf jpeg

That is 43 implemented pairs. Check them programmatically:

from enconvert import IMPLEMENTED_CONVERSIONS, valid_outputs_for

valid_outputs_for("json")                  # ['csv', 'toml', 'xml', 'yaml']
"heic-to-webp" in IMPLEMENTED_CONVERSIONS  # True

The complete parameter reference lives in Parameters and options.


Web intelligence (V2)#

Every V2 read carries render_quality, a float from 0.0 to 1.0 that says how honestly the page rendered. A challenge screen, cookie wall, login gate, HTTP error page, or empty SPA shell comes back with a low score, a deductions map naming what fired, and a warnings list. The content is still returned, just flagged, so a bad read never quietly enters your agent's context. Treat the score as a gate and check it before you use the text. PerceiveResult, PerceiveDirectResult, DistillItem, LookupItem.perceive, and WatcherSnapshot all expose it. Concepts are in the V2 overview.

Perceive#

Render one URL into agent-ready artifacts. Synchronous: the call returns the completed operation with signed artifact URLs.

op = client.v2.perceive(
    "https://example.com",
    outputs=["markdown", "screenshot", "structured"],
    extract=["tables", "metadata"],
    only_main_content=True,
)

print(op.render_quality)            # 0.0 to 1.0
print(op.status_code)               # HTTP status of the page itself
print(op.deductions)                # e.g. {"login_wall": 0.65}
print(op.outputs["markdown"].url)   # signed URL, 15 minutes
print(op.structured)

if (op.render_quality or 0) < 0.6:
    print("Low-confidence read:", op.warnings)
Option Type Default Description
outputs list[str] ["markdown", "structured"] Any of markdown, html_cleaned, html_raw, screenshot, screenshot_full_page, pdf, links, images, structured.
extract list[str] -- Any of tables, prices, contacts, metadata, main_content, headings, structured_data, technologies, all.
wait_for str -- CSS selector to wait for before capture.
viewport PerceiveViewport 1920 x 1080 width 320 to 3840, height 240 to 2160.
cache_mode "enabled" \| "bypass" \| "refresh" server default Reuse, skip, or rewrite the cached render.
block_resources list[str] -- Any of image, media, font, stylesheet, script, xhr, fetch, websocket, manifest, other.
only_main_content bool server default Strip navigation, headers, footers, and other page chrome from the extracted content.
direct_download bool server default Return raw bytes instead of a signed URL. Unset by default, so the key is omitted from the request and the server default (false) applies. Accepted by perceive only; the batch endpoint rejects it.

Also accepted: schema (dict, a free-form extraction schema passed through untouched), js_code (str, run in the page before capture), wait_timeout_ms (int), headers (dict[str, str]), cookies (list[BrowserCookie]), auth (HttpBasicAuth), proxy_url (str), geolocation (dict), action_chain (list[dict] of scripted click, type, and scroll steps), pdf_options (PdfOptions, applied to the pdf output), mobile (bool), and respect_robots (bool). perceive_batch takes the identical set apart from direct_download.

Artifact URLs are re-signed on every read, so call client.v2.get_perceive_operation(op.operation_id) for a fresh one instead of caching the string.

Raw bytes, no signed-URL round-trip. perceive_direct forces direct_download on and hands you the artifact itself, with metadata parsed from the response headers.

direct = client.v2.perceive_direct("https://example.com", outputs=["markdown"])
print(direct.filename, direct.content_type, len(direct.content), direct.render_quality)

raw = client.v2.download_perceive_artifact(op.operation_id, output="markdown")

perceive_direct requires exactly one artifact-producing output out of markdown, html_cleaned, html_raw, screenshot, screenshot_full_page, pdf, links, and images. structured may ride along but stays inline server-side; pass anything else and the SDK raises EnconvertError before sending the request. On download_perceive_artifact, output may be omitted when the operation produced exactly one artifact, and an artifact past its retention window answers 410.

Batches. Up to 1000 URLs share one options block. Small batches complete inline; larger ones come back with status "queued", so poll the job id.

batch = client.v2.perceive_batch(
    ["https://a.example.com", "https://b.example.com"], outputs=["markdown"], output_mode="zip"
)

done = client.v2.get_perceive_batch(batch.job_id)
print(done.status, done.completed, done.failed, done.pending)
for item in done.items:
    print(item.url, item.render_quality)

output_mode is "manifest" (default) or "zip"; when it is "zip", the finished bundle is on done.zip.url. Details in Perceive.

Discover#

Enumerate a site's URLs with no browser rendering, which makes it the cheap first step before you perceive anything.

found = client.v2.discover(
    "https://example.com", mode="hybrid", max_urls=200, exclude_patterns=["/tag/"]
)
print(found.total, found.truncated, found.sources)
Option Type Default Description
mode "sitemap" \| "crawl" \| "hybrid" server default Sitemap parsing, an HTTP crawl, or both merged and deduplicated.
max_urls / max_depth int server default Cap on returned URLs (truncated is True when more existed), and crawl depth from the seed.
same_domain_only bool server default Stay on the seed host.
respect_robots bool server default Honor robots.txt.

include_patterns and exclude_patterns (both list[str]) filter the result set. DiscoverResult.sources holds the raw per-source counts before dedup, such as {"sitemap": 42, "crawl": 30}. More in Discover.

Lookup#

Run a categorized web search, optionally rendering the top hits in the same call.

search = client.v2.lookup(
    "best static site generators", category="web", num_results=10, perceive_top=3
)

for hit in search.results:
    quality = hit.perceive.render_quality if hit.perceive else None
    print(hit.position, hit.title, hit.url, quality)
Option Type Default Description
category "web" \| "news" \| "images" \| "scholar" \| "patents" \| "maps" server default Search vertical.
time_filter "hour" \| "day" \| "week" \| "month" \| "year" -- Recency window.
num_results / page int server default Results per page, and the 1-based page number.
perceive_top int 0 Auto-render the top N result URLs. Each arrives with its full PerceiveResult on hit.perceive.

Also accepted: country (str), locale (str), location (str) for geo-sensitive queries, and autocorrect (bool). LookupResult carries answer_box, knowledge_graph, perceive_operation_ids, and perceive_top, the last reporting how many results were actually rendered, which can be lower than what you asked for. Reference in Lookup.

Distill#

Schema-driven structured extraction across one or many pages. Supply exactly one of urls or discover_from; schema is always required. Both rules are enforced client-side and raise EnconvertError before a request goes out.

from enconvert import CssField, CssSchema

extraction = client.v2.distill(
    urls=["https://example.com/pricing"],
    schema={"plans": "list of plan names with monthly prices"},
    css_schema=CssSchema(
        base_selector=".plan-card",
        fields=[
            CssField(name="name", type="text", selector="h3"),
            CssField(name="price", type="text", selector=".price"),
        ],
        target_field="plans",
    ),
)

for item in extraction.results:
    print(item.url, item.extraction_tier, item.fields_from_css, item.fields_from_llm, item.data)

The optional css_schema runs first and answers whatever the selectors can reach. Only the fields it misses escalate to the language-model tier, and extraction_tier on each item reports which path ran: css, llm, mixed, or none.

Discover and distill in one call:

from enconvert import DistillDiscoverFrom

client.v2.distill(
    discover_from=DistillDiscoverFrom(url="https://example.com", mode="sitemap", max_pages=10),
    schema={"title": "page title"},
)
Option Type Default Description
schema dict -- (required) Output shape you want back, passed through untouched.
urls list[str] -- Explicit page list. Mutually exclusive with discover_from.
discover_from DistillDiscoverFrom -- url, optional mode, optional max_pages (1 to 50, default 10).
css_schema CssSchema -- Selector pass run before any model call.

Also accepted: wait_for (str), wait_timeout_ms (int), headers (dict[str, str]), cookies (list[BrowserCookie]), and respect_robots (bool). CssField supports the types text, attribute, html, regex, nested, list, and nested_list, with optional attribute, pattern, default, transform (lowercase, uppercase, strip), and nested fields up to five levels deep. See Distill.

Ingest#

Turn a whole site, an explicit URL list, or a pile of uploaded documents into chunked, RAG-ready JSONL. Ingest is always asynchronous.

import time

from enconvert import IngestChunkOptions

job = client.v2.ingest(
    mode="sitemap",
    url="https://docs.example.com",
    max_pages=100,
    chunk=IngestChunkOptions(max_words=512, sentence_overlap=1),
    webhook_url="https://my.app/hooks/enconvert",
)

status = client.v2.get_ingest_job(job.job_id)
while status.status in ("queued", "discovering", "processing"):
    time.sleep(10)
    status = client.v2.get_ingest_job(job.job_id)

print(status.status, status.pages_processed, status.total_chunks, status.output_url)

Uploaded files run through the same job lifecycle under mode files. PDF, DOCX, PPTX, XLSX, CSV, HTML, EPUB, TXT and MD, plus legacy and ODF office formats are accepted, and each entry may be a path, an os.PathLike, raw bytes, or a FileData.

file_job = client.v2.ingest_files(["handbook.pdf", "notes.docx"])
Option Type Default Description
mode "urls" \| "sitemap" \| "crawl" \| "files" "urls" urls needs a non-empty urls list and rejects url. Every other mode needs a seed url and rejects urls.
url / urls str / list[str] -- Seed URL for sitemap and crawl, or the explicit page list for mode urls.
max_pages int server default Cap on pages ingested.
chunk IngestChunkOptions -- max_words 32 to 4000, default 512. sentence_overlap 0 to 10, default 1. Pair it with webhook_url (str), called once the job reaches a terminal state.

Crawl shaping and render options are accepted too: max_depth, same_domain_only, include_patterns, exclude_patterns, respect_robots, wait_for, and wait_timeout_ms. ingest_files takes only chunk and webhook_url.

for summary in client.v2.list_ingest_jobs(limit=20).jobs:
    print(summary.job_id, summary.status, summary.total_chunks)

client.v2.cancel_ingest_job(job.job_id)
client.v2.retry_ingest_webhook(job.job_id)

secret = client.v2.get_webhook_secret()
print(secret.signature_header, secret.signature_scheme, secret.replay_tolerance_seconds)
client.v2.rotate_webhook_secret()

cancel_ingest_job is idempotent; cancelling an already-terminal job returns it unchanged. retry_ingest_webhook answers 409 when the job is not completed and 400 when no webhook was configured. rotate_webhook_secret invalidates the previous secret immediately. Deeper coverage in Ingest.

Watch#

Re-render a URL on a fixed cadence and get told when it changes.

watcher = client.v2.create_watcher(
    "https://example.com/pricing",
    frequency_minutes=60,
    diff_mode="auto",
    webhook_url="https://my.app/hooks/changes",
    notify_email=True,
)

for snap in client.v2.get_watcher_snapshots(watcher.watcher_id, limit=10).snapshots:
    print(snap.checked_at, snap.has_changes, snap.similarity, snap.change_count)

client.v2.list_watchers(limit=20)
client.v2.update_watcher(watcher.watcher_id, status="paused")
client.v2.update_watcher(watcher.watcher_id, webhook_url="")
client.v2.delete_watcher(watcher.watcher_id)
Option Type Default Description
frequency_minutes int server default Check interval. Hourly is the floor.
diff_mode "auto" \| "text" \| "structured" \| "tables" \| "metadata" server default Which layer of the page the diff engine compares.
track_fields dict -- Named fields to track, passed through untouched.
webhook_url str -- Called on every detected change.
notify_email bool server default Send change notifications by email.

update_watcher needs at least one field and raises EnconvertError otherwise; status accepts "active" or "paused", and an empty webhook_url clears the webhook. delete_watcher is a soft, idempotent delete that returns the tombstoned watcher with status "deleted", after which the watcher reads as 404.

Snapshot diffs carry untrusted page content. The dicts in WatcherSnapshot.changes come straight from the watched site. Escape them before rendering into HTML, an email body, or a chat message.

Diff engine and notification payloads are documented in Watch.


PDF options#

PdfOptions is a frozen dataclass shared by convert_url_to_pdf, convert_website_to_pdf, convert_document, convert_to_pdf, and the v2.perceive family.

from enconvert import BrowserCookie, HttpBasicAuth, PdfHeaderFooter, PdfMargins, PdfOptions

client.convert_url_to_pdf(
    "https://internal.example.com/report",
    pdf_options=PdfOptions(
        page_size="A4",
        orientation="landscape",
        margins=PdfMargins(top=10, bottom=10, left=15, right=15),
        scale=0.9,
        header=PdfHeaderFooter(content="Quarterly Report", height=15),
        footer=PdfHeaderFooter(content="Confidential", height=12),
    ),
    auth=HttpBasicAuth(username="user", password="pass"),
    cookies=[BrowserCookie(name="session", value="abc123", domain="internal.example.com")],
    save_to="report.pdf",
)
Field Type Description
page_size str "A4", "A3", "Letter", "Legal", and friends.
page_width / page_height float Custom page geometry. Together they override page_size.
orientation "portrait" \| "landscape" Defaults to portrait.
margins PdfMargins top, bottom, left, right, all optional.
scale float Render scale, for example 0.9 for 90 percent.
grayscale bool Post-process the PDF to grayscale.
header PdfHeaderFooter content (up to 2000 characters) and height.
footer PdfHeaderFooter content (up to 2000 characters) and height.

BrowserCookie takes name, value, and either domain or url, plus optional path, expires, http_only, secure, and same_site ("Strict", "Lax", "None"). The SDK maps http_only and same_site onto their camelCase wire keys for you.


Error handling#

Errors are exception classes, so match them with except. Order handlers from specific to general: AuthenticationError, QuotaError, and RateLimitError all subclass APIError, which subclasses EnconvertError.

from enconvert import APIError, AuthenticationError, EnconvertError, QuotaError, RateLimitError

try:
    op = client.v2.perceive("https://example.com")
except AuthenticationError:
    print("Invalid or missing API key. Check ENCONVERT_API_KEY.")
except QuotaError as e:
    print(f"402 from the API: {e.message}")
except RateLimitError:
    print("Too many requests. Back off and retry.")
except APIError as e:
    print(f"API error [{e.status_code}]: {e.message}")
except EnconvertError as e:
    print(f"Rejected before the request was sent: {e}")
Class Raised on Status code
AuthenticationError Invalid, missing, or revoked key 401, 403
QuotaError HTTP 402 402
RateLimitError Too many requests 429
APIError Any other 4xx or 5xx the actual code
EnconvertError Base class, and client-side validation that never reaches the network --

APIError carries status_code and message, and its str() reads [404] Not found. Client-side raises you may hit: an empty api_key, an unsupported file extension, an unimplemented conversion pair, a distill call with both or neither of urls and discover_from, an ingest mode and argument mismatch, an empty update_watcher payload, and a perceive_direct output list that does not resolve to exactly one artifact. The full message map is in Error codes.


Timeout recovery#

Long URL and document conversions can outlive a reverse-proxy timeout even when the server finishes the job. The Python SDK recovers transparently:

  1. Before each single-file or single-URL conversion, the SDK generates a UUID4 hex string and sends it as job_id.
  2. If that request comes back 5xx, the SDK switches to polling GET /v1/convert/status/{job_id} every 3 seconds. A 404 there means the job row is not written yet, so polling continues.
  3. On success the SDK returns the result as if nothing happened; on failed it raises APIError(500, ...) with the server's error message; and past the 300-second polling deadline it raises APIError(504, "Conversion timed out").

You write no code for this. When a response arrives through the recovery path, ConversionResult.job_id is populated so you can correlate it in your logs.

Two carve-outs. convert_website_to_pdf and convert_website_to_screenshot opt out, because a website submission has no per-job row, so a 5xx there means the submission itself failed and surfaces directly. V2 endpoints are not polled either; use their own job ids with get_perceive_batch or get_ingest_job.


Configuration#

client = Enconvert(
    api_key=os.environ["ENCONVERT_API_KEY"], timeout=300.0, base_url="https://api.enconvert.com"
)
Option Type Default Description
api_key str -- (required) Private API key. An empty value raises EnconvertError at construction.
timeout float 300.0 Per-request timeout in seconds, applied to every call including uploads.
base_url str https://api.enconvert.com API base URL. Trailing slashes are stripped.

The client keeps a requests.Session, so connections pool across calls; reuse one client instead of building one per request. Your key travels as the X-API-Key header, and save_to downloads hit the signed storage URL as a plain unauthenticated GET streamed to disk in 64 KB chunks, so the key never leaves the API host.

Never hardcode the API key. Read it from an environment variable or your secret manager, and keep it server-side. Anyone holding your private key can run work against your project.

Result shape#

Conversion methods return a frozen ConversionResult:

@dataclass(frozen=True)
class ConversionResult:
    presigned_url: str
    object_key: str
    filename: str
    file_size: int | None = None
    conversion_time_seconds: float | None = None
    job_id: str | None = None

The presigned URL is short-lived. Pass save_to, or fetch the URL yourself, and store the bytes in your own bucket when you need durable access.

Type Returned by Key fields
JobStatus get_job_status status (processing, success, failed), presigned_url, object_key, error
BatchSubmission, BatchStatus convert_website_to_*, get_batch_status, wait_for_batch batch_id, status, url_count, total, completed, failed, discovery_method, zip_download_url, items
PerceiveResult v2.perceive, v2.get_perceive_operation operation_id, render_quality, status_code, deductions, outputs, structured, warnings, cache_hit
PerceiveDirectResult v2.perceive_direct, v2.download_perceive_artifact content (raw bytes), content_type, filename, render_quality, source_status_code
IngestJob v2.ingest, v2.ingest_files, v2.get_ingest_job job_id, status, mode, pages_processed, total_chunks, output_url, error_message
Watcher the v2 watch methods watcher_id, status, frequency_minutes, diff_mode, checks_count, next_check_at, last_change_at

V2 artifact URLs are signed for 15 minutes and re-signed each time you read the operation, so call get_perceive_operation rather than caching a URL string.


Source and issues#


Frequently asked questions#

How do I convert files in Python?#

Run pip install enconvert, build a client with Enconvert(api_key=os.environ["ENCONVERT_API_KEY"]), and call a typed method such as convert_document, convert_image, convert_to_pdf, or convert_to_markdown. Pass save_to="out.pdf" and the SDK streams the finished file straight to disk instead of handing you a URL to fetch yourself.

How do I convert a URL to PDF in Python?#

Call client.convert_url_to_pdf("https://example.com", save_to="page.pdf"). Set single_page=False plus pdf_options=PdfOptions(page_size="A4") for paginated output, and adjust viewport_width, load_media, or enable_scroll when a page needs a wider canvas or lazy-loaded images.

How do I convert DOCX to PDF in Python?#

Either client.convert_document("report.docx", save_to="report.pdf"), since output_format already defaults to "pdf", or client.convert_to_pdf("report.docx", save_to="report.pdf") when you want server-side format auto-detection. The same call handles XLSX, PPTX, ODT, ODS, ODP, OTS, Pages, and Numbers.

How do I convert HEIC to WebP in Python?#

client.convert_image("photo.heic", output_format="webp", save_to="photo.webp"). The input format comes from the filename extension, and jpeg, png, svg, heic, and webp all convert to each other. Working from memory instead of disk? Wrap the bytes in FileData(data=blob, filename="photo.heic") so the extension survives.

How do I scrape a web page into clean Markdown in Python?#

Use client.v2.perceive(url, outputs=["markdown"]) and read op.outputs["markdown"].url, or client.v2.perceive_direct(url, outputs=["markdown"]) to get the bytes back in the response body. Add only_main_content=True to drop navigation and footers. For a plain conversion with no scoring or extraction, convert_url_to_markdown is the lighter call.

What does render_quality mean, and when should I retry a page?#

It is a float from 0.0 to 1.0 on every V2 read that says how honestly the page rendered. A challenge screen, cookie wall, login gate, HTTP error page, or empty SPA shell scores low and names what fired in deductions, with detail in warnings and the page's own HTTP status in status_code. Gate on it: treat a low score as a signal to retry with cache_mode="refresh", a wait_for selector that only real content matches, or different cookies, rather than feeding the text to your model.

How do I turn a whole documentation site into RAG chunks in Python?#

Call client.v2.ingest(mode="sitemap", url="https://docs.example.com", max_pages=100, chunk=IngestChunkOptions(max_words=512, sentence_overlap=1)). The job is asynchronous, so either poll get_ingest_job(job_id) until the status leaves queued, discovering, and processing, or pass webhook_url and wait to be called. The finished JSONL sits at output_url. For local documents rather than a site, ingest_files runs the same pipeline.

Does the Python SDK support asyncio?#

Not natively. Every method is synchronous and built on requests. Inside an async application, wrap calls in await asyncio.to_thread(client.convert_to_pdf, "report.docx") or hand them to a concurrent.futures.ThreadPoolExecutor so the event loop keeps running. The client is safe to share across threads because it holds a pooled requests.Session.

How does the SDK survive a proxy timeout on a long conversion?#

Each single-file and single-URL conversion sends a generated job_id. If the request returns 5xx, the SDK polls GET /v1/convert/status/{job_id} every 3 seconds for up to 300 seconds, returns normally once the job reports success, raises APIError(500, ...) if it reports failed, and raises APIError(504, "Conversion timed out") if the deadline passes. Website batch submissions skip this path, since a failure there means the submission itself did not land.