Changelog

New features, improvements, and fixes — everything we ship, as we ship it.

Subscribe via RSS
Clear Pick a year to narrow by month; a date range overrides both pickers.
Fixed

POST /v2/ingest/files orphaned staged uploads on rejection and cancel #

POST /v2/ingest/files staged uploads one at a time, interleaved with per-file validation — file N was written to object storage before file N+1 was gated. A gate failure on any file after the first rejected the whole request, but the files already staged stayed in storage: the request never created a job row, so nothing referenced those objects and nothing removed them. Separately, a job canceled before assembly returned without deleting its staged sources — only the assembled-successfully path cleaned up.

Staging is now two passes. Every file in the batch is validated first; only then is any object written. A rejected batch stages nothing, so there is no orphan window to clean up. Peak memory is unchanged — one file's bytes at a time in both passes.

Every staged object is registered for a timed cleanup sweep at stage time, with a TTL of 24 hours. This is a backstop, not the prompt path: it covers a worker crash and any terminal path that never reaches the eager delete. Registration is idempotent and never fails a submit. If staging itself faults partway (a storage or database error), the partial batch is discarded on a best-effort basis immediately rather than waiting on the sweep; a discard failure is logged and never masks the underlying error.

Source cleanup now runs on every terminal path — assembled, all pages failed or skipped, all completed pages lost before assembly, canceled during the page loop, and canceled at assembly. Only the assembled path did this before. Each delete is now individually guarded: previously the cleanup loop was unguarded, so a storage error on page 3 of 10 propagated out and left pages 4–10 orphaned. A cleanup failure is logged and never sinks an otherwise-finished job. Cleanup is gated on source type — URL-mode pages are never touched.

Object-key collision on path-like filenames. The per-file {job}_{index}_ uniqueness prefix was applied before filename sanitization, and sanitization takes the basename — so for a filename containing a path separator (../../etc/passwd.csv) the prefix was stripped off again and two such uploads in one job collided onto a single key, one silently overwriting the other. The filename is sanitized before the prefix is applied. For filenames without a separator the resulting key is byte-identical to before.

Two new filename gates, both 400. A filename longer than 255 characters is now rejected with Filename exceeds the 255-character limit. — the filename is deliberately not echoed back. A filename containing a null byte is rejected with Filename contains an invalid null byte.; this previously surfaced as an opaque 500, so it is a 500 → 400 status change on an existing condition. Both are checked before the extension and content-type gates, which are unchanged and still echo the filename.

Residual — a job canceled in the queued window relies on the backstop. Eager cleanup runs where cancellation is observed: inside the page loop and at assembly. A job canceled before the worker begins that loop, or one whose job row is already gone when cleanup would run, has no eager path — its staged sources are removed by the 24-hour sweep, not immediately.

Constraints: scoped to POST /v2/ingest/files. No request or response fields changed and no schema change — the backstop writes rows through the existing retention path. Status codes are unchanged for every existing condition except the null-byte filename noted above.

Security Fixed

Upload size limit on anything-to-pdf enforced against received bytes, not Content-Length #

POST /v1/convert/anything-to-pdf enforced the upload size ceiling by reading the content-length request header. When the header was absent — a chunked or HTTP/2 client — the check was skipped in its entirety, not merely relaxed: an over-ceiling file was accepted and converted. The ceiling is now measured against the bytes actually received.

Scope: this changed POST /v1/convert/anything-to-pdf and nothing else. Every other upload route that calls the shared size check still measures content-length and still performs no size check at all when the header is absent. The defect class is fixed at one endpoint, not across the API.

What the 413 does and does not do: the multipart body is fully parsed — and spooled, to disk above 1 MB — before the handler runs, so the size check happens after receipt. It prevents the oversized file being loaded and converted; it does not abort the upload. An oversized body still crosses the wire and still occupies the spool. Do not treat the 413 as a bandwidth or ingress control.

Behavior changes on anything-to-pdf:

  • Chunked/HTTP2 request with no content-length, file over the ceiling — was 200, now 413.
  • Request with content-length where the file is under the ceiling but the multipart envelope is over it — was 413, now 200. The old check counted boundaries and other form fields against the file's budget; the new one measures the file. This is a relaxation, and near-limit uploads that previously failed will now go through.
  • detail.file_size now reports the file's byte count rather than the envelope's.

The 413 body shape, the default ceiling applied when the subscription carries none, and the rejection analytics event are unchanged.

Separately, a non-numeric content-length on the fallback path previously surfaced as a 500; it is now ignored.

Fixed

pdf_options no longer silently discarded on anything-to-pdf and the office PDF endpoints #

POST /v1/convert/anything-to-pdf and the nine dedicated office conversion endpoints accepted and validated a pdf_options payload, then discarded everything except grayscale — page geometry never reached the converter. Geometry is now honoured wherever the engine can apply it, and rejected with a 400 wherever it cannot. No endpoint validates an option it then throws away.

Geometry now applied on anything-to-pdf. The geometry fields — page_size, page_width, page_height, orientation, margins, scale, header, footer — are honoured for HTML (.html, .htm, .xhtml), Markdown (.md, .markdown, .mdown, .mkd), plain text (.txt, .text), .epub, raster images (.png, .jpg, .jpeg, .gif, .bmp, .tiff, .tif, .webp, .heic, .heif), and .svg.

Geometry now rejected with 400 on anything-to-pdf. For office/ODF/iWork/RTF/CSV input (.doc, .docx, .xls, .xlsx, .ppt, .pptx, .odt, .ods, .odp, .ots, .pages, .numbers, .rtf, .csv) and .pdf passthrough, page layout comes from the source document. Previously these options were silently dropped and the request returned 200 with the source document's layout. An explicitly-set geometry option now returns 400 before any conversion runs:

pdf_options ['page_size', 'scale'] cannot be applied to '.docx' input: page geometry for this format is determined by the source document (LibreOffice), not by the converter. Remove these options (only 'grayscale' is supported for '.docx'), or set the page layout in the source file before uploading.

detail is a plain string. The engine clause reads LibreOffice for office input and PDF passthrough for .pdf.

Constraints: the trigger is explicitness, not value. A field counts as set if it appears in the request body at all — sending {"page_size": "A4"} (the default value) on a .docx returns 400, while {"grayscale": true} returns 200. Unsupported extensions still fail as unsupported formats, unchanged.

grayscale is unaffected on every input family, including office and .pdf. It was never broken and is applied as a post-process regardless of input format.


Behaviour change: geometry options rejected on the dedicated office endpoints

The same rejection now applies to POST /v1/convert/doc-to-pdf, excel-to-pdf, ppt-to-pdf, odt-to-pdf, ods-to-pdf, odp-to-pdf, ots-to-pdf, pages-to-pdf and numbers-to-pdf. All nine render through LibreOffice, which takes no page-geometry arguments — page size, orientation and margins are properties of the source document's page style, applied at layout time before PDF export. These endpoints previously accepted and validated the geometry fields, then dropped them and returned 200 with the source document's own layout.

This is breaking for callers who send geometry to these endpoints. Such a request returns 400 before conversion instead of 200, with the same message shape shown above. The options never had any effect, so the PDF a caller receives is unchanged for every request that does not set geometry — but a caller that has been sending page_size and ignoring the fact that it did nothing must now remove it.

grayscale is unchanged on all nine and still applies. Omitting pdf_options entirely is unchanged. A wrong-extension upload (for example .html sent to doc-to-pdf) still fails with the unsupported-format error first — the format check takes precedence over the options check.


Behaviour change: scale now takes effect on html-to-pdf and markdown-to-pdf

POST /v1/convert/html-to-pdf and POST /v1/convert/markdown-to-pdf already honoured page_size, orientation, margins, header and footer. scale was parsed and range-checked, then dropped — nothing consumed it. It is now applied.

This changes the output of existing requests. A caller that has been sending scale on either endpoint has been receiving unscaled PDFs and will now receive scaled ones. Callers relying on the old no-op must remove scale from the request. The accepted range is unchanged: 0.1 to 2.0 inclusive, default 1.0; scale: 1.0 remains a no-op.

Semantics: page geometry is fixed, content scales. An A4 request at scale: 2.0 returns an A4 page with content rendered at 2×, not an A2 page. This matches the scale semantics already in effect on the url-to-pdf path. The same semantics apply to scale on anything-to-pdf for every geometry-capable input family listed above.


Other observable effects

  • Requesting geometry on an image or SVG changes the rendering engine. Raster input moves from the direct image path to an HTML layout pass, with a PNG re-encode of the raster; .svg moves from CairoSVG to the same layout pass (still vector). Output is not byte- or pixel-comparable to a no-options run on the same file. A request with no geometry options is byte-identical to before on both paths.
  • SVG with geometry: external resources are hard-blocked. Neither path fetches external resources, but on the geometry path an .svg that references one now fails with a 400 — SVG to PDF conversion failed: External resources are not fetched: <url>.
  • .txt default margin. Plain text with no geometry option (including a grayscale-only request) keeps its hardcoded 2cm margin. Sending any geometry option drops that default so the supplied margins win.
  • header / footer content containing a double quote or backslash now renders literally on anything-to-pdf, html-to-pdf and markdown-to-pdf. On html-to-pdf and markdown-to-pdf — the two endpoints that honoured header/footer before this change — such content terminated the generated CSS string early, breaking the render or injecting caller-supplied rules into the document. That is fixed. On anything-to-pdf the fields were never applied before, so the escaping is in place from the first request that reaches it.

No new request or response fields. pdf_options is unchanged as a schema.

New

New endpoint: POST /v1/convert/anything-to-markdown #

Summary

POST /v1/convert/anything-to-markdown accepts an uploaded document and returns a single UTF-8 Markdown file, routing to a per-format extractor by the uploaded filename's extension. 22 extensions are accepted. Conversion is synchronous and in-process — no job queue, no polling requirement. This is a v1 conversion route and is not part of the /v2 surface; it authenticates via the existing X-API-Key / Authorization: Bearer mechanism used by the other /v1/convert/* routes.

POST /v1/convert/anything-to-markdown

Accepted extensions, by dispatch group:

  • Native extractors — .pdf, .docx, .pptx, .xlsx, .csv, .epub
  • Markup — .html, .htm, .xhtml
  • Legacy / ODF office — .doc, .ppt, .xls, .odt, .ods, .odp, .rtf — converted through headless LibreOffice to HTML first, then to Markdown. Fidelity for this group is LibreOffice's HTML export, not a native reader.
  • Plain text / Markdown — .txt, .text, .md, .markdown, .mdown, .mkd — encoding-normalised passthrough (tries utf-8-sig, utf-8, cp1252, latin-1, then lossy utf-8).

Known catalog gap — client-side filtering. The public endpoint catalog that drives the dashboard playground's file picker currently declares only 17 of these 22. .xhtml, .text, .markdown, .mdown and .mkd are accepted by the API but may be filtered out by the picker's accept= list until the catalog is refreshed. Direct API calls with those extensions succeed today. The catalog is a strict subset of what the code accepts, so no advertised extension is rejected — the failure mode is only that five working extensions are hidden.

Request: multipart/form-data. file (required). output_filename (default null — falls back to the input basename). job_id (optional, client-supplied; echoed back in the response and readable afterwards via GET /v1/convert/status/{job_id} if the connection drops mid-request). direct_download (accepted, default true, but inert on this route — the response is always the JSON body below; the parameter exists for request-shape parity with the *-to-pdf routes). Unlike html-to-pdf / markdown-to-pdf / anything-to-pdf, this route takes no pdf_options.

Response: 200 with JSON containing a pre-signed URL — never raw bytes, never a server-minted job handle.

{
  "presigned_url": "…",
  "object_key": "…",
  "filename": "report_20260715_142233123.md",
  "file_size": 18432,
  "conversion_time_seconds": 1.84,
  "job_id": null
}

The same values are mirrored on X-Object-Key, X-File-Size, X-Conversion-Time and X-Filename headers. The .md is written to object storage before the response returns; output name is {base}_{YYYYmmdd_HHMMSSfff}.md. job_id echoes the caller's value (null if none was sent) — it does not indicate async processing.

Constraints: request timeout 300 s → 504 {"error": "Request timeout"}. The legacy/ODF/RTF conversion subprocess is capped at 120 s, and both its failure and its timeout surface as 400, not 504. Upload size is checked from Content-Length before the body is read → 413. PDFs are capped at 2000 pages (400 above that). ZIP-backed formats (.docx, .pptx, .xlsx, .epub) reject archives declaring more than 400 MB uncompressed or more than 10000 entries → 400. Concurrency is bounded in-process; a full pending queue returns 503 "Server is at capacity. Please retry shortly." with Retry-After: 10.

Silent truncation points — these do not error, they quietly shorten output: PDF words beyond 20000 per page are dropped; PDF pages carrying more than 4000 vector edges skip table detection entirely, so tables on those pages are simply not extracted; table cells are truncated at 500 chars (… appended); table body rows are capped at 5000 (_(table truncated to 5000 rows)_ appended); EPUB chapter text is bounded at 100 MB total, and exceeding it stops the chapter loop.

Fidelity

Tables — preserved as GFM pipe tables across PDF, .xlsx, .csv, .pptx and .docx. Cells are flattened to a single line (\r/\n → space) and | is escaped to \|. Spreadsheet and CSV values are read as strings, so leading zeros (007), long integers (no scientific notation) and trailing-zero decimals (1.50) survive verbatim.

Headings — preserved as real ATX, with a per-format mechanism:

  • .docx — Word Heading 1–6 styles map to #–######.
  • PDF — inferred from font size, not from structure. The most common rounded size is treated as body text; larger sizes are ranked and mapped largest → #, and the result is capped at ###. A line is only promoted if it is ≤14 words, ≥3 characters and contains a letter. Prose lines beginning with # are escaped to \# so they are not misread as headings.
  • .pptx — synthetic ## Slide N: Title per slide, plus ### Notes for speaker notes.
  • .xlsx — synthetic ## SheetName per sheet; an empty sheet emits _(empty sheet)_.

Images — effectively dropped. Do not rely on them. There is no asset extraction and nothing is uploaded. An <img> becomes a Markdown image link to its original src, which for file input is relative and unresolvable by the caller; an <img> with no src is removed. PDF extraction is word- and table-only, so images never appear at all. .pptx picture shapes are skipped. One exception: .docx embedded images arrive as base64 data: URIs and pass straight through as ![](data:image/png;base64,…). Nothing strips them — expect substantially inflated output for image-heavy Word documents.

OCR — not performed. Scanned or image-only PDFs raise rather than returning an empty file: 400 "No extractable text found in the PDF. Scanned or image-only PDFs require OCR, which this endpoint does not perform." Image uploads (.png, .jpg, .jpeg, .gif, .bmp, .webp, .tiff, .tif) are not in the allowlist and are rejected with 400 by the extension gate, carrying the standard invalid-format message below.

PDF layout — multi-column PDFs read column-interleaved; this is a known limitation of the v1 extractor. Running headers/footers repeating on ≥60% of pages (minimum 3 pages) are dropped. Words inside a detected table region are removed from the prose stream so content is not duplicated. Consecutive prose lines merge into a paragraph, breaking on a heading, a table, or a vertical gap greater than 1.7× the previous font size.

HTML, EPUB, DOCX and legacy office — whole-document and faithful. Only <script>, <style>, <noscript>, <iframe>, <svg>, <canvas> and <template> are stripped — <nav>, <footer>, <aside>, <form> and <button> are kept. style/class/id and all on* handlers are removed; XML processing instructions are removed rather than leaking as literal text.

EPUB — chapters in spine (reading) order. The EPUB3 nav document is skipped so the table of contents does not pollute the output. Only .xhtml/.html/.htm spine items are read, and XML declaring a DTD or entity is rejected outright.

Links and code — anchor-only (#…), javascript: and empty-href links are unwrapped to plain text; bare URLs become autolinks (<url>); relative links stay relative. Fenced code blocks carry a language hint when one can be sniffed from class="language-*" or data-lang.

Output — always UTF-8. Runs of three or more blank lines collapse to one, CRLF/CR normalise to LF, and the file ends with exactly one newline. No YAML frontmatter.

Errors

400 — extension outside the allowlist, checked before any extractor runs: "Invalid file format '{ext}' for anything-to-markdown. Allowed: {list}". This is also what image uploads receive. Also returned for a filename with no extension; for corrupt, empty or text-free documents ("The document contains no extractable text.", "Invalid EPUB: the file is not a valid ZIP archive.", "The CSV file contains no rows.", "Archive rejected: decompressed content is too large.", "Invalid EPUB: XML declaring a DTD/entity is not allowed."); for a scanned PDF; and for legacy-office conversion failure or its 120 s timeout ("Document conversion timed out.").

400 — content/extension mismatch. Magic bytes are checked for .pdf, .docx, .xlsx, .pptx, .odt, .ods, .odp, .epub, .doc, .xls and .ppt; a high-confidence mismatch returns "File content does not match the 'anything-to-markdown' input type." .rtf, .csv, .html, .txt and .md are not sniffed at all.

413 — Content-Length above the permitted request size. 503 — capacity gate, with Retry-After: 10. 504 — the 300 s request timeout. 500 — anything else, as "Conversion failed: {message}"; extractor-library exceptions surface here rather than as a typed error.

Choosing between anything-to-markdown and url-to-markdown

Separate routes, separate implementations, different inputs. POST /v1/convert/url-to-markdown is unchanged by this release.

POST /v1/convert/anything-to-markdown POST /v1/convert/url-to-markdown
Input uploaded file (multipart/form-data) JSON body with URL(s) — no file input
Rendering no browser; per-format extractor headless browser render, sees JS-injected content
Modes synchronous only sync, async and batch (ZIP)
Extra params none viewport, scroll, cookies, auth, headers
HTML handling whole document, faithful main-article extraction — nav/footer/aside dropped
Output body Markdown, no frontmatter Markdown with YAML frontmatter (title, description, url, links, images)

The two share only their HTML-to-Markdown conversion. Uploaded files sent to /v2/ingest run through the same converter as this route; /v1/convert/anything-to-markdown is the standalone one-file-in, one-file-out form of it.

New

New endpoint: POST /v1/convert/anything-to-pdf #

Summary

POST /v1/convert/anything-to-pdf accepts 36 input file extensions and returns a PDF. The engine is selected from the file extension — callers do not specify a source format. Synchronous multipart/form-data; file upload only, no URL input.

POST /v1/convert/anything-to-pdf

Accepted extensions (36):

Family Extensions Engine
Office .doc .docx .xls .xlsx .ppt .pptx LibreOffice headless
OpenDocument .odt .ods .odp .ots LibreOffice headless
Apple iWork .pages .numbers LibreOffice headless
Other document .rtf .csv LibreOffice headless
Markup .html .htm .xhtml WeasyPrint
Markdown .md .markdown .mdown .mkd Markdown → HTML → WeasyPrint
Plain text .txt .text WeasyPrint — HTML-escaped, wrapped in a monospace <pre>, 2cm page margins
Ebook .epub text extraction → Markdown → WeasyPrint
Raster image .png .jpg .jpeg .gif .bmp .tiff .tif .webp .heic .heif Pillow, embedded at 100 DPI
Vector image .svg CairoSVG — true vector output, not rasterized
PDF .pdf validated passthrough

For the LibreOffice families, conversion fidelity is bounded by LibreOffice's — relevant in particular for iWork and legacy binary Office formats.

Request: file (required). output_filename (default: input basename). job_id (default null, used only for timeout-recovery polling). pdf_options (JSON string). direct_download (default true).

pdf_options is parsed and validated, but only grayscale has any effect on this endpoint — it is applied to the finished PDF. page_size, orientation, margins, scale, header, and footer are accepted and validated (an invalid page_size returns 400), but are never passed to the conversion engine and have no effect on the output. Do not rely on them here. direct_download is likewise accepted for request-shape parity with other convert routes and has no effect — the response is always JSON.

Response: 200 with JSON — presigned_url, object_key, filename (timestamped, e.g. report_20260714_101530123.pdf), file_size, conversion_time_seconds, job_id. The same values are mirrored on the X-Object-Key, X-File-Size, X-Conversion-Time, and X-Filename headers. This is a completed conversion, not a queued job.

Constraints:

  • Routing is by extension alone. A file with no extension returns 400. Magic bytes never select an engine — content checking is a veto that runs before dispatch and rejects only a high-confidence mismatch between a declared binary type and recognizably different bytes (400, "File content does not match the 'anything-to-pdf' input type."). It fails open: unrecognized signatures pass.
  • 21 of the 36 extensions are byte-checked (.doc .docx .epub .gif .heic .heif .jpeg .jpg .numbers .odp .ods .odt .ots .pages .pdf .png .ppt .pptx .webp .xls .xlsx). The other 15 are not, and .bmp, .tif, and .tiff are binary formats that are currently not sniffed despite having well-known signatures. Practical consequence: a .docx that is really a PNG is caught by the veto (400 content mismatch), but a .rtf that is really a PNG is not — it passes to the LibreOffice engine and surfaces as a 400 conversion error instead.
  • Animated GIF and multi-page TIFF convert first frame only — a multi-frame raster always yields a one-page PDF. Transparency (RGBA/LA/palette-alpha) is flattened onto white.
  • No page-count limit. Images beyond Pillow's decompression-bomb threshold return 400 ("Image is too large to process safely.").
  • Max upload size is read from the caller's subscription and enforced against the content-length header, not against bytes read. A request that omits content-length is not size-checked. Oversized returns 413.
  • Concurrent conversions are capped; requests beyond the pending queue depth return 503 with Retry-After: 10.
  • Global request timeout is 300s → 504 {"error": "Request timeout"}. The LibreOffice subprocess has its own 120s cap, which surfaces as 500, not 504 — a slow Office/iWork/RTF/CSV conversion fails with Conversion failed: ... before the global timeout is reached. The WeasyPrint, Pillow, and CairoSVG paths have no per-conversion timeout and are bounded only by the 300s global.
  • Corrupt or malformed input (broken image, invalid SVG, malformed EPUB, a .pdf that isn't a PDF, LibreOffice failure) returns 400 with the underlying reason. Unsupported extension returns 400 with the full allowed list in detail. Malformed pdf_options JSON returns 400.
New

File uploads: POST /v2/ingest/files #

/v2/ingest previously accepted only URLs — an explicit list, a declared sitemap, or a crawl. It now also accepts uploaded documents, via a new multipart route. File jobs reuse the URL path's job model, chunker, assembled JSONL deliverable, and signed webhook; the mode field is widened to "urls" | "sitemap" | "crawl" | "files", and every existing /v2/ingest route operates on file jobs unchanged.

POST /v2/ingest/files

Takes N uploaded documents, extracts text from each, chunks each with the same semantic chunker applied to crawled pages, and assembles one JSONL containing every chunk from every upload. Returns 202 with the same ing_-prefixed job_id and the same response body as the JSON path. No browser process is spun up — file jobs are CPU-bound extraction, not renders, and are not subject to browser serialization.

Request: multipart/form-data only — no JSON, no base64. Documents go in files (repeated part, named files, not files[]). Remaining fields are form fields, not a body model: max_words (default 512), sentence_overlap (default 1), webhook_url (optional).

Accepted extensions (allowlist on extension — the declared part Content-Type is not consulted): .txt .text .md .markdown .mdown .mkd .html .htm .xhtml .pdf .docx .pptx .xlsx .csv .epub .doc .ppt .xls .odt .ods .odp .rtf. The last seven are not handled by a native extractor — they are converted through an unoserver/LibreOffice subprocess first. Images are rejected at the extension gate with 400; there is no OCR path.

No mode field — the route is hard-wired to mode="files". None of the URL-path fields exist here (max_pages, max_depth, same_domain_only, include_patterns, exclude_patterns, respect_robots, wait_for, wait_timeout_ms). Because the route takes discrete form fields rather than a strict body model, unknown fields are silently ignored instead of returning 422 — a misspelled max_words falls back to the default with no error.

Response (IngestJobResponse, unchanged): job_id, status, mode, pages_discovered, pages_processed, pages_failed, total_chunks, output_url (signed URL to the assembled JSONL, present once completed), error_message, webhook_url, webhook_delivered, created_at, completed_at, warnings. For file jobs pages_discovered is the uploaded file count, and the job moves queued → processing — it never enters discovering, since pages exist at submit. Existing routes apply as-is: GET /v2/ingest/{job_id}, DELETE /v2/ingest/{job_id}, POST /v2/ingest/{job_id}/retry-webhook, GET /v2/ingest (summaries render mode: "files"), GET /v2/ingest/webhook-secret, POST /v2/ingest/webhook-secret/rotate.

Output: chunk records are the same shape as URL jobs — id, content, metadata{source_url, title, headings_path, section, word_count, chunk_index}. For file chunks, metadata.source_url and metadata.title carry the original filename rather than a URL: same key, different semantics per mode. There is no per-chunk marker distinguishing a file chunk from a URL chunk — job mode is the only discriminator. id derives from a hash of metadata.source_url, so two uploads sharing a filename in one request produce colliding chunk ids in the assembled JSONL; both files still process. Filenames are emitted as supplied — unsanitized in metadata; sanitization applies to storage keys only.

Constraints: 200 files per request maximum (400 above it) — the JSON path's 1000-page ceiling does not apply here. Empty file → 400. Per-file byte ceiling is read from the caller's account configuration; over it → 413. No aggregate request-size cap exists — 200 files each at the per-file ceiling is accepted. max_words clamps to 32–4000 and sentence_overlap to 0–10: out-of-range values are silently clamped, not rejected with 422 as on the JSON path (max_words=999999 becomes 4000). max_words remains a soft cap — code blocks and tables stay atomic and may exceed it.

Submit is not instant. Each file is read, magic-byte checked, and staged to storage one at a time inside the request before the 202 returns, so submit wall-clock scales with file count × size and the 300 s request window applies to the upload. Content sniffing runs after the extension gate and rejects only when the bytes resolve to a recognizably different known binary type — indeterminate bytes pass. A single unsupported, empty, oversized, or mismatched file rejects the whole request: no job is created and no file in the batch is processed, but files already staged before the failing one are left orphaned in storage — no job row exists to clean them up. Retrying a rejected batch re-uploads everything.

No conversion timeout except 120 s on the unoserver-routed formats — .pdf, .docx, .pptx, .xlsx, .csv, .epub, text, and HTML extraction is uncapped, and there is no job-level deadline. Uploaded source files are deleted once the JSONL is assembled; that cleanup runs at the tail of assembly only, so a job canceled before it reaches assembly leaves its uploads in storage. webhook_url is scheme-checked at submit (http:// or https://, else 400) and SSRF-screened at delivery; a file job fetches no other URL, so nothing else is screened.

Activity rows for file jobs record zero input bytes and log the same /v2/ingest endpoint as the JSON path, so the two are indistinguishable in activity history.

New Improved

New v2 API surface: perceive, discover, lookup, distill, ingest, and watch endpoints #

Summary

Six new endpoint groups (20 routes total) ship under the /v2 prefix: perceive, discover, lookup, distill, ingest, and watch. These cover headless-browser page capture, site URL enumeration, live web search, schema-driven structured extraction, async URL-to-JSONL ingestion, and scheduled page-change monitoring. All routes authenticate via the existing X-API-Key / Authorization: Bearer mechanism used elsewhere in the API. Rate limiting applies to billable POST routes; GET status-polling routes are exempt. There is no shared response envelope — each endpoint defines its own response model; check status/error fields per-endpoint rather than assuming a common shape.


/v2/perceive

  • POST /v2/perceive — single-URL headless-browser render producing any combination of requested outputs in one page load: markdown, markdown_fit, html_cleaned, html_raw, screenshot, screenshot_full_page, pdf, links, images, structured (via outputs[]).
  • GET /v2/perceive/{operation_id} — poll/fetch a single operation.
  • POST /v2/perceive/batch — same flow fanned out over urls[] (max 1000, de-duplicated) under one shared options block.
  • GET /v2/perceive/batch/{job_id} — poll aggregate batch status and per-URL results.

Request fields: url (max 2048 chars, http(s):// only), outputs[] (default ["markdown","structured"]), extract[] (tables, metadata, main_content, headings, structured_data implemented; prices, contacts, technologies, all are accepted by the schema but degrade to a warning — not implemented), extraction_schema (aliased schema), wait_for (css:/js: expression), wait_timeout_ms (0–60000, default 30000), js_code (max 20000 chars), viewport, headers/cookies/auth, cache_mode (enabled|bypass|refresh, default enabled), pdf_options, block_resources[], respect_robots (default false), mobile (default false). proxy_url, geolocation, action_chain are accepted by the schema but return 422 — not implemented.

Response (PerceiveResponse): operation_id, status (queued|processing|completed|failed), url/url_final, content_hash, render_quality (0.0–1.0), cache_hit, outputs{} (each a pre-signed URL artifact, default 900s expiry), structured, extraction_tier (heuristic|css|llm), tokens{input,output}, duration_ms, error, warnings[]. Structured extraction always runs a heuristic pass (metadata/JSON-LD/headings/tables/main content); the LLM tier fires only when a schema is supplied, the render isn't flagged bot-blocked, and heuristic fields are unfilled.

Batch specifics: output_mode (manifest|zip, default manifest) — zip bundles all successful artifacts into one archive, exposed via a zip field on PerceiveBatchResponse. Batches process strictly sequentially through a single in-process worker — no concurrency. Batches of ≤10 URLs attempt an inline response, waiting up to 240s before degrading to 202 + job_id; batches of >10 URLs always return 202 immediately. PerceiveBatchResponse fields: job_id, status (queued|processing|completed|failed|partial), output_mode, total/completed/failed/pending, zip, items[] (PerceiveResponse[]), warnings[]. The batch queue is in-memory only — a server restart mid-batch marks remaining rows failed ("interrupted by server restart"); batches must be resubmitted.

Constraints: main_content extraction truncated to 50,000 chars; every URL is SSRF-screened before any row is created; respect_robots=true adds a robots.txt check (403 if disallowed); pdf output always wins hook-chain selection over screenshot when both are requested; cache keyed by a fingerprint of render-affecting fields (URL, outputs, extract list, schema, pdf_options, viewport, mobile, js_code, wait config, block_resources, headers, sorted cookies, auth), 1-hour window.


/v2/discover

  • POST /v2/discover — enumerates a site's URLs without rendering (no browser process spun up). Three modes: sitemap (robots.txt-declared sitemaps, probed sitemap.xml/<sitemapindex> recursion, RSS/Atom feeds), crawl (HTTP-only breadth-first crawl harvesting <a href> links from raw HTML — cannot see JS-injected routes on client-rendered SPAs, by design), hybrid (default — union of both, deduplicated).

Request: url (max 2048), mode (default hybrid), max_urls (1–1000, default 100), max_depth (1–5, default 2, crawl mode only), include_patterns[]/exclude_patterns[] (regex, max 50 entries each, compiled at request-validation time — malformed pattern returns 422), same_domain_only (default true), respect_robots (default false; filters the output list post-hoc rather than gating fetches).

Response (DiscoverResponse): url, mode, total, urls[], pages_crawled, truncated, robots_respected, sources{} (raw per-source counts before dedup/filter), warnings[].

Constraints: crawl mode hard-caps actual HTTP fetches at min(max_urls, 50) regardless of max_urls; seed URL and every BFS-followed link are SSRF-screened independently; synchronous request/response — no operation row, no job queue, no persisted artifact. Handler catches all internal errors and returns a generic 500 to avoid leaking library/path details.


/v2/lookup

  • POST /v2/lookup — live web search proxy (Serper/Google SERP) across six categories. No caching and no retrieval against previously ingested/watched content — each call is a fresh outbound query. Optionally auto-renders the top-N result URLs through the /v2/perceive flow (markdown-only) in the same round trip via perceive_top.

Request: query (1–512 chars), category (web|news|images|scholar|patents|maps, default web), country/locale/time_filter/location, num_results (1–100, default 10), page (1–10, default 1), autocorrect (default true), perceive_top (0–10, default 0).

Response (LookupResponse): lookup_id, echoed query/category/filters, total, results[] (title, url, snippet, position, source, date, image_url, thumbnail_url, extra{}, optional perceive), perceive_top, perceive_operation_ids[], answer_box, knowledge_graph, warnings[].

Constraints: 15s timeout per attempt, up to 3 attempts with 0.5s/1.0s backoff on 429/500/502/503/504; 401/403 are not retried; a shared circuit breaker gates the provider — an open breaker returns 503 immediately without attempting the call; auto-perceive runs sequentially, not in parallel, and degrades to a per-result warning rather than failing the whole search on a single render failure; result order is passed through from the provider unmodified (no re-ranking); raw provider error bodies are never surfaced to the client (mapped to generic 502/503).


/v2/distill

  • POST /v2/distill — schema-driven structured extraction from one or more URLs via a two-pass engine. Pass 1: a free, no-LLM CSS selector extraction against a caller-supplied css_schema. Pass 2: escalates only fields left empty by pass 1 to an LLM extractor (claude-haiku-4-5) with a reduced schema covering just those fields. Results merge and normalize to exactly the caller's schema (missing fields become null/[]).

Request: urls[] (max 50 entries, mutually exclusive with discover_from) or discover_from ({url, mode, max_pages: 1–50}), schema (aliased extraction_schema, required, max 200 top-level properties, JSON-Schema or flat {field: description} form), css_schema (baseSelector, fields[] — name/type(text|attribute|html|regex|nested|list|nested_list)/selector/attribute/pattern/transform, nested recursion depth max 5), wait_for (max 1024 chars), wait_timeout_ms (0–60000, default 30000), headers/cookies, respect_robots (default false). Regex patterns in css_schema are compiled at request time and rejected on a nested-quantifier ReDoS heuristic.

Response (DistillResponse): operation_id, total/completed/failed, results[] (url/url_final, status (completed|failed), data, extraction_tier (css|llm|mixed|none), fields_from_css, fields_from_llm, render_quality, tokens{input,output}, error, warnings[]), warnings[].

Constraints: 50-URL hard ceiling per request; each URL renders sequentially through a shared browser instance (not concurrent); CSS pass is wall-clock bounded at 10s and falls through to the LLM pass with a warning on timeout; LLM pass is skipped (CSS-only fallback, no error) when the render is flagged bot-blocked or a per-request escalation limit is reached (at most one LLM call per URL); LLM request timeout is 60s; SSRF/robots protection is inherited per-URL from the underlying render step — a rejected render fails only that URL (status: "failed") without aborting the rest of the batch.


/v2/ingest

  • POST /v2/ingest — creates an ingestion job.
  • GET /v2/ingest — paginated job list.
  • GET /v2/ingest/{job_id} — job status/progress.
  • DELETE /v2/ingest/{job_id} — cancel (idempotent).
  • POST /v2/ingest/{job_id}/retry-webhook — manual webhook redelivery.
  • GET /v2/ingest/webhook-secret / POST /v2/ingest/webhook-secret/rotate — signing-secret management.

Discovers and/or renders a URL or set of URLs (mode: urls explicit list, or sitemap/crawl — seed URL expanded via the same crawler used by /v2/discover, up to max_pages), renders each page (credential-free — no auth/cookies/custom headers accepted on any URL-mode request), converts HTML to Markdown, and splits it with a heading-aware chunker. On completion, all pages' chunks are concatenated into a single JSONL deliverable (LangChain JSONLoader / LlamaIndex SimpleDirectoryReader / vector-DB bulk-import compatible), returned as a signed download URL. Always asynchronous — POST /v2/ingest returns 202.

Request (IngestRequest, extra="forbid" — unknown fields reject the request): mode (default urls), url (seed, required for sitemap/crawl, max 2048 chars), urls[] (required for urls mode, max 1000 entries, each max 2048 chars), max_pages (1–1000, default 50), max_depth (1–5, default 2), same_domain_only (default true), include_patterns[]/exclude_patterns[] (max 50 each, validated at request time), respect_robots (default false), wait_for (max 1024 chars), wait_timeout_ms (0–60000, default 30000), chunk: {max_words, sentence_overlap} (both bounded), webhook_url (max 2048 chars, http(s):// only at submit).

Response (IngestJobResponse): job_id, status (queued|discovering|processing|completed|failed|canceled), mode, pages_discovered/pages_processed/pages_failed/total_chunks, output_url (present only once status == "completed"), error_message, webhook_url, webhook_delivered, created_at/completed_at, warnings[]. GET /v2/ingest paginates via skip/limit (default 20, max 100) + has_more. DELETE cancellation is observed between page renders — the worker stops without assembling output; terminal jobs are unchanged by a repeat call.

Webhooks: fired automatically on completion if webhook_url is set; best-effort — delivery failure never fails the job. SSRF-screened at delivery time (not at submit time). Signed with a per-project HMAC-SHA256 secret over "<unix_ts>.<raw body>", sent as X-Enconvert-Signature: sha256=<hex hmac> and X-Enconvert-Timestamp: <unix seconds> (timestamp bound into the MAC for replay protection; default consumer-side freshness tolerance 300s). Payload: {job_id, status, output_url, pages_processed, total_chunks}. Retries at 1.0s, 4.0s, 16.0s backoff (4 attempts worst case), each retry re-signed with a fresh timestamp; any 2xx counts as success. POST /v2/ingest/{job_id}/retry-webhook returns 400 if no webhook is configured or the target now resolves to a private address, 409 if the job isn't completed.


/v2/watch

  • POST /v2/watch (201) — create a watcher.
  • GET /v2/watch — list watchers.
  • GET /v2/watch/{watcher_id} — get one.
  • GET /v2/watch/{watcher_id}/snapshots — capture history.
  • PATCH /v2/watch/{watcher_id} — update.
  • DELETE /v2/watch/{watcher_id} — soft delete (idempotent).

Schedules recurring headless-browser captures of a URL and diffs each capture against the prior one, persisting change records and optionally delivering a signed webhook and/or email when a change is detected.

Request (POST, extra="forbid"): url (max 2048, http(s):// only), frequency_minutes (60–43,200, default 60), diff_mode (auto|text|structured|tables|metadata, default auto), track_fields (optional field/selector subset), webhook_url (max 2048, http(s):// only), notify_email (default true). PATCH rejects an all-null body with 422 (no silent no-op).

Response (WatcherResponse): watcher_id (wat_<uuid4hex>), url, status (active|paused|deleted), frequency_minutes, diff_mode, track_fields, webhook_url, notify_email, consecutive_errors, checks_count, last_check_at, next_check_at, last_change_at, created_at/updated_at. GET /v2/watch list returns a leaner WatcherSummary (same minus diff_mode/track_fields/webhook_url/notify_email/updated_at), limit clamped to 1–100.

Diffing: four strategies run under diff_mode=auto — text diffing (similarity-ratio comparison on main-content text, flagged changed below 0.98 similarity, one capped unified diff), structured-list diffing (key-matched add/remove/modify for links by href, JSON-LD structured_data, prices, contacts), table diffing (matched by heading/caption/header signature/position), and metadata diffing (key-by-key). Each Change carries section, kind (added|removed|modified), key, field, before/after — string values >2000 chars truncated; before/after are documented as untrusted page content requiring HTML-escaping by any consumer. Capped at 500 changes per diff (overflow appends a summary record with the true count). A render flagged blocked or scoring below the quality floor is persisted as an audit-only snapshot — no diff, no baseline eligibility, no notification. GET /v2/watch/{watcher_id}/snapshots returns newest-first, limit clamped 1–100 at the handler (200 hard cap in the store).

Webhooks: same signing scheme as /v2/ingest (X-Enconvert-Signature/X-Enconvert-Timestamp, HMAC-SHA256 over "<ts>.<body>") — the signing secret is shared per-project across /v2/ingest and /v2/watch deliveries and is re-screened for SSRF immediately before each send (independent of the scheme-only check performed at create/PATCH time). Retries at 1.0s/4.0s/16.0s, each re-signed with a fresh timestamp; any 2xx is success; a dead endpoint is logged as a non-delivery, never raised. Payload: {event: "change_detected", watcher_id, url, checked_at, similarity, change_count, changes}.

Constraints: frequency_minutes floor of 60 is enforced redundantly at schema validation, flow computation, and poller-claim time; ceiling 43,200 minutes (30 days); 3 consecutive check failures auto-pauses a watcher (status="paused", next_check_at cleared, optional owner email); the poller claims due watchers in batches of ≤50; prior-snapshot reads are scoped to a project-namespaced storage path as a cross-project isolation guard independent of write-time checks.