Changelog

New features, improvements, and fixes — everything we ship, as we ship it.

Subscribe via RSS
Clear Pick a year to narrow by month; a date range overrides both pickers.
Fixed Improved Security

Pages with strict CSP render again #

The renderer injects a small stylesheet to fix print colors and hide modals. That injection raced Chromium's CSP error stream and threw whenever any frame logged a violation, measured at 9 of 20 renders on a page whose own main frame carries no policy at all, and it failed the entire conversion over a cosmetic style. Injection now runs inside the page and a failure is logged and skipped. CSP is also disabled per page over the DevTools protocol, which covers policies delivered by that no response rewrite can reach. The request interceptor no longer refetches and buffers every image, script and stylesheet just to strip a header that only matters on the main document. That removes 100 to 200 MB of copying per render on media heavy pages, and subresource requests now reach the SSRF re-validation check that never saw them before.

Fixed

Renders killed by a timeout or a disconnect get cleaned up #

When a request hit the 300 second gateway timeout or the client went away, the page being rendered stayed alive and kept executing the site's JavaScript, so every later conversion competed with it. A cancelled render now queues cleanup that waits for the conversion slot to free: a healthy browser gets its leftover pages parked, a browser that fails its health probe gets relaunched. Browser shutdown is bounded at 10 seconds, so a wedged Chromium can no longer hang recovery or leave the browser manager half torn down and failing every request after it. The client still gets its 504 immediately.

New

The EnConvert CLI: the whole API from your terminal #

Summary

EnConvert now has an official command-line interface. enconvert is a single, open-source (MIT) binary that covers the whole API from your terminal: file conversion across 40+ formats, URL and website rendering to PDF, screenshot or markdown, and the full v2 web-data surface (perceive, discover, lookup, distill, ingest). It installs natively on macOS, Linux and Windows, needs nothing but your existing secret API key, and sends no telemetry. Nothing in the API itself changes.

What changed

  • One command for every conversion. enconvert convert report.docx --to pdf infers the endpoint from the input extension and target format across all 46 file-conversion routes (documents, spreadsheets, presentations, images, data formats, compression). Batch globs, -O output directories, --skip-existing and per-file progress are built in.
  • URL and website rendering. enconvert url pdf|screenshot|markdown <url> exposes every render option (viewport, selectors, cookies, headers, basic auth, ad blocking, PDF geometry) 1:1 with the API, and enconvert site pdf|screenshot drives the async website crawls with --wait polling and ZIP downloads.
  • The v2 web-data verbs, first class. perceive (including batches with 200/202 handling), discover, lookup (with enrichment and answer synthesis), distill (schema, prompt and CSS-schema modes) and the complete ingest family including file uploads and webhook-secret management.
  • Scripting-grade plumbing. Stable documented exit codes, --json with the gateway's raw response, a bundled --jq filter (no jq install required), --jsonl streaming, paths-on-stdout output rules, jobs wait <id> for any job kind, and enconvert api — a gh-style passthrough that reaches every endpoint, including ones without a typed command.
  • Native installs on every platform. Homebrew (brew install enconvert/tap/enconvert), Scoop, Winget, a checksum-verifying curl -fsSL https://get.enconvert.com/install.sh | sh, and npm i -g @enconvert/cli. Standalone binaries for macOS (Intel and Apple Silicon), Linux (glibc and musl, x64 and arm64) and Windows.
  • Profiles, config and safe credentials. ~/.config/enconvert/config.toml profiles, project-level .enconvertrc.toml, a credential_helper hook for 1Password/pass/Vault, 0600-permission key storage, and automatic migration of keys already saved by npx @enconvert/mcp setup.
  • No telemetry. The CLI makes no requests other than the API calls you ask for and an optional once-daily version check that a single environment variable disables.

Scope

This is a new client only. No /v1/* or /v2/* endpoint behavior changed. The CLI covers every working v1 endpoint and every v2 endpoint except /v2/watch (watchers remain dashboard-managed for now; typed commands for them will follow). Source: https://github.com/enconvert/cli.

Docs

A new CLI documentation page was added in all five languages (en, fr, de, es, it), the CLI joined the Integrations section of the homepage with a dedicated /integrations/cli page, and the SDKs section of the docs sidebar now lists the CLI alongside the Node.js SDK and MCP server.

New Improved

Image Compression and SVG Sizing #

Summary

Two additions to the V1 conversion API, both synchronous multipart/form-data file conversions gated only by the global plan conversion limit. POST /v1/convert/compress-image compresses PNG, JPEG and WebP files without ever changing their format — lossless-first, with an optional target_size_kb budget met by aspect-ratio-locked downscaling. The three SVG rasterization endpoints (svg-to-png, svg-to-jpeg, svg-to-webp) gain optional width and height parameters that control the output dimensions in pixels; the playground exposes matching size inputs with an aspect-ratio lock derived from the uploaded SVG.

POST /v1/convert/compress-image

Accepted extensions (4): .png .jpg .jpeg .webp — magic-byte checked. The output keeps the input's extension and format; a file whose content does not match its extension is rejected with 400 rather than silently converted. Animated inputs (APNG, animated WebP) are rejected with 400 rather than silently flattened to their first frame.

Request: file (required). target_size_kb (optional; a zero/negative value is a 400 rejected before quota is burned, a non-numeric value is a 422 from request validation). output_filename (default: input basename; the input's extension is preserved). job_id (optional, enables status polling). direct_download (accepted for parity with the other file endpoints; responses currently always take the JSON path below, same as every V1 file conversion).

Behavior — stage 1, lossless (always runs): metadata (EXIF, XMP, PNG text chunks) is stripped; the ICC color profile and the EXIF orientation flag are preserved — orientation is re-emitted as a minimal single-tag EXIF block instead of being baked into pixels, which keeps the JPEG path free of extra quantization loss. PNG re-encodes at zlib level 9 + optimize, plus a palette candidate accepted only when the palette roundtrip is provably pixel-identical. JPEG re-encodes reusing the original quantization tables (quality='keep') with optimized progressive Huffman coding. WebP re-encodes as true lossless VP8L at maximum effort. The smallest of the original bytes and all candidates wins, so the output is never larger than the input.

Behavior — stage 2, dimension reduction (only when target_size_kb is set and stage 1 missed it): LANCZOS downscale with the aspect ratio locked; the scale factor is binary-searched (up to 8 iterations, minimum scale 1%) for the largest dimensions that fit the budget. Downscaled JPEG and lossy-source WebP re-encode at quality 85; PNG and lossless-source WebP stay lossless at the reduced size (lossless vs lossy WebP sources are told apart by walking the RIFF chunk list for VP8L). An unreachable target returns the smallest file achieved, not an error — check file_size / X-File-Size to see what was reached.

Response: 200 with JSON — presigned_url, object_key, filename (timestamped, input extension preserved, e.g. photo_20260717_101530123.png), file_size, conversion_time_seconds, job_id. The same values are mirrored on the X-Object-Key, X-File-Size, X-Conversion-Time, X-Filename headers.

Constraints:

  • The decoded canvas is capped at 40,000,000 pixels (e.g. 8000x5000), checked from the image header before any pixels are decoded, so a decompression bomb is rejected with 400 without allocating the full surface.
  • Whole-request size is validated against the plan's max file size via Content-Length (413).
  • The shared conversion concurrency gate applies: at capacity the endpoint returns 503 with Retry-After: 10.
  • WebP lossless encoding uses maximum effort (method=6) up to 4 MP and drops to method=4 above it, bounding CPU on large canvases; all requests are still bounded by the 300s gateway timeout (504, then poll with job_id).
  • CMYK JPEGs are re-encoded in CMYK (no mode change); 16-bit PNGs skip the palette candidate and only get the plain lossless re-encode.

width / height on svg-to-png, svg-to-jpeg, svg-to-webp

Request: optional width and height integer form fields, 1–10000 each. One dimension alone scales the render proportionally — the other is derived from the SVG's own aspect ratio (CairoSVG native behavior). Both together set the exact canvas size, which may change the aspect ratio. Omitting both keeps the previous behavior (the SVG's width/height/viewBox attributes decide).

Constraints:

  • Total output is capped at 25,000,000 pixels (400). For single-dimension requests the derived dimension is estimated server-side from the root <svg> width/height attributes (absolute units only) or viewBox, so an extreme-ratio SVG cannot request an unbounded render surface. If the ratio cannot be determined from those (e.g. the SVG sizes itself in em/ex/%, which CairoSVG can still resolve into a large canvas), a single-dimension request is rejected with 400 and asked to supply both width and height — the size must then be explicit and self-bounding.
  • Validation runs before quota is burned or an activity row is logged; out-of-range dimensions are pure client errors.
  • The converters' signatures stay backward-compatible — existing calls without width/height are byte-for-byte unaffected.

Playground: the three SVG conversions now show width/height inputs whose placeholders display the uploaded SVG's intrinsic size, plus a Lock aspect ratio toggle (on by default). The lock reads the SVG's width/height attributes client-side, falling back to the viewBox, and disables itself when neither yields a usable ratio. Blank fields mean "the SVG's own size". The controls are localized in all five languages.

Docs

Every affected page was updated in all five languages (en, fr, de, es, it): the three SVG endpoint pages now document width/height (parameter table, resolution notes, FAQs), a new compress-image endpoint page follows the standard skeleton, and the image-conversions category page, endpoints overview, parameters reference ("Image Options" section) and endpoint counts (48 → 49) were updated accordingly.

Improved New

Render Engine Ladder Added #

Summary

The V2 rendering endpoints — perceive, distill, ingest, and watch — now render each URL through an automatic multi-engine fallback instead of a single headless-browser pass. A page is first fetched over a real-browser TLS fingerprint; if that path is blocked or the page needs JavaScript, the request escalates to a headless-Chrome render, and a page that still looks blocked by anti-bot protection escalates once more to a stealth-hardened render. The result: more real-world pages return usable content, and simple pages come back faster. Nothing in your request changes — this is automatic, and every existing call behaves the same or better.

What changed

  • TLS-first fast path. Static and server-rendered pages are fetched over a real-browser TLS/HTTP-2 fingerprint with no headless browser involved. These pages return faster and free browser capacity for the pages that genuinely need it.
  • Automatic browser fallback. If the fast path is blocked, hits a bot wall, or the page renders empty (a client-side JavaScript app), the request transparently escalates to the headless-Chrome render — the same capture pipeline (cookie-banner dismissal, lazy-load scrolling, sticky-header handling, image waiting) as before.
  • Stealth escalation for anti-bot pages. A render that still looks blocked by anti-bot protection is retried once with additional browser-fingerprint hardening, recovering pages a plain render could not reach.
  • Same signals, same shape. The render_quality score and the blocked-page warning still tell a real render from a challenge page, and no response field changes. When every engine is still blocked, the best attempt is returned and flagged — exactly as before.

Scope

This applies to the V2 rendering endpoints only: perceive, distill, ingest, and watch. The V1 url-to-pdf, url-to-screenshot, and url-to-markdown endpoints are unchanged.

Improved Security

V2 Competitive Hardening #

Summary

A round of V2 improvements closing the biggest gaps against dedicated web-data APIs: durable batch perception that survives restarts, high-value search enrichment (concurrent multi-format rendering + a synthesized cited answer + structured extraction across results), prompt-only structured extraction, more capable site discovery (gzipped sitemaps, WAF-resilient fetches, a far higher crawl ceiling), and a security-hardening pass (DNS-rebind protection on the render path, an opt-in local URL threat policy, and CSP respected by default).


/v2/perceive/batch — durable, resumable batches + cancel

Batches are now restart-safe. Previously the batch envelope (the shared render options and output mode) lived only in memory, so a server restart mid-batch failed the whole batch and you had to resubmit. The envelope is now persisted, so an interrupted batch resumes automatically and re-renders only the URLs that had not finished — already-completed URLs keep their artifacts.

  • New: DELETE /v2/perceive/batch/{job_id} — cancel a running batch. The worker stops between URLs; already-completed URLs remain available. Idempotent; a terminal batch is unchanged. Project-scoped 404.
  • New batch status: canceled (joins queued|processing|completed|partial|failed).
  • The batch status endpoint now reports the true output_mode and the canceled state from the durable batch record.

No request-shape change for submitting a batch — existing calls keep working.


/v2/lookup — high-value enrichment (enrich)

A new optional enrich object turns "search, then read the top results" from a slow, markdown-only, one-at-a-time pass into a fast, configurable one — and can synthesize a grounded answer across the results.

enrich fields: - outputs[] — which perceive outputs to produce per enriched result (e.g. markdown, html_cleaned, links, screenshot, structured). Defaults to ["markdown"]. - concurrency (1–5, default 3) — how many result URLs to enrich in parallel. Markdown/HTML renders use the no-browser path and truly parallelize; screenshot/PDF renders serialize on the shared browser. - schema (aliased) — run schema-driven structured extraction against each enriched result; the data appears under each result's perceive.structured. - synthesize_answer (default false) — synthesize one cited, grounded answer to the query across the enriched results, returned on the response as answer (with answer_sources). Uses the perceived page content when available, otherwise result snippets. - answer_prompt — an optional question to answer instead of the raw query.

New response fields: answer, answer_sources[]. When enrich is omitted, perceive_top keeps its previous behavior (markdown-only, sequential). Structured extraction and answer synthesis require an LLM-enabled plan; on plans without it they degrade to a warning.


/v2/distill — prompt-only extraction

You can now distill without writing a schema. Supply a natural-language prompt and the extraction schema is synthesized from it (single model), then the normal two-pass CSS→LLM engine runs.

  • schema is now optional; provide either schema or prompt (schema wins if both are given).
  • New response field synthesized_schema echoes the fields that were derived from the prompt.
  • Prompt-only mode requires an LLM-enabled plan; without one it returns a clear warning.

/v2/discover — gzipped sitemaps, WAF resilience, higher ceiling

  • Gzipped sitemaps (sitemap.xml.gz and any Content-Type: application/gzip sitemap) are now decompressed and parsed instead of silently failing.
  • WAF resilience: sitemap and robots fetches now send realistic browser headers, so sites that 403/503 a bare client are far more likely to return their sitemap.
  • Higher crawl ceiling: crawl-mode HTTP fetches are no longer capped at 50. The ceiling now equals your max_urls (up to 1000), so large sites are enumerated far more completely. Operators can lower it with DISCOVER_CRAWL_MAX_PAGES if needed. This also lifts the same ceiling for /v2/ingest crawl-mode jobs.

Security hardening

  • DNS-rebind / redirect protection on the render path: every request the browser makes (the initial navigation, redirects, and subresources) is re-validated against the SSRF rules at request time — not only the seed URL before navigation — so a hostname that rebinds to a private address, or a redirect to an internal one, is blocked before the browser connects. The no-browser TLS engine additionally pins each connection to the validated IP.
  • Local URL threat policy (opt-in, no external dependency): a configurable denylist of domains, TLDs, and host patterns enforced on every fetched URL, with an audit trail of blocks. Configure via THREAT_BLOCKED_DOMAINS, THREAT_BLOCKED_TLDS, or a THREAT_POLICY_FILE. Empty by default (blocks nothing).
  • Content-Security-Policy respected by default for the raw browser-context path (previously bypassed). Callers that need CSP bypass for scripted DOM interaction opt in explicitly.

These changes require no API changes on your side; blocked URLs return 400.

New Security Improved

Hardened URL rendering: SSRF screening, ad/media blocking, selector waits, and clearer errors #

Summary

The URL rendering endpoints — url-to-pdf, url-to-screenshot, url-to-markdown, website-to-pdf, and website-to-screenshot — get new rendering controls, stronger URL-safety screening, and clearer error responses. Existing requests are unaffected: every new option defaults to its previous behavior.

New rendering options

  • wait_for_selector (+ wait_for_selector_timeout, default 10000 ms, max 60000): wait for a CSS selector before capture so single-page apps that hydrate after load are captured fully. Returns 422 if the element never appears.
  • block_ads: abort requests to known ad and tracker domains so they never load, render, or slow the capture.
  • block_media: abort image and audio/video requests entirely for a faster, lighter render (distinct from load_media, which only controls waiting).

Stronger URL safety (SSRF)

Every target URL is now screened before the browser fetches it. Requests to private, loopback, link-local, reserved, or cloud-metadata addresses — plus embedded-credential URLs, non-http(s) schemes, and non-standard IP notations — are rejected with 400. This applies to single URLs, every URL in a batch, and pages discovered by the website-to-* crawl endpoints.

Credential scoping

An Authorization header (for example a Bearer token) passed via headers, or credentials from the auth object, are now sent only to the target origin — never to the third-party ad, analytics, or CDN subresources a page requests.

Clearer error responses

URL conversions now tell a target/input problem apart from an engine problem, using a structured {error, code, detail} body:

  • 415 unsupported_content_type — a non-HTML response such as JSON sent to url-to-pdf/url-to-screenshot (use url-to-markdown for JSON).
  • 422 selector_not_found — a wait_for_selector that never appeared.
  • 502 upstream_unreachable / empty_render — the target could not be reached, or rendered nothing.
  • 504 upstream_timeout — the target took too long.
  • 503 with a Retry-After header when the render pool is momentarily at capacity.

A 500 now specifically means our engine faulted, rather than a catch-all.

Non-HTML URLs

url-to-markdown now returns application/json and text/plain bodies verbatim inside a fenced code block instead of running article extraction over the browser's viewer.

Fixed Security

LibreOffice conversion timeouts now return 504 instead of 500 or 400 #

A LibreOffice/unoserver conversion that exceeds its 120-second subprocess budget now returns 504. The same timeout previously surfaced as 500 on the PDF path and 400 on the markdown path.

The 120-second budget itself is unchanged — only the status, the message, and the analytics fields changed. Nothing got faster and nothing new times out.

Endpoints moving 500 → 504:

Endpoint Before After
POST /v1/convert/doc-to-pdf 500 504
POST /v1/convert/excel-to-pdf 500 504
POST /v1/convert/ppt-to-pdf 500 504
POST /v1/convert/odt-to-pdf 500 504
POST /v1/convert/ods-to-pdf 500 504
POST /v1/convert/odp-to-pdf 500 504
POST /v1/convert/ots-to-pdf 500 504
POST /v1/convert/pages-to-pdf 500 504
POST /v1/convert/numbers-to-pdf 500 504
POST /v1/convert/anything-to-pdf — for .doc .docx .xls .xlsx .ppt .pptx .odt .ods .odp .ots .pages .numbers .rtf .csv input 500 504

Endpoint moving 400 → 504:

Endpoint Before After
POST /v1/convert/anything-to-markdown — for .doc .ppt .xls .odt .ods .odp .rtf input only 400 504

The two paths' format sets differ. .docx, .pptx, .xlsx and .csv go through LibreOffice on the PDF path and can time out at 504 there — on the markdown path they use pure-Python readers, never reach unoserver, and are unaffected, as are .pdf and .epub.

Response: {"detail": "Document conversion timed out after 120 seconds."} — detail is a plain string, identical text from both paths. Unlike the adjacent 500, it carries no Conversion failed: prefix.

Breaking — retry semantics on POST /v1/convert/anything-to-markdown. Moving 400 → 504 reclassifies these responses from client error to server error. A client that treats 4xx as terminal and 5xx as retryable will now retry a timed-out legacy-office markdown conversion instead of failing it outright. That is the intent — the document was valid and the request was not the caller's fault — but it is a change in your request volume against these endpoints and in what your own users see on failure. If you retry 5xx with a backoff policy, check it before deploying against large .doc/.xls/.odt inputs.

Breaking — message text. anything-to-markdown previously returned "Document conversion timed out.". The text is now "Document conversion timed out after 120 seconds." A client string-matching the old message breaks.

Fixed — the previous 500 leaked server internals. On the PDF path the subprocess timeout propagated uncaught and its string form was interpolated into the response body, so callers received the full unoconvert argv and both server-side temporary file paths — e.g. Conversion failed: Command '['unoconvert', '--convert-to', 'pdf', '/tmp/tmpXXXX.docx', '/tmp/tmpXXXX.pdf']' timed out after 120 seconds. The 504 body carries the fixed message and nothing else.

Analytics. Failure events for these timeouts now carry error_type="conversion_timeout" and error_code=504. They previously reported error_type="TimeoutExpired" / error_code=500 on the PDF path and error_type="value_error" / error_code=400 on the markdown path. Dashboards or alerts filtering on those old values will stop seeing these events.

Scope. v1 conversion endpoints only. POST /v2/ingest/files reaches the same LibreOffice path for legacy-office input but does not use the v1 error mapping — a timeout there is still recorded as a failed page on the job, not returned as a 504. Unsupported input formats continue to return 400.

Security Fixed

Filename and title normalization in JSONL deliverable metadata #

Uploaded filenames flowed verbatim into the metadata.source_url and metadata.title fields of the JSONL deliverable produced by POST /v2/ingest/files. The same path also covered the URL side of ingest, where metadata.title is the <title> lifted from an arbitrary remote page.

This was not a JSON injection. JSON encoding already neutralizes quotes, newlines and C0 controls inside the file — the deliverable always parsed. The problems were what a downstream consumer does with the decoded value:

  • Line framing. U+2028 and U+2029 survive non-ASCII JSON encoding and are emitted raw, and Python's str.splitlines() treats them as line breaks — a label containing one splits a record for any line-oriented reader.
  • Display spoofing. Bidi override characters let invoice<U+202E>fdp.exe render as invoice.pdf.
  • Path prefixes. ../../../../etc/passwd.csv reached the label intact, carrying a traversal prefix into the deliverable.
  • Index bloat. The label had no length cap and is duplicated into every chunk record of the page.

Both labels are now normalized at emit time. Stripped: bidi formatting controls (U+200E, U+200F, U+202A–U+202E, U+2066–U+2069), line/paragraph separators (U+2028, U+2029), and all Unicode category Cc — which covers C0 controls, DEL, C1 controls and U+0085 — except \t. Whitespace runs collapse to a single space; the value is NFC-normalized first, so a decomposed name like café.pdf now emits as the composed café.pdf.

Kept deliberately. Spaces and punctuation survive — Q3 Report (final) v2.pdf is emitted unchanged. U+200C (ZWNJ) and U+200D (ZWJ) are not stripped: the filter targets Cc and the specific bidi/separator codepoints, not all of category Cf, because ZWNJ/ZWJ are required in legitimate Devanagari/Indic and emoji filenames. Clean URLs pass through untouched.

Caps. metadata.title is capped at 512 characters. metadata.source_url is capped at 2048 — generous enough that a real URL is never truncated, since truncating source_url destroys URL-page provenance. For file pages the label is built at the 512 cap, so a file page's source_url is effectively 512-capped.

Path stripping is file-only. An uploaded filename is reduced to its basename across both / and \ before normalization, which removes traversal prefixes without losing legitimate information. URL-sourced source_url is not basenamed — a URL keeps its full path.

Two other file-label changes. A file page whose filename is empty — or which normalizes to empty, such as a bidi-only name or a trailing-separator path like a/b/ — previously emitted metadata.title: "" and fell back to the internal storage object key for metadata.source_url, leaking an internal key into the deliverable. Both fields now emit upload in that case.

New 400 conditions on POST /v2/ingest/files:

  • Filename longer than 255 characters → 400 {"detail": "Filename exceeds the 255-character limit."}. The filename is deliberately not echoed back.
  • Filename containing a null byte → 400 {"detail": "Filename contains an invalid null byte."}. Previously an opaque 500 — Postgres TEXT cannot hold NUL and the insert failed deep in the write path.

Both checks run before extension parsing. The pre-existing content-type and unsupported-extension 400s are unchanged and still echo the filename.

Scope and limitations — read these:

  • A file page's source_url is a display label, not an identifier. It is not unique across files, and basenaming means two uploads sharing a basename — including two different paths — now emit an identical source_url. Do not key deduplication or provenance on it for file pages.
  • Normalization is emit-time only. The uploaded filename is still stored raw in the ingest page record for audit. Any consumer of that value — admin UI, support tooling, log readers — is unprotected and must escape for itself.
  • Chunk content is not sanitized. Only metadata.source_url and metadata.title pass through the normalizer. The document body is emitted as extracted.
  • Filename gates are POST /v2/ingest/files only. POST /v1/convert/anything-to-markdown and the other v1 convert endpoints gain no length or null-byte gate.
  • The storage object key is derived by a separate, more aggressive filename sanitizer and is unaffected by this change.
Fixed

Chunk ids for uploaded files on POST /v2/ingest/files no longer collide for same-named files #

Chunk ids in the assembled JSONL from POST /v2/ingest/files were derived from the uploaded filename. Two files sharing a filename in one job produced byte-identical ids across both pages' chunks — a downstream upsert keyed on id silently overwrote one document with the other, and no error was raised. Ids are now seeded from the page's storage object key, which is unique per file per job.

The id format is unchanged: 12 lowercase hex characters, a hyphen, a 4-digit zero-padded chunk index. Only the seed changed.

Breaking: every file-page chunk id changes. The seed moves from the filename (report.pdf) to the storage object key (prod/files/42/v2-ingest-uploads/ing_abc_1_report.pdf) — a different input to the same hash, so a completely different id. There is no overlap with the old ids and no migration path. Consequences:

  • Re-ingesting a file corpus yields new ids. An upsert keyed on id will insert duplicates rather than replace. Delete the old vectors first, or re-index.
  • The object key embeds the job id, so the same file uploaded in a later job gets different ids. File-page ids are stable within one job, not across jobs.
  • A file-page id is not reproducible from the deliverable. The seed never appears in the JSONL (metadata.source_url carries the filename label), and the source object is deleted at assembly.
  • A file-mode job that is in flight across the deploy assembles a mixed-id JSONL — pages completed before the deploy keep filename-seeded ids, pages completed after get key-seeded ids. Per-page JSONL is concatenated verbatim at assembly; nothing rewrites it. Drain in-flight file-mode jobs before deploying.

URL-mode ids are byte-identical to before. For URL pages the seed and the previous derivation are the same value, and the seed is deliberately hashed raw. A re-ingested URL corpus produces the same ids and existing vector-store rows upsert in place — no re-index. This covers ids only: the metadata-label sanitization shipping alongside it can still change metadata.title on URL pages, since that title is lifted from the remote page's HTML.

Determinism is preserved for both modes: the seed is a stored column, so a resumed or re-run page within a job re-derives the same id.