Skip to main content
Polyvia: Multimodal Document Retrieval API. We’re releasing Polyvia 1, as two products:
  • Polyvia API: Multimodal Document Retrieval API (for developers of AI agents).
  • Polyvia Platform: Search & Exploration over multimodal docs (for knowledge workers in enterprises).
Agentic, file-by-file search (e.g. Claude Code, Claude Cowork, Codex) works only up to ~100 multimodal files — past that it’s too slow. So at scale, when you’re connecting an enterprise’s large internal datasets, you still need retrieval. And the multimodal infra tools today stop at visual extractors / PDF parsers (e.g. Reducto, LlamaIndex). We built Polyvia Engine — an end-to-end pipeline for multimodal document retrieval: VLM Visual Extractor → Multimodal Knowledge Ontology → Self-Improving Retrieval Agent. We index your unstructured & visual & multimodal docs (PDFs, charts, slides, complex tables, infographics, scans, handwriting, invoices, and more) into multimodal knowledge ontology, and provide you with a retrieval endpoint.

Start in 30 seconds

Get your API key in the Polyvia Platform — open API in the sidebar and click Create API Key. It’s shown only once, and all keys start with poly_.
Ingest a batch into a group, then ask one question across the whole corpus.
See also

Full quickstart

Get a key, ingest a batch, scope queries to a group — in a couple of minutes.

Python SDK

Typed client with sync + async, MCP and agent-tools support.

Polyvia 1

Polyvia 1 ships as two products — the Polyvia API and the Polyvia Platform.

Polyvia API

Multimodal Document Retrieval API, for developers of AI agents.

Polyvia Platform

Search & Exploration over multimodal docs, for knowledge workers in enterprises.
See also

Release notes

Faster ingestion, Office/Google formats, knowledge-graph view, and more — see what’s shipped.

Polyvia Engine

One pipeline turns scattered visual & multimodal files into a queryable knowledge layer — then answers in sub-200ms, grounded in a visual citation.

VLM Visual Extractor

Reads the hardest visual documents — charts, infographics, complex multi-page tables, slides, scans, handwriting, pictures — into structured facts.

Multimodal Knowledge Ontology

Disambiguates and connects every extracted fact into one semantic ontology over your whole corpus — a single, queryable source of truth.

Self-Improving Retrieval Agent

Agentic, multi-hop retrieval that self-improves over time. Every answer grounded in a visual citation tied to the exact source page.
See also

Core concepts

Documents, groups, ingestion, querying and citations.

Supported formats

Every modality and file type Polyvia ingests.

Integrations

Import from Drive, Dropbox, OneDrive, Notion, S3 and Slack.

Build with Polyvia

Python SDK

Typed client with sync + async, MCP and agent-tools support.

Polyvia API

Multimodal Document Retrieval API — for developers of AI agents. Ingest, groups, query and usage.

TypeScript SDK

Fully-typed ESM/CJS client for Node and modern JS.

MCP Server

Connect Polyvia to Claude, Cursor and other MCP tools.

Agent Skills

Drop-in skills for Claude Code, Cursor and agent environments.

Stay connected

Homepage

Blog