← Managed inference

DeepSeek OCR

Turns document images into markdown with layout preserved

Overview

DeepSeek-OCR is a roughly 3B vision-language model from DeepSeek's "Contexts Optical Compression" work, aimed at reading document images. A grounding prompt makes it emit markdown that keeps the page's structure, and it offers several resolution modes (Tiny through Large, plus a tiled Gundam mode) so you can trade vision tokens against fidelity on dense pages. It is supported natively in vLLM.

Strengths

  • Markdown output with layout grounding
  • Resolution modes trade cost against fidelity
  • MIT licensed

Use cases

  • PDF and scan to markdown conversion
  • Reading dense multi-column pages

Good to know

A conversion utility rather than an assistant -- it is deliberately kept out of the agent and chat model pickers.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

deepseek-ocr · curl
# OCR models are called through the
# platform APIs and workflows; list what is deployed:
curl https://gateway.graphn.ai/v1/models \
  -H "Authorization: Bearer $GRAPHN_API_KEY"