Overview
DeepSeek-OCR is a roughly 3B vision-language model from DeepSeek's "Contexts Optical Compression" work, aimed at reading document images. A grounding prompt makes it emit markdown that keeps the page's structure, and it offers several resolution modes (Tiny through Large, plus a tiled Gundam mode) so you can trade vision tokens against fidelity on dense pages. It is supported natively in vLLM.
Strengths
- Markdown output with layout grounding
- Resolution modes trade cost against fidelity
- MIT licensed
Use cases
- PDF and scan to markdown conversion
- Reading dense multi-column pages
Good to know
A conversion utility rather than an assistant -- it is deliberately kept out of the agent and chat model pickers.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
# OCR models are called through the # platform APIs and workflows; list what is deployed: curl https://gateway.graphn.ai/v1/models \ -H "Authorization: Bearer $GRAPHN_API_KEY"
graphn model get deepseek-ocr
@function
async def read_document(image_url: str) -> str:
return await vision.extract_text(image_url)