← Managed inference

HunyuanOCR

Small OCR model that also spots text, extracts fields and translates

Overview

HunyuanOCR is Tencent's lightweight end-to-end OCR vision-language model, around 1.1B parameters, unifying document parsing, text spotting, information extraction and text-image translation in a single model. The checkpoint at the repository root is now HunyuanOCR-1.5, which extends maximum image resolution to 4K and context to 128K and adds long-tail coverage such as low-resource and ancient scripts; version 1.0 is archived in a subfolder.

Strengths

  • Four text-centric tasks in one small model
  • Handles pages up to 4K resolution
  • Text-image translation built in

Use cases

  • Structured field extraction from forms and receipts
  • High-volume document parsing on a small footprint
  • Translating text embedded in images

Good to know

Serves HunyuanOCR-1.5, the current release. Released under Tencent's Hunyuan Community terms, not a standard open licence.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

hunyuan-ocr · curl
# OCR models are called through the
# platform APIs and workflows; list what is deployed:
curl https://gateway.graphn.ai/v1/models \
  -H "Authorization: Bearer $GRAPHN_API_KEY"