Overview
HunyuanOCR is Tencent's lightweight end-to-end OCR vision-language model, around 1.1B parameters, unifying document parsing, text spotting, information extraction and text-image translation in a single model. The checkpoint at the repository root is now HunyuanOCR-1.5, which extends maximum image resolution to 4K and context to 128K and adds long-tail coverage such as low-resource and ancient scripts; version 1.0 is archived in a subfolder.
Strengths
- Four text-centric tasks in one small model
- Handles pages up to 4K resolution
- Text-image translation built in
Use cases
- Structured field extraction from forms and receipts
- High-volume document parsing on a small footprint
- Translating text embedded in images
Good to know
Serves HunyuanOCR-1.5, the current release. Released under Tencent's Hunyuan Community terms, not a standard open licence.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
# OCR models are called through the # platform APIs and workflows; list what is deployed: curl https://gateway.graphn.ai/v1/models \ -H "Authorization: Bearer $GRAPHN_API_KEY"
graphn model get hunyuan-ocr
@function
async def read_document(image_url: str) -> str:
return await vision.extract_text(image_url)