Qwen3.8-27B vs Claude Opus 4.6 on real document work
Published
August 15, 2026
Reading time
8 min
We tested the launch claim with a production workflow instead of a benchmark harness: same pages, same prompt, same scoring, and public list prices for both models.
When Qwen3.8-27B launched on GraphN, Qwen's own comparison table put it against Anthropic's Claude Opus 4.6 and reported mixed results on document and vision benchmarks. Those are vendor-run numbers with no independent reproduction. Document extraction is also the most commercially deployed vision workload, which makes the claim worth checking against work shaped like production.
So we checked it: a production-shaped document workflow, both models, same prompt, same pages, same scoring.
The workflow
Document to Ledger is a GraphN workflow that turns document images into audit-ready structured data. One vision-native model reads the raw pages, no OCR stage, and returns vendor, dates, line items, and totals as strict JSON. The function then verifies the arithmetic (line items against subtotal, subtotal plus tax against total) so a human only reviews documents that flag. Structured extraction from documents is currently the most commercially deployed vision-model workload, which is why we picked it.
Here it is as it looks in the product. Switch tabs for the canvas, the workflow YAML, and a replay of the run:
graphn.ai/ws_cbe340ec9d0f
GraphN/Personal/Getting Started ▾/Document to LedgerLive v1 ▾0 ▰▱▱ / 1,000LE
</> EditDelete
model qwen3.8-27b · input document_urls · verifies totals arithmetic
The function behind that single step is the entire implementation:
python
import base64
import json
import re
import httpx
from foundry_helpers import vision
EXTRACT_PROMPT ="""Extract the financial data from the attached document image(s).
Respond with ONLY a JSON object, no prose, exactly this shape:
{"vendor_name": str, "invoice_number": str, "invoice_date": "YYYY-MM-DD",
"due_date": "YYYY-MM-DD" or null, "currency": ISO-4217 str, "po_number": str or null,
"line_item_count": int, "line_items": [{"description": str, "quantity": num, "unit_price": num, "amount": num}],
"subtotal": num, "tax": num, "total": num}
Rules: numbers as plain JSON numbers with a dot decimal separator regardless of the document's locale.
Amounts in parentheses are negative. line_items excludes subtotal/tax/total rows. The document may span multiple pages."""@functionasyncdefextract_ledger(document_urls:list, model:str="qwen3.8-27b",**kwargs)->dict:"""Read document pages with a vision model and return a verified ledger.""" pages =[]asyncwith httpx.AsyncClient(timeout=60)as client:for url in document_urls: resp =await client.get(url) resp.raise_for_status() pages.append(base64.b64encode(resp.content).decode()) raw =await vision.analyze( pages, EXTRACT_PROMPT, model=model, temperature=0.0, max_tokens=4000) text = re.sub(r"^```(json)?|```
quot;,"", raw.strip(), flags=re.M).strip()match= re.search(r"\{.*\}", text, re.S)ifnotmatch:return{"error":"model returned no JSON","raw": raw[:2000]}try: ledger = json.loads(match.group(0))except json.JSONDecodeError as exc:return{"error":f"invalid JSON from model: {exc}","raw": raw[:2000]} flags =[]try: line_sum =round(sum(float(li["amount"])for li in ledger.get("line_items",[])),2)ifabs(line_sum -float(ledger["subtotal"]))>0.01: flags.append(f"line items sum to {line_sum}, subtotal reads {ledger['subtotal']}") expected_total =round(float(ledger["subtotal"])+float(ledger["tax"]),2)ifabs(expected_total -float(ledger["total"]))>0.01: flags.append(f"subtotal+tax = {expected_total}, total reads {ledger['total']}")except(KeyError, TypeError, ValueError)as exc: flags.append(f"could not verify arithmetic: {exc}")return{"ledger": ledger,"verified":not flags,"flags": flags,"model": model}
The model input is the whole comparison harness: pass qwen3.8-27b for the launch model, or the imported anthropic/claude-opus-4.6 endpoint for the closed comparison. Registering an external OpenAI-compatible endpoint takes a minute in the import wizard; weight-based imports from Hugging Face or S3 have their own walkthrough. Both run behind GraphN's one OpenAI-compatible API, which is what made this test a short function instead of a project.
The protocol
Every measurement below is a workflow execution: 200 runs of Document to Ledger per condition, visible in the workspace's execution history, each traceable to an execution id.
Documents: 100 synthetic invoices and receipts (112 pages) generated from a seeded random process across 8 template families: clean invoices, dense 9-12 row parts tables with discounts, German VAT invoices with comma decimals, thermal receipts, two-page invoices, handwritten annotations, stamps across the totals, and credit memos with negative amounts. Synthetic on purpose: we own the ground truth and there are no privacy questions.
Two conditions: clean renders (1,400px JPEG at quality 92), and a degraded scan pass (1,000px, randomized skew, blur, and heavy JPEG compression) to approximate real scanner output.
Scoring: up to 11 units per document: vendor, invoice number, both dates, currency, PO number, line-item count, the full set of line amounts, subtotal, tax, total. 1,087 scorable units per model per condition.
Settings: temperature 0, one pass, identical prompt, 8,000-token completion budget, and both models at their serving defaults through the production path (Qwen's thinking on by default; Opus at its default reasoning effort).
Scoring rules we committed to before reading results by model: where a discount row makes "subtotal" legitimately two values, either reading counts; credit-memo totals compare by magnitude because the template renders parentheses either way; two documents whose pages carry no currency marker under a European vendor name are excluded from currency scoring as undecidable for any reader.
Two harness lessons worth passing on. First, thinking models need output headroom: at a 4,000-token budget, Qwen occasionally spent nearly all of it reasoning before emitting JSON. Second, serving stacks have opinions: the content filter on the Qwen route refused 2 of the 100 degraded scans outright (a stamped invoice and a credit memo that both passed at clean quality). Those two documents are excluded from Qwen's scan scoring and counted in its favor nowhere; in production that failure mode needs a retry-or-route-around plan just like any other.
The hardest pages in the set, skewed and heavily compressed scans with stamps across the totals, are the condition the theatre's Run tab replays:
Degraded scan of an invoice with a PAID stamp overlapping the subtotal, tax, and total rows
The results
Scroll horizontally to compare
Qwen3.8-27B
Claude Opus 4.6
Clean renders, fields correct
1,075 / 1,087 (98.9%)
1,081 / 1,087 (99.4%)
Scan quality, fields correct
1,045 / 1,066 (98.0%)
1,076 / 1,087 (99.0%)
Average seconds per document, end to end (scan)
11.6
10.1
Cost per 1,000 scanned documents, list price
$3.73
8.98
Latency is the full workflow execution: fetch the pages, run the model, verify the arithmetic. Cost uses each model's public list price on OpenRouter as of August 15, 2026, applied to measured token consumption for identical requests, so both models are priced from the same public sheet. The Qwen side runs on GraphN managed inference; Opus ran as an imported OpenRouter endpoint, and its share of the entire experiment billed under eleven dollars.
The aggregate hides the interesting part: the two models have different failure personalities.
Scroll horizontally to compare
Family (scan condition)
Qwen3.8-27B
Claude Opus 4.6
Clean invoices
143 / 143
143 / 143
Dense tables with discounts
139 / 143
139 / 143
German VAT, comma decimals
132 / 132
132 / 132
Thermal receipts
133 / 143
143 / 143
Two-page invoices
132 / 132
132 / 132
Handwritten annotations
125 / 132
125 / 132
Stamped totals
131 / 131
142 / 142
Credit memos
110 / 110
120 / 120
Qwen's weakness is small monospace type: thermal-receipt check numbers, in both conditions. Opus's weaknesses run the other way: handwritten purchase-order numbers, even on clean renders, and line-item counts on dense scans. On the hardest category, cursive handwriting at scan quality, the two models are identical: 125 of 132 each, failing largely the same documents. German decimal commas, negative credit amounts, stamps over totals, and two-page stitching bothered neither model.
What we conclude, and what we do not
Two hundred workflow executions per condition is a serious field test, though still not a leaderboard. What it supports: on the most commercially common vision workload, run through a real production path, the open 27B model and the closed frontier model its launch table names finish within half a point on clean documents and within a point on degraded scans, with complementary rather than nested failure modes. At list prices, the same job costs about a fifth as much on the open model. That is consistent with Qwen's "mixed results" framing on document benchmarks, now checked against production-shaped work with every data point traceable to an execution id.
If your documents are harder than these, run the same test yourself: the workflow above is the harness, and swapping model is the entire experiment.
Independent coverage noting the launch table's mixed document and vision results against Claude Opus 4.6 and the absence of third-party reproduction at publication.