At What Volume Does OCR Pricing Actually Get Cheaper?

Jul 23, 2026 7 min read

AWS and Azure halve their basic OCR rate above one million pages a month. Google does not until five million. All three list at the same $1.50, which makes a list price comparison the wrong tool at volume.

// Try it now

PDF, JPG, PNG, BMP, HEIC, TIFF

Upload a document to extract

Above one million pages a month on AWS Textract and Azure Document Intelligence, where basic OCR drops from $1.50 to $0.60 per 1,000 pages. On Google Document AI the same discount does not arrive until above five million pages a month. All three list at exactly $1.50, so the vendor that looks identical on a pricing page can cost two and a half times as much once you are actually running volume.

That five times difference in where the boundary sits is the single most consequential number in high volume document processing, and it is invisible in every list price comparison including the ones we publish. Here is the full picture, verified from the vendors' own sources in July 2026.

Where each tier boundary sits

Google's pricing page sets Enterprise Document OCR tier one at 1 to 5,000,000 pages a month at $1.50 per 1,000, and tier two at 5,000,001 and above at $0.60. AWS Textract Detect Document Text and Azure Read both put that same $0.60 rate above one million pages. Every one of those numbers is public and none is surprising on its own. Together they mean the band between one and five million pages a month, which is where a great many real pipelines sit, is exactly where the three vendors are furthest apart despite listing identically.

Worth noting: Google's structured meters do not behave this way. Form Parser and Custom Extractor both tier at 1,000,001 pages, the same boundary the other two use. It is specifically the basic OCR meter where the boundary sits five times further out.

What that costs in practice

Monthly cost for text extraction only, on the meter each vendor positions for it. The Azure commitment column is our arithmetic over Microsoft's published tier and overage rates, picking the cheapest combination at each volume.

Pages per monthAWS TextractAzure, pay as you goAzure, best commitmentGoogle
100,000$150$150$150$150
500,000$750$750$375$750
1,000,000$1,500$1,500$750$1,500
2,000,000$2,100$2,100$1,200$3,000
5,000,000$3,900$3,900$3,000$7,500
10,000,000$6,900$6,900$5,260$10,500

The three are level until a million pages and then they separate. At two million the gap between Google and AWS is $900 a month. At ten million it is $3,600 a month, or over $43,000 a year, for output that is broadly comparable on clean documents. None of that is hidden pricing; it is four published numbers that nobody puts on the same page.

The Azure commitment tiers most buyers never see

Azure publishes commitment tiers through its Retail Prices API and its enterprise documentation, but not on the marketing pricing page most people read. For Read, the published tiers are $375 a month for 500,000 pages, $1,200 for 2 million, $4,200 for 8 million and $7,200 for 16 million, with overage rates of $0.75, $0.60, $0.53 and $0.45 per 1,000 respectively.

Two things stand out. The 500,000 page commitment reaches $0.75 per 1,000 at half the volume where the pay-as-you-go discount even begins, so a company running half a million pages a month is paying twice what it needs to if it never asked. And the 16 million tier at $0.45 is 25 percent below the best pay-as-you-go rate reachable on any of the three clouds. If you take one action from this article, it is to check whether your Azure spend already qualifies for a tier nobody offered you.

The meters that never get cheaper

This is where a scaling budget usually breaks, because the assumption that everything discounts eventually is wrong.

MeterVendorList rateAt volume
QueriesAWS Textract$15 / 1kNo discount at any volume
Layout ParserGoogle Document AI$10 / 1kNo discount at any volume
Documents, standard outputBedrock Data Automation$10 / 1kNo discount at any volume
Documents, custom outputBedrock Data Automation$40 / 1kNo discount at any volume

Amazon Bedrock Data Automation is the outlier worth naming. It has no volume discount on any meter at any scale, and its $40 per 1,000 pages custom output rate is the most expensive flat custom rate among the big three. It also charges an extra $0.0005 per field per page above 30 fields, so a wide schema raises the rate rather than leaving it flat. A team that models its year on the basic OCR tier curve and then ships a pipeline built on Queries, Layout Parser or Bedrock gets a straight line no amount of growth bends.

Three things that quietly cost you the discount

Split billing splits your volume. Tiers count per billing account and usually per region. Two AWS accounts each running 600,000 pages a month both sit in the expensive band, while one account running 1.2 million would not. Consolidate billing before you model the curve, not after.

Batch does not discount everywhere. On the language model readers, batch is about half price. On the traditional OCR APIs it usually is not: Azure Document Intelligence has no batch discount at all, and asynchronous calls on AWS and Google bill the same per page as synchronous ones. And on Anthropic and Mistral, batch is explicitly excluded from zero data retention arrangements, so if your contract depends on one, the discounted path may not be open to you.

Throughput ceilings arrive before the price does. At these volumes the binding constraint is often pages per minute rather than dollars per page. Google caps an online request at 15 pages while Azure accepts 2,000 in a single call, and provisioned capacity on Google costs $300 per extra page-per-minute per month on top of the per-page rate. That is a real line item that never appears in a per-1,000-pages comparison.

When does running your own model win?

Once volume is genuinely large, the alternative to a per-page meter is a GPU rented by the hour, and the arithmetic turns on utilization rather than on the model being free. A per-page API bills nothing when no documents arrive. A GPU bills all 730 hours in a month whether it is busy or idle, so the real rate is monthly instance cost divided by pages actually processed.

On the cheapest current AWS card suited to the work, break-even against the $1.50 cloud rate lands around 390,000 pages a month. Below that the GPU is dramatically worse, and at 10,000 pages a month it works out to roughly 39 times the cloud rate. Above it the GPU pulls ahead and keeps going. The catch is that break-even assumes sustained utilization and most real document volume is spiky, so the honest comparison uses your trough month rather than your peak. We worked the whole calculation through on self hosted OCR cost.

What to do with this

Model your actual monthly page count against each vendor's tier boundary rather than against its headline rate, and do it per meter, because a pipeline that runs basic OCR on everything and structured extraction on a subset has two different curves running at once. Check whether a commitment tier already applies to you. Then negotiate, treating the published curve as the floor of the conversation rather than the answer, because every vendor here does private pricing above a certain spend and none of them publishes the discount. Walking in with a modeled annual page count, a named alternative and the tier math is what makes that conversation short.

It is also worth putting document extraction spend somewhere it gets looked at monthly. This is a line item that grows with usage rather than with headcount, which is exactly the profile that goes unnoticed until it is large, and the same discipline that catches cloud and SaaS spend drifting upward catches an OCR meter quietly crossing a tier boundary. The full breakdown, including every meter and the sources for each figure, is on high volume OCR API pricing.

Google's tier boundaries were read off Google's own pricing page and the Azure commitment tiers off the Azure Retail Prices API, the feed the portal bills from, both on July 23, 2026. Rates change, so re-verify before you build a business case on them.

Extract your documents with DocuOCR

DocuOCR's AI OCR software turns any document into clean, structured data in seconds. No template setup required.

Start free

← Back to all articles

From the same family of tools