// Re-verified from vendor sources, July 23, 2026

High Volume OCR API Pricing: What Textract, Azure and Google Really Cost Above 1 Million Pages

All three hyperscalers list basic OCR at $1.50 per 1,000 pages. Their volume tiers start in different places, and one of them does not discount until five million pages a month. Here is the real curve from 100,000 to 10 million pages, with the commitment tiers that are not on the pricing pages.

  • The cost curve at six volumes
  • Where each tier boundary sits
  • Meters that never get cheaper
  • The commitment tiers nobody shows
Upload a document, no signup

PDF, JPG, PNG, BMP, HEIC, TIFF

Upload a document to extract

Run a real document from your pipeline and check the field accuracy before you model the volume.

SOC 2 Type II
Parallel batch processing
US data handling
Fields, not just text
$1.50
the identical rate all three list
5x
how much later Google's tier begins
$0.45
the lowest published rate, Azure 16M tier
4
meters with no discount at any volume
// The short answer

What a high volume OCR API actually costs

At high volume, basic OCR runs between $0.45 and $1.50 per 1,000 pages depending on vendor and contract, and structured field extraction runs between $18 and $50. The spread inside that range is decided almost entirely by where each vendor's volume tier begins, not by the list price, because AWS Textract Detect Document Text, Azure Read and Google Enterprise Document OCR all list at exactly the same $1.50 per 1,000 pages.

AWS and Azure begin their discounted tier above one million pages a month. Google does not begin its discounted tier until above five million. That single difference means a company running two million pages a month pays about $2,100 on AWS, about $1,200 on an Azure commitment tier, and about $3,000 on Google, for output that is broadly comparable. The list-price comparison that put them level was accurate and useless.

Last updated July 2026. Google's tier boundaries were read off Google's own pricing page and the Azure commitment tiers off the Azure Retail Prices API, which is the feed the portal bills from, both on July 23, 2026.

// Monthly cost, basic OCR

What basic OCR costs per month from 100,000 to 10 million pages

Text extraction only, on the meter each vendor positions for it: Textract Detect Document Text, Azure Read, and Google Enterprise Document OCR. The three are level until a million pages, then they separate.

Pages per month AWS Textract Azure Read, pay as you go Azure, best commitment tier Google Enterprise Document OCR
100,000 pages $150 $150 $150 $150
500,000 pages $750 $750 $375 $750
1,000,000 pages $1,500 $1,500 $750 $1,500
2,000,000 pages $2,100 $2,100 $1,200 $3,000
5,000,000 pages $3,900 $3,900 $3,000 $7,500
10,000,000 pages $6,900 $6,900 $5,260 $10,500

How to read this. The AWS and Azure pay-as-you-go columns apply the published graduated tiers, so a 2 million page month bills the first million at $1.50 and the second at $0.60. The Azure commitment column is our arithmetic over Microsoft's published commitment and overage rates, picking the cheapest combination at each volume, and it is our calculation rather than a figure Microsoft quotes as a package. The Google column stays at $1.50 all the way to five million pages because that is where its first tier ends. Negotiated enterprise pricing exists on every vendor here and none of them publishes it, so treat this curve as the floor of the conversation.

// The finding

Google's discount starts five times later than everyone else's

Google's pricing page sets Enterprise Document OCR tier one at 1 to 5,000,000 pages a month at $1.50 per 1,000, and tier two at 5,000,001 and above at $0.60. AWS Textract and Azure Document Intelligence both put the same $0.60 rate above one million pages. Every one of those four numbers is public and none of them is surprising in isolation. Put side by side, they mean that the band between one and five million pages a month, which is where a great many real document pipelines actually sit, is the band where the three vendors are furthest apart despite listing identically.

At two million pages a month the gap is $900 against AWS. At five million it is $3,600, and Google is still on its opening rate while the others have been discounted for four million pages. Only above five million does Google's rate finally match, and by then the cumulative difference for the year is substantial. Notably, Google's structured meters do not behave this way: Form Parser and Custom Extractor both tier at 1,000,001 pages, the same boundary the other two use. It is specifically the basic OCR meter where the boundary sits five times further out.

This is not an argument that Google is the wrong choice. Its per-request page ceiling, its parser lineup and the fact that it does not bill failed requests all matter, and we cover those on the OCR pricing per 1,000 pages reference. It is an argument that a list-price comparison is the wrong tool once you are past a million pages, and that the tier boundary belongs in your model before the vendor logo does.

// Not on the marketing page

The Azure commitment tiers most buyers never see

Read off the Azure Retail Prices API for East US on July 23, 2026. This is the feed the portal bills from, and it exposes tiers the marketing pricing page does not show.

Commitment tier Monthly fee Overage rate Effective rate at the allowance
Read 500K $375 / mo $0.75 / 1k $0.75 per 1,000
Read 2M $1,200 / mo $0.60 / 1k $0.60 per 1,000
Read 8M $4,200 / mo $0.53 / 1k $0.53 per 1,000
Read 16M $7,200 / mo $0.45 / 1k $0.45 per 1,000
Prebuilt 500K $4,000 / mo $8.00 / 1k $8.00 per 1,000
Prebuilt 1M $7,500 / mo $7.50 / 1k $7.50 per 1,000
Custom 500K $10,500 / mo $21.00 / 1k $21.00 per 1,000
Custom 1M $18,000 / mo $18.00 / 1k $18.00 per 1,000

Two things stand out. The 500,000 page Read commitment reaches $0.75 per 1,000 pages at half the volume where the pay-as-you-go discount even begins, so a company running half a million pages a month is paying twice what it needs to if it never asked. And the 16 million tier at $0.45 is 25 percent below the best pay-as-you-go rate anyone can reach on any of the three clouds. The connected container tiers run lower still, 15 to 20 percent below the equivalent Azure tier, though those bill through a deployment you host and operate yourself.

// Every meter, list rate and volume rate

Which OCR meters get cheaper at volume, and which never do

Four of these fourteen meters are flat at every volume. If your pipeline is built on one of them, growth does not bend your cost curve at all.

Meter Vendor List rate At volume
Detect Document Text AWS Textract $1.50 / 1k $0.60 above 1M pages/mo
Read Azure AI Document Intelligence $1.50 / 1k $0.60 above 1M, or $0.45 on the 16M commitment
Enterprise Document OCR Google Document AI $1.50 / 1k $0.60, but only above 5M pages/mo
Analyze Expense AWS Textract $10 / 1k $8 above 1M pages/mo
Tables AWS Textract $15 / 1k $10 above 1M pages/mo
Forms AWS Textract $50 / 1k $40 above 1M pages/mo
Analyze ID AWS Textract $25 / 1k $10 after the first 100K pages
Queries AWS Textract $15 / 1k No discount at any volume
Custom extraction Azure AI Document Intelligence $30 / 1k $20 above 1M, or $18 on the 1M commitment
Custom Extractor Google Document AI $30 / 1k $20 above 1M pages/mo
Form Parser Google Document AI $30 / 1k $20 above 1M pages/mo
Layout Parser Google Document AI $10 / 1k No discount at any volume
Documents, standard output Amazon Bedrock Data Automation $10 / 1k No discount at any volume
Documents, custom output Amazon Bedrock Data Automation $40 / 1k No discount at any volume

Amazon Bedrock Data Automation is the outlier worth naming. It has no volume discount on any meter at any scale, and its custom output rate of $40 per 1,000 pages is the most expensive flat custom rate among the big three. It also charges an extra $0.0005 per field per page above 30 fields, so a wide schema raises the rate rather than leaving it flat. We break that down on the Amazon Bedrock Data Automation pricing reference.

// What breaks a volume forecast

Six things that make a high volume OCR budget wrong

01

The tier boundary, not the list price, decides the winner

All three hyperscalers list basic OCR at $1.50 per 1,000 pages. Between one and five million pages a month, two of them have already halved and one has not. Comparing list prices at that volume gives you exactly the wrong answer.

02

Some meters never get cheaper

Google Layout Parser and AWS Textract Queries are single tier at every volume, and no Bedrock Data Automation meter discounts at any scale. If your pipeline is built on one of those, your cost curve is a straight line and no amount of growth bends it.

03

Split billing splits your discount

Volume tiers count per billing account and usually per region. Two accounts each running 600,000 pages a month both sit in the expensive band, while one account running 1.2 million would not. Consolidate before you model.

04

The commitment tier is not on the pricing page

Azure publishes commitment tiers through its Retail Prices API and its enterprise documentation, not on the marketing pricing page most people read. The 500,000 page commitment reaches $0.75 per 1,000 at half the volume where the pay-as-you-go discount even begins.

05

Throughput ceilings arrive before the price does

At these volumes the binding constraint is often pages per minute rather than dollars per page. Google caps an online request at 15 pages, Azure takes 2,000 in a single call, and provisioned capacity on Google costs $300 per extra page-per-minute per month on top of the per-page rate.

06

The cheapest run may cost you your retention terms

Batch is where the discount lives on the language model readers, and it is also named in the zero data retention exclusion lists. If your contract depends on that arrangement, the discounted path may not be available to you.

// The other end of the curve

At what point does running your own model get cheaper

Once volume is genuinely large, the alternative to a per-page meter is a GPU you rent by the hour. The arithmetic turns on utilization rather than on the model being free. A per-page API bills nothing when no documents arrive; a GPU instance bills all 730 hours in a month whether it is busy or idle, so the real rate is the monthly instance cost divided by the pages you actually processed.

On the cheapest current AWS card suited to the work, that break-even against the $1.50 per 1,000 cloud rate lands somewhere around 390,000 pages a month. Below it the GPU is dramatically more expensive, and at 10,000 pages a month it is roughly 39 times the cloud rate. Above it the GPU pulls ahead and keeps going. What makes the decision hard is that the break-even assumes sustained utilization, and most real document volume is spiky. We worked the whole calculation through, including the hidden costs, on self hosted OCR cost.

There is also a constraint that arrives before the money does. At these volumes the ceiling is often pages per minute rather than dollars per page, and the per-request limits differ enormously: Google caps an online request at 15 pages while Azure accepts 2,000 in a single call. Those ceilings, and what it costs to raise them, are on the OCR API limits comparison.

// Where we fit

How DocuOCR prices at volume, and when a raw API is the better buy

DocuOCR plans work out to roughly $14 to $20 per 1,000 pages across the published tiers, and we are not going to pretend that is cheaper than $0.60 for raw text extraction, because it is not and it is not the same product. What you get for the difference is classification of a mixed batch, named field extraction with per-field confidence, validation rules, a review queue for the values that fail them, and an export that lands in your system rather than a JSON blob you still have to reconcile.

Here is the honest test. If what you need is text off a page and you already have engineers who will build classification, field mapping, validation and human review around it, a hyperscaler meter at volume is the cheaper path and you should take it. If what you need is extracted, checked data and you would otherwise be building that layer yourself, compare our rate against the meter plus the engineering, because the meter is rarely the expensive part. For very large or steady volume, talk to us before you model off the published tiers, which is the same advice this page gives about every vendor on it.

Frequently asked questions

Do OCR APIs get cheaper at high volume?
Some meters do and several do not. Basic OCR drops from $1.50 to $0.60 per 1,000 pages on AWS Textract and Azure Read once you pass a million pages a month. But Google Layout Parser is a single tier at $10 per 1,000 with no discount at any volume, AWS Textract Queries is single tier at $15, and Amazon Bedrock Data Automation has no volume discount on any meter at any scale.
At what volume does OCR pricing get cheaper?
It depends on the vendor, and this is the detail that costs money. AWS Textract and Azure Document Intelligence both begin their discounted tier above 1,000,000 pages per month. Google Enterprise Document OCR does not begin its discounted tier until above 5,000,000 pages per month. All three list at the same $1.50, so the difference only appears once you are actually running volume.
Which OCR API is cheapest at 2 million pages a month?
Azure, then AWS, then Google, even though all three list at $1.50 per 1,000 pages. At 2 million pages a month AWS Textract Detect Document Text works out to about $2,100 and Azure Read pay-as-you-go matches it, while Azure's 2M commitment tier is $1,200. Google Enterprise Document OCR is still inside its first tier at that volume, so it bills about $3,000.
What is an Azure Document Intelligence commitment tier?
A fixed monthly fee that buys a page allowance at a lower effective rate, with a defined overage rate above it. For Read the published tiers are $375 a month for 500,000 pages, $1,200 for 2 million, $4,200 for 8 million and $7,200 for 16 million. They are not shown on the marketing pricing page, which is why most buyers compare pay-as-you-go rates and conclude the market is flat.
Is there a volume discount on AWS Textract?
On most meters, yes, above 1 million pages a month. Detect Document Text goes from $1.50 to $0.60 per 1,000 pages, Tables from $15 to $10, Forms from $50 to $40 and Analyze Expense from $10 to $8. Analyze ID drops from $25 to $10 after the first 100,000 pages. Queries is the exception: it stays at $15 per 1,000 pages at every volume.
How much does it cost to OCR 10 million pages a month?
For basic text extraction, roughly $5,300 to $10,500 a month depending on vendor and contract. Azure on its 8 million page commitment tier plus overage works out near $5,260, AWS Textract near $6,900 on published tiers, and Google Enterprise Document OCR near $10,500 because its first tier runs all the way to 5 million pages. Structured field extraction costs several times more on all three.
Should I negotiate OCR pricing at high volume?
Yes, and treat the published curve as the floor of the conversation rather than the answer. Every vendor in this market does private pricing above a certain spend and none of them publishes the discount. Going in with a modeled annual page count, a named alternative and the published tier math is what makes that conversation short.
Is a per-page API still the right choice at very high volume?
Usually yes, until utilization gets high and stays there. A per-page API bills nothing when idle, while a GPU bills every hour of the month whether or not documents arrive. On current AWS GPU rates a self-hosted open model breaks even against the $1.50 cloud rate somewhere around 390,000 pages a month on the cheapest card, and only wins clearly when volume is both large and steady.
Do OCR volume discounts apply per account or per region?
Per billing account and usually per region, which quietly costs money at scale. Splitting a workload across two AWS accounts or two Azure subscriptions splits the volume that counts toward the tier, so each half can sit in the expensive band while the combined total would have qualified. Consolidate billing before you model the curve.
Does batch processing reduce the cost per page?
On the language model readers, yes, by about half. Mistral and Anthropic both discount batch by 50 percent and OpenAI's batch tier is roughly half price. The traditional OCR APIs are different: Azure Document Intelligence has no batch discount at all, and asynchronous calls on AWS and Google bill the same per page as synchronous ones.

Model the volume before you sign the tier

Run a real document from your pipeline through DocuOCR, check the fields and the confidence scores, and see what the extraction layer is worth before you price the meter underneath it.

From the same family of tools