AWS price feeds read September 9, 2026

AWS OCR and Intelligent Document Processing: Textract, Bedrock Data Automation and Rekognition Pricing Compared

The same page costs $1.50 or $80.00 per 1,000 on Amazon Textract depending on which feature flags your code sends. Here is the whole matrix, read straight from Amazon's price list feed, plus where Bedrock Data Automation and Rekognition genuinely fit.

Written for US engineering and finance teams costing a document pipeline on AWS. Every rate is US East (N. Virginia) on-demand, taken from Amazon's machine-readable price list, not from a blog. Last updated September 2026.

  • All 16 Textract feature rates
  • Where BDA is cheaper than Textract
  • The 100-word Rekognition trap
  • A 50% saving in one code change
Fields out, no Block parsing

PDF, JPG, PNG, BMP, HEIC, TIFF

Upload a document to extract

Send a real document through before you model a per-page rate on a spreadsheet.

Encrypted in transit
Source files auto-purged
US data handling
Seconds per document
53x
is the spread between the cheapest and dearest Textract rate for the same page
33x
is what adding the FORMS flag does to your per-page cost, $1.50 to $50.00
0
times the word confidence appears in Amazon's Bedrock Data Automation output docs
100
words is all Rekognition DetectText will read from an image, per AWS's own quota page
// The short answer

What AWS OCR costs, in one paragraph

AWS has four services that touch text and only one of them reads documents. Amazon Textract is it, and its price depends entirely on which feature flags your code sends: $1.50 per 1,000 pages for plain text, $15.00 with TABLES, $50.00 with FORMS, and $80.00 for FORMS plus Custom Queries plus TABLES in a single call. Bedrock Data Automation charges a flat $10.00 per 1,000 pages for standard document output. Rekognition DetectText is the cheapest thing on the list at $1.00 per 1,000 images and it will not read a page. Comprehend does not do OCR at all; it analyzes text something else already extracted.

The comparison that gets repeated everywhere, "Textract is $1.50 and Bedrock Data Automation is $10.00", is misleading, and this is the part worth taking away. It puts BDA's full document output next to Textract's plainest operation. Priced against the Textract calls that return the same thing, BDA Standard at $10.00 is cheaper than Textract TABLES at $15.00 and much cheaper than Textract FORMS at $50.00 for the first million pages. What BDA actually lacks is a volume tier, which it has none of at any volume, and a confidence score, which it does not return.

Where this page is not neutral, said up front. DocuOCR sells a document extraction API, so we compete with Textract. We have still written the cases where staying on AWS is the right call, and there are three of them below, including the one where Textract is simply the cheapest option on earth. Every figure here comes from Amazon's own machine-readable price list feeds and its own documentation, so you can check all of it. Our own rates sit on the pricing page and the cross-vendor view is on OCR pricing per 1,000 pages.

// Pick the right box first

The four AWS services people mean by "AWS OCR"

Most of the wasted engineering time in this category happens before any code is written, when a team picks the wrong service because two of them sound like OCR and one of them is much cheaper. The table below is the whole decision. Note that Rekognition is priced per image while Textract and BDA are priced per page, so those numbers are not directly interchangeable even though they look it.

Service Accepts US East rate Confidence score? Volume discount? What it is actually for
Amazon Textract Documents: PDF, TIFF, PNG, JPEG $1.50 to $80.00 per 1,000 pages Yes, on every Block, 0 to 100 Yes, above 1M pages The default choice for pages
Bedrock Data Automation Documents, images, audio, video $10.00 per 1,000 pages standard, $40.00 custom No, the word does not appear in the output docs No tier at any volume Structured output without writing the parser
Amazon Rekognition DetectText PNG and JPEG only, no PDF $1.00 per 1,000 images, $0.40 above 35M Yes, per detection Yes, four tiers Text in photographs, capped at 100 words
Amazon Comprehend Text you already extracted Priced per unit of text, not per page Yes, per entity Not applicable Entities, classification and PII after OCR

The Rekognition trap, quoted from AWS

Rekognition DetectText is the cheapest text extraction on AWS and teams find it first because of the price. Then they discover what AWS's own quota page says: "DetectText can detect up to 100 words in an image", and "Amazon Rekognition supports the PNG and JPEG image formats". A one-page invoice is well over 100 words and usually arrives as a PDF. Rekognition reads text in photographs, which is a different job from reading a page, and no amount of tuning changes that cap.

Comprehend is a second stage, not an alternative

Amazon Comprehend never sees your scan. It takes text that already exists and returns entities, sentiment, language and PII, which is why AWS publishes a pattern for piping Textract output into it rather than choosing between them. If your source is a scanned PDF, Comprehend cannot start until Textract has finished, and you pay for both stages. Anyone comparing "Comprehend against Textract" is comparing consecutive steps.

// Read from the price feed, not the marketing page

Every Amazon Textract rate, from $1.50 to $80.00 per 1,000 pages

These are the synchronous US East (N. Virginia) on-demand rates from Amazon's price list feed for Textract, offer file 20260831092230, published August 31, 2026 and read on September 9, 2026. Asynchronous rates match. The last column is the multiplier against plain text detection, which is the number nobody puts in front of an engineer before they add a feature flag.

Operation or feature set What you get back First 1M pages / 1,000 Above 1M / 1,000 vs plain text
DetectDocumentText Plain text and lines $1.50 $0.60 1x
AnalyzeDocument SIGNATURES Signature locations $3.50 $1.40 2.3x
AnalyzeDocument LAYOUT Titles, headers, paragraphs, lists $4.00 $3.00 2.7x
AnalyzeExpense Receipt and invoice fields $10.00 $8.00 6.7x
AnalyzeDocument TABLES Table structure as cells $15.00 $10.00 10x
AnalyzeDocument QUERIES Answers to questions you ask $15.00 $10.00 10x
QUERIES + TABLES Both, in one call $20.00 $15.00 13.3x
AnalyzeID Driver licenses and passports $25.00 $10.00 above 100k 16.7x
Custom Queries Queries against your adapter $25.00 $15.00 16.7x
Custom Queries + TABLES Both, in one call $30.00 $20.00 20x
AnalyzeDocument FORMS Key-value pairs $50.00 $40.00 33.3x
FORMS + QUERIES Both, in one call $55.00 $45.00 36.7x
FORMS + Custom Queries Both, in one call $65.00 $50.00 43.3x
Analyze Lending Mortgage document package $70.00 $55.00 46.7x
FORMS + QUERIES + TABLES All three, in one call $70.00 $55.00 46.7x
FORMS + Custom Queries + TABLES All three, in one call $80.00 $60.00 53.3x

Two details in that table are worth pulling out. SIGNATURES has the steepest volume discount in the whole service, falling 60% from $3.50 to $1.40. And AnalyzeID is the only Textract operation whose discount does not wait for a million pages: its tier boundary sits at 100,000, where the rate drops from $25.00 to $10.00 per 1,000. If you process identity documents, that is a genuinely favourable tier and it is not advertised anywhere. The wider page-limit and file-size rules that go with these operations are on the Textract limits page, and the operation-by-operation breakdown sits on OCR API operations.

// A one-line code change

Splitting one AnalyzeDocument call into two costs you 50% more

Textract prices feature combinations below the sum of their parts, and the AWS pricing page does not spell this out, so plenty of production code asks for tables in one call and queries in another. Same document, same result, more money. Here is what the price feed says the difference is worth.

Features you need As separate calls, per 1,000 pages In one call You save
QUERIES and TABLES $15.00 + $15.00 = $30.00 $20.00 $10.00 per 1,000 pages, 33% off
FORMS and QUERIES $50.00 + $15.00 = $65.00 $55.00 $10.00 per 1,000 pages, 15% off
Custom Queries and TABLES $25.00 + $15.00 = $40.00 $30.00 $10.00 per 1,000 pages, 25% off
FORMS, QUERIES and TABLES $50.00 + $15.00 + $15.00 = $80.00 $70.00 $10.00 per 1,000 pages, 12.5% off

The saving is a flat $10.00 per 1,000 pages every time, which at two million pages a year is $20,000 for passing an array with two entries instead of calling twice. Worth ten minutes of somebody's afternoon. The larger version of the same problem is routing: if one document type in your pipeline needs FORMS and the rest do not, sending everything through FORMS costs $50.00 per 1,000 on all of it. Splitting the route by document type is usually the biggest single saving available on an AWS document pipeline, and we work through the other levers in reducing AWS document extraction costs.

// Correcting the comparison everyone repeats

Bedrock Data Automation is not the expensive one, it is the untiered one

Bedrock Data Automation charges $0.0100 per page for standard document output and $0.0400 per page for custom output driven by a blueprint, plus $0.0005 per additional field per page. That is $10.00 and $40.00 per 1,000 pages. Read against Textract's plain text rate it looks like a 6.7x penalty, which is how every article frames it.

Read against what it replaces, the picture inverts. BDA Standard returns text, markdown, tables and CSV in one call for $10.00. The nearest Textract equivalent, AnalyzeDocument with TABLES, is $15.00. BDA Custom returns your named fields for $40.00; Textract FORMS returns key-value pairs you still have to map for $50.00. Below a million pages, BDA is the cheaper of the two on structured output, and almost nobody says so.

The real problem with BDA is the shape of its price, not the level. In Amazon's Bedrock price feed, offer file 20260901205051 published September 1, 2026, every BDA document rate is a single price dimension running from zero to infinity. There is no tier. Textract tiers everything at a million pages. So the gap against plain text detection widens from 6.7x to 16.7x exactly at the volume where a buyer most needs it to narrow.

Annual bill at four rates, US East on-demand

Pages Textract text BDA standard Textract tables Textract forms
100,000 $150 $1,000 $1,500 $5,000
500,000 $750 $5,000 $7,500 $25,000
1,000,000 $1,500 $10,000 $15,000 $50,000
5,000,000 $3,000 $50,000 $50,000 $200,000
10,000,000 $6,000 $100,000 $100,000 $400,000

Tiered rates applied where they exist: Textract text drops to $0.60 per 1,000 above a million pages, tables to $10.00, forms to $40.00. BDA stays at $10.00 throughout. At five million pages BDA and Textract TABLES cost the same $50,000, and above that Textract keeps falling while BDA does not. Full BDA rate card on the Bedrock Data Automation pricing page.

The migration detail that breaks pipelines: BDA does not return a confidence score

Textract puts a Confidence value on every Block it returns. Amazon's API reference defines it as "The confidence score that Amazon Textract has in the accuracy of the recognized text and the accuracy of the geometry points around the recognized text", a float with a valid range of 0 to 100. Teams build their entire human-review queue on that number: above the threshold auto-approve, below it route to a person. We counted the word "confidence" in Amazon's Bedrock Data Automation document-output documentation again on September 9, 2026 and it appears zero times across 23,637 characters of rendered text. A team that migrates from Textract to BDA for the structured output silently loses the input its approval logic runs on, and rebuilding it means scoring one model's output with a second model. The thresholds themselves, and why a 98% per-field pass rate still sends a third of your invoices to a human, are worked through on the OCR confidence score page.

// There is no SKU called this

AWS intelligent document processing is a pipeline you build, not a product you buy

Searching for AWS intelligent document processing turns up solution guides, reference architectures and a GitHub accelerator, which is the honest answer: AWS sells the parts and you assemble them. That matters for costing, because the per-page rate on this page is only one line of the bill. The rest looks like this.

1

Ingest and storage

S3 for the documents, plus whatever puts them there: an SFTP transfer family endpoint, an email intake, an upload form. Storage is cheap; the request charges on a high-volume pipeline are not nothing.

2

Extraction

Textract or Bedrock Data Automation, at the per-page rates above. This is the line everybody budgets and it is frequently not the largest one.

3

Orchestration and reconstruction

Lambda or Step Functions to drive the calls, retries and pagination, plus the code that turns Block objects into fields your database understands. This is engineering time, and it recurs every time a vendor changes a layout.

4

Human review

A queue, a UI and the people staffing it. On AWS this is Augmented AI or something you build. Its size is set by your confidence threshold, which is why losing the confidence score changes the cost of the whole pipeline, not just one call.

Step three is the one that surprises people. Textract returns an array of Block objects with types like WORD, LINE, CELL, MERGED_CELL, KEY_VALUE_SET and QUERY_RESULT, connected by a relationships graph. Getting from that to a row in your ledger is a parser, and it has a well-known sharp edge: Amazon's documentation states that a CELL block always reports a row span and column span of one, with real spans living on separate MERGED_CELL blocks, so the obvious loop over cells silently flattens every merged header and shifts the columns after it. We wrote that up in detail on whether Textract outputs CSV, and the broader build-or-buy arithmetic sits on build or buy for document extraction.

// Including where we lose

When to stay on AWS, and when a document API is the cheaper answer

Plain text at very high volume

Above a million pages a month Textract DetectDocumentText falls to $0.60 per 1,000 pages, which is the cheapest published document OCR rate anywhere. If all you need is characters off a page and the documents already sit in S3, nothing on this page beats it and we will not pretend otherwise.

Stay on Textract

Identity documents in the hundreds of thousands

AnalyzeID is the only Textract operation whose volume discount starts at 100,000 pages instead of a million, dropping from $25.00 to $10.00 per 1,000. A team processing driver licenses at that scale gets a tier nobody else on AWS gets.

Stay on Textract

A pipeline your team already operates

If Step Functions, Lambda, SQS and a human review queue are already built and running, the migration cost of moving extraction elsewhere is real and the per-page saving may not cover it. Price the switch honestly, not just the rates.

Stay on AWS

Typed fields rather than Block objects

Textract returns geometry and relationships, not values. Getting from a KEY_VALUE_SET array to invoice_total is a parser you write, test against every vendor layout you receive, and maintain forever. That engineering cost is invisible in a per-page comparison and it is usually the largest line in the real budget.

Compare an extraction API

Confidence-routed human review on BDA

Bedrock Data Automation returns structured output and no confidence score. If your approve-or-review decision runs on a per-field threshold, BDA removes the input that decision needs, and rebuilding it means a second model scoring the first one.

Compare an extraction API

Mixed document types under one FORMS call

A pipeline that routes every page through FORMS because one document type needs key-value pairs pays $50.00 per 1,000 on all of them. If splitting the routing is more work than moving, the arithmetic favours a service that returns fields at one rate.

Compare an extraction API

If you land on the right-hand column, the two comparisons worth reading next are Amazon Textract alternatives for the vendor view and high-volume OCR pricing for the arithmetic at scale. If your constraint is that documents cannot leave your own network, no cloud service on this page solves it, ours included, and the OCR SDK licensing comparison covers what that actually costs.

// Questions people actually ask

AWS OCR pricing questions, answered

Which AWS service should I use for OCR?

Amazon Textract for documents, and nothing else. Rekognition reads text in photographs and caps at 100 words per image with no PDF input. Comprehend does not read images at all, it analyzes text you already extracted. Bedrock Data Automation is the newer document service and is worth pricing against Textract, but for plain page-to-text on AWS, Textract DetectDocumentText at $1.50 per 1,000 pages is the answer.

How much does AWS OCR cost per 1,000 pages?

Between $1.50 and $80.00, on the same service, for the same page. Amazon's price list feed puts Textract DetectDocumentText at $1.50 per 1,000 pages in US East. Add the FORMS feature and it becomes $50.00. Ask for FORMS, Custom Queries and TABLES in one call and it is $80.00. Bedrock Data Automation charges a flat $10.00 per 1,000 pages for standard document output.

Is Bedrock Data Automation more expensive than Textract?

Only against Textract's cheapest operation. The comparison everybody repeats, $1.50 against $10.00, puts BDA's full document output next to plain text detection. Against the Textract operations that return the same thing, BDA Standard at $10.00 per 1,000 pages is cheaper than Textract TABLES at $15.00 and far cheaper than Textract FORMS at $50.00, for the first million pages.

Does Bedrock Data Automation return confidence scores?

No. We counted again on September 9, 2026: the word "confidence" appears zero times in Amazon's BDA document-output documentation, across 23,637 characters of text. Textract puts a Confidence float on every single Block, valid range 0 to 100. If your pipeline routes documents to a human on a confidence threshold, that threshold has nothing to read after you migrate.

Can Amazon Rekognition read documents?

No, and AWS says so in its own quota page. "DetectText can detect up to 100 words in an image", and "Amazon Rekognition supports the PNG and JPEG image formats". A single-page invoice runs well past 100 words and usually arrives as a PDF, which Rekognition will not accept. Rekognition is built for text in photographs: signs, labels, packaging, license plates.

Why is my AWS Textract bill higher than I expected?

Almost always because a feature flag was added to the call. Going from DetectDocumentText to AnalyzeDocument with FORMS multiplies the page rate by 33, from $1.50 to $50.00 per 1,000 pages, and nothing in the SDK warns you. TABLES and QUERIES are each a 10x step. The bill line item is per page, so the multiplier applies to every page you have ever sent.

Does Textract charge less at high volume?

Yes, but only past one million pages a month, and only on Textract. Text drops from $1.50 to $0.60 per 1,000, TABLES from $15.00 to $10.00, FORMS from $50.00 to $40.00. AnalyzeID is the exception and the only good one: its discount starts at 100,000 pages, dropping from $25.00 to $10.00 per 1,000. Bedrock Data Automation has no volume tier at all.

What is AWS intelligent document processing?

It is AWS's umbrella name for a pipeline rather than a product you buy. In practice it means Textract or Bedrock Data Automation for extraction, optionally Comprehend for classification and entity detection, S3 for storage, Lambda or Step Functions for orchestration, and a human review layer you build yourself. There is no single SKU called AWS IDP, which is why costing one is difficult.

Can I cut my Textract bill without losing any output?

Yes, and it is usually a one-line change. Textract prices feature combinations below the sum of the parts, so asking for QUERIES and TABLES in one AnalyzeDocument call costs $20.00 per 1,000 pages while asking in two separate calls costs $30.00. Same result, 50% more money. The same holds for FORMS with QUERIES, $55.00 together against $65.00 split.

Does AWS Textract read handwriting?

Yes, and it labels it. Every Block carries a TextType field whose valid values are HANDWRITING and PRINTED, so you can tell which parts of a page were written by hand and treat them differently. It costs nothing extra: handwriting is covered by the same DetectDocumentText and AnalyzeDocument rates as printed text.

What is the difference between Textract and Comprehend?

Textract turns pixels into text and structure. Comprehend takes text that already exists and finds entities, sentiment, language and PII in it. They are stages, not alternatives, which is why AWS documents a pattern for sending Textract output into Comprehend. If your document is a scan, Comprehend cannot start until Textract has finished.

Is there an AWS OCR free tier worth planning around?

Not for a production workload. AWS free tiers on these services are small and time-limited, and any real document pipeline exhausts them in the first week. Budget from the per-page rates instead. The number that matters is your monthly page count multiplied by the rate for the specific feature combination your code sends, not the trial allowance.

How do I estimate an AWS document pipeline before building it?

Take your monthly pages, decide the exact feature combination each document type needs, and price each type separately. A pipeline that sends every page through FORMS because one document type needs key-value pairs pays $50.00 per 1,000 on all of them. Splitting the routing so only the forms go through FORMS is usually the single largest saving available.

Does Textract return JSON I can use directly?

It returns JSON, but not fields. Textract gives you an array of Block objects with types like WORD, LINE, CELL, MERGED_CELL, KEY_VALUE_SET and QUERY_RESULT, plus relationships between them. Turning that into "invoice_total: 1420.00" is code you write and maintain. That reconstruction layer is a real cost that never appears in a per-page rate comparison.

Should I stay on AWS for document extraction?

Often yes, and we will say so. If your documents already live in S3, your team knows IAM, and you need plain text at very high volume, Textract at $0.60 per 1,000 pages above a million is hard to beat on price. The case for leaving is when you need typed fields with confidence per field, and you would rather not build and own the reconstruction layer.

One rate, typed fields, confidence included

DocuOCR has no feature flags to multiply your bill, no Block graph to reconstruct and no missing confidence score. You send a document and get back named fields with a score on each one, at a published per-page rate. Upload something real above and compare the output against what a Textract AnalyzeDocument response would need before it becomes a database row. If plain text at eight figures of volume is genuinely all you need, stay on Textract and spend the afternoon elsewhere.