The same page costs $1.50 or $80.00 per 1,000 on Amazon Textract depending on which feature flags your code sends. Here is the whole matrix, read straight from Amazon's price list feed, plus where Bedrock Data Automation and Rekognition genuinely fit.
Written for US engineering and finance teams costing a document pipeline on AWS. Every rate is US East (N. Virginia) on-demand, taken from Amazon's machine-readable price list, not from a blog. Last updated September 2026.
Upload a document to extract
Drop files here or click to upload
Up to 50 files
Free plan extracts the first 5 files, the rest unlock when you upgrade
Uploading...
Send a real document through before you model a per-page rate on a spreadsheet.
AWS has four services that touch text and only one of them reads documents. Amazon Textract is it, and its price depends entirely on which feature flags your code sends: $1.50 per 1,000 pages for plain text, $15.00 with TABLES, $50.00 with FORMS, and $80.00 for FORMS plus Custom Queries plus TABLES in a single call. Bedrock Data Automation charges a flat $10.00 per 1,000 pages for standard document output. Rekognition DetectText is the cheapest thing on the list at $1.00 per 1,000 images and it will not read a page. Comprehend does not do OCR at all; it analyzes text something else already extracted.
The comparison that gets repeated everywhere, "Textract is $1.50 and Bedrock Data Automation is $10.00", is misleading, and this is the part worth taking away. It puts BDA's full document output next to Textract's plainest operation. Priced against the Textract calls that return the same thing, BDA Standard at $10.00 is cheaper than Textract TABLES at $15.00 and much cheaper than Textract FORMS at $50.00 for the first million pages. What BDA actually lacks is a volume tier, which it has none of at any volume, and a confidence score, which it does not return.
Where this page is not neutral, said up front. DocuOCR sells a document extraction API, so we compete with Textract. We have still written the cases where staying on AWS is the right call, and there are three of them below, including the one where Textract is simply the cheapest option on earth. Every figure here comes from Amazon's own machine-readable price list feeds and its own documentation, so you can check all of it. Our own rates sit on the pricing page and the cross-vendor view is on OCR pricing per 1,000 pages.
Most of the wasted engineering time in this category happens before any code is written, when a team picks the wrong service because two of them sound like OCR and one of them is much cheaper. The table below is the whole decision. Note that Rekognition is priced per image while Textract and BDA are priced per page, so those numbers are not directly interchangeable even though they look it.
| Service | Accepts | US East rate | Confidence score? | Volume discount? | What it is actually for |
|---|---|---|---|---|---|
| Amazon Textract | Documents: PDF, TIFF, PNG, JPEG | $1.50 to $80.00 per 1,000 pages | Yes, on every Block, 0 to 100 | Yes, above 1M pages | The default choice for pages |
| Bedrock Data Automation | Documents, images, audio, video | $10.00 per 1,000 pages standard, $40.00 custom | No, the word does not appear in the output docs | No tier at any volume | Structured output without writing the parser |
| Amazon Rekognition DetectText | PNG and JPEG only, no PDF | $1.00 per 1,000 images, $0.40 above 35M | Yes, per detection | Yes, four tiers | Text in photographs, capped at 100 words |
| Amazon Comprehend | Text you already extracted | Priced per unit of text, not per page | Yes, per entity | Not applicable | Entities, classification and PII after OCR |
Rekognition DetectText is the cheapest text extraction on AWS and teams find it first because of the price. Then they discover what AWS's own quota page says: "DetectText can detect up to 100 words in an image", and "Amazon Rekognition supports the PNG and JPEG image formats". A one-page invoice is well over 100 words and usually arrives as a PDF. Rekognition reads text in photographs, which is a different job from reading a page, and no amount of tuning changes that cap.
Amazon Comprehend never sees your scan. It takes text that already exists and returns entities, sentiment, language and PII, which is why AWS publishes a pattern for piping Textract output into it rather than choosing between them. If your source is a scanned PDF, Comprehend cannot start until Textract has finished, and you pay for both stages. Anyone comparing "Comprehend against Textract" is comparing consecutive steps.
These are the synchronous US East (N. Virginia) on-demand rates from Amazon's price list feed for Textract, offer file 20260831092230, published August 31, 2026 and read on September 9, 2026. Asynchronous rates match. The last column is the multiplier against plain text detection, which is the number nobody puts in front of an engineer before they add a feature flag.
| Operation or feature set | What you get back | First 1M pages / 1,000 | Above 1M / 1,000 | vs plain text |
|---|---|---|---|---|
| DetectDocumentText | Plain text and lines | $1.50 | $0.60 | 1x |
| AnalyzeDocument SIGNATURES | Signature locations | $3.50 | $1.40 | 2.3x |
| AnalyzeDocument LAYOUT | Titles, headers, paragraphs, lists | $4.00 | $3.00 | 2.7x |
| AnalyzeExpense | Receipt and invoice fields | $10.00 | $8.00 | 6.7x |
| AnalyzeDocument TABLES | Table structure as cells | $15.00 | $10.00 | 10x |
| AnalyzeDocument QUERIES | Answers to questions you ask | $15.00 | $10.00 | 10x |
| QUERIES + TABLES | Both, in one call | $20.00 | $15.00 | 13.3x |
| AnalyzeID | Driver licenses and passports | $25.00 | $10.00 above 100k | 16.7x |
| Custom Queries | Queries against your adapter | $25.00 | $15.00 | 16.7x |
| Custom Queries + TABLES | Both, in one call | $30.00 | $20.00 | 20x |
| AnalyzeDocument FORMS | Key-value pairs | $50.00 | $40.00 | 33.3x |
| FORMS + QUERIES | Both, in one call | $55.00 | $45.00 | 36.7x |
| FORMS + Custom Queries | Both, in one call | $65.00 | $50.00 | 43.3x |
| Analyze Lending | Mortgage document package | $70.00 | $55.00 | 46.7x |
| FORMS + QUERIES + TABLES | All three, in one call | $70.00 | $55.00 | 46.7x |
| FORMS + Custom Queries + TABLES | All three, in one call | $80.00 | $60.00 | 53.3x |
Two details in that table are worth pulling out. SIGNATURES has the steepest volume discount in the whole service, falling 60% from $3.50 to $1.40. And AnalyzeID is the only Textract operation whose discount does not wait for a million pages: its tier boundary sits at 100,000, where the rate drops from $25.00 to $10.00 per 1,000. If you process identity documents, that is a genuinely favourable tier and it is not advertised anywhere. The wider page-limit and file-size rules that go with these operations are on the Textract limits page, and the operation-by-operation breakdown sits on OCR API operations.
Textract prices feature combinations below the sum of their parts, and the AWS pricing page does not spell this out, so plenty of production code asks for tables in one call and queries in another. Same document, same result, more money. Here is what the price feed says the difference is worth.
| Features you need | As separate calls, per 1,000 pages | In one call | You save |
|---|---|---|---|
| QUERIES and TABLES | $15.00 + $15.00 = $30.00 | $20.00 | $10.00 per 1,000 pages, 33% off |
| FORMS and QUERIES | $50.00 + $15.00 = $65.00 | $55.00 | $10.00 per 1,000 pages, 15% off |
| Custom Queries and TABLES | $25.00 + $15.00 = $40.00 | $30.00 | $10.00 per 1,000 pages, 25% off |
| FORMS, QUERIES and TABLES | $50.00 + $15.00 + $15.00 = $80.00 | $70.00 | $10.00 per 1,000 pages, 12.5% off |
The saving is a flat $10.00 per 1,000 pages every time, which at two million pages a year is $20,000 for passing an array with two entries instead of calling twice. Worth ten minutes of somebody's afternoon. The larger version of the same problem is routing: if one document type in your pipeline needs FORMS and the rest do not, sending everything through FORMS costs $50.00 per 1,000 on all of it. Splitting the route by document type is usually the biggest single saving available on an AWS document pipeline, and we work through the other levers in reducing AWS document extraction costs.
Bedrock Data Automation charges $0.0100 per page for standard document output and $0.0400 per page for custom output driven by a blueprint, plus $0.0005 per additional field per page. That is $10.00 and $40.00 per 1,000 pages. Read against Textract's plain text rate it looks like a 6.7x penalty, which is how every article frames it.
Read against what it replaces, the picture inverts. BDA Standard returns text, markdown, tables and CSV in one call for $10.00. The nearest Textract equivalent, AnalyzeDocument with TABLES, is $15.00. BDA Custom returns your named fields for $40.00; Textract FORMS returns key-value pairs you still have to map for $50.00. Below a million pages, BDA is the cheaper of the two on structured output, and almost nobody says so.
The real problem with BDA is the shape of its price, not the level. In Amazon's Bedrock price feed, offer file 20260901205051 published September 1, 2026, every BDA document rate is a single price dimension running from zero to infinity. There is no tier. Textract tiers everything at a million pages. So the gap against plain text detection widens from 6.7x to 16.7x exactly at the volume where a buyer most needs it to narrow.
| Pages | Textract text | BDA standard | Textract tables | Textract forms |
|---|---|---|---|---|
| 100,000 | $150 | $1,000 | $1,500 | $5,000 |
| 500,000 | $750 | $5,000 | $7,500 | $25,000 |
| 1,000,000 | $1,500 | $10,000 | $15,000 | $50,000 |
| 5,000,000 | $3,000 | $50,000 | $50,000 | $200,000 |
| 10,000,000 | $6,000 | $100,000 | $100,000 | $400,000 |
Tiered rates applied where they exist: Textract text drops to $0.60 per 1,000 above a million pages, tables to $10.00, forms to $40.00. BDA stays at $10.00 throughout. At five million pages BDA and Textract TABLES cost the same $50,000, and above that Textract keeps falling while BDA does not. Full BDA rate card on the Bedrock Data Automation pricing page.
Textract puts a Confidence value on every Block it returns. Amazon's API reference defines it as "The confidence score that Amazon Textract has in the accuracy of the recognized text and the accuracy of the geometry points around the recognized text", a float with a valid range of 0 to 100. Teams build their entire human-review queue on that number: above the threshold auto-approve, below it route to a person. We counted the word "confidence" in Amazon's Bedrock Data Automation document-output documentation again on September 9, 2026 and it appears zero times across 23,637 characters of rendered text. A team that migrates from Textract to BDA for the structured output silently loses the input its approval logic runs on, and rebuilding it means scoring one model's output with a second model. The thresholds themselves, and why a 98% per-field pass rate still sends a third of your invoices to a human, are worked through on the OCR confidence score page.
Searching for AWS intelligent document processing turns up solution guides, reference architectures and a GitHub accelerator, which is the honest answer: AWS sells the parts and you assemble them. That matters for costing, because the per-page rate on this page is only one line of the bill. The rest looks like this.
S3 for the documents, plus whatever puts them there: an SFTP transfer family endpoint, an email intake, an upload form. Storage is cheap; the request charges on a high-volume pipeline are not nothing.
Textract or Bedrock Data Automation, at the per-page rates above. This is the line everybody budgets and it is frequently not the largest one.
Lambda or Step Functions to drive the calls, retries and pagination, plus the code that turns Block objects into fields your database understands. This is engineering time, and it recurs every time a vendor changes a layout.
A queue, a UI and the people staffing it. On AWS this is Augmented AI or something you build. Its size is set by your confidence threshold, which is why losing the confidence score changes the cost of the whole pipeline, not just one call.
Step three is the one that surprises people. Textract returns an array of Block objects with types like WORD, LINE, CELL, MERGED_CELL, KEY_VALUE_SET and QUERY_RESULT, connected by a relationships graph. Getting from that to a row in your ledger is a parser, and it has a well-known sharp edge: Amazon's documentation states that a CELL block always reports a row span and column span of one, with real spans living on separate MERGED_CELL blocks, so the obvious loop over cells silently flattens every merged header and shifts the columns after it. We wrote that up in detail on whether Textract outputs CSV, and the broader build-or-buy arithmetic sits on build or buy for document extraction.
Above a million pages a month Textract DetectDocumentText falls to $0.60 per 1,000 pages, which is the cheapest published document OCR rate anywhere. If all you need is characters off a page and the documents already sit in S3, nothing on this page beats it and we will not pretend otherwise.
Stay on Textract
AnalyzeID is the only Textract operation whose volume discount starts at 100,000 pages instead of a million, dropping from $25.00 to $10.00 per 1,000. A team processing driver licenses at that scale gets a tier nobody else on AWS gets.
Stay on Textract
If Step Functions, Lambda, SQS and a human review queue are already built and running, the migration cost of moving extraction elsewhere is real and the per-page saving may not cover it. Price the switch honestly, not just the rates.
Stay on AWS
Textract returns geometry and relationships, not values. Getting from a KEY_VALUE_SET array to invoice_total is a parser you write, test against every vendor layout you receive, and maintain forever. That engineering cost is invisible in a per-page comparison and it is usually the largest line in the real budget.
Compare an extraction API
Bedrock Data Automation returns structured output and no confidence score. If your approve-or-review decision runs on a per-field threshold, BDA removes the input that decision needs, and rebuilding it means a second model scoring the first one.
Compare an extraction API
A pipeline that routes every page through FORMS because one document type needs key-value pairs pays $50.00 per 1,000 on all of them. If splitting the routing is more work than moving, the arithmetic favours a service that returns fields at one rate.
Compare an extraction API
If you land on the right-hand column, the two comparisons worth reading next are Amazon Textract alternatives for the vendor view and high-volume OCR pricing for the arithmetic at scale. If your constraint is that documents cannot leave your own network, no cloud service on this page solves it, ours included, and the OCR SDK licensing comparison covers what that actually costs.
Amazon Textract for documents, and nothing else. Rekognition reads text in photographs and caps at 100 words per image with no PDF input. Comprehend does not read images at all, it analyzes text you already extracted. Bedrock Data Automation is the newer document service and is worth pricing against Textract, but for plain page-to-text on AWS, Textract DetectDocumentText at $1.50 per 1,000 pages is the answer.
Between $1.50 and $80.00, on the same service, for the same page. Amazon's price list feed puts Textract DetectDocumentText at $1.50 per 1,000 pages in US East. Add the FORMS feature and it becomes $50.00. Ask for FORMS, Custom Queries and TABLES in one call and it is $80.00. Bedrock Data Automation charges a flat $10.00 per 1,000 pages for standard document output.
Only against Textract's cheapest operation. The comparison everybody repeats, $1.50 against $10.00, puts BDA's full document output next to plain text detection. Against the Textract operations that return the same thing, BDA Standard at $10.00 per 1,000 pages is cheaper than Textract TABLES at $15.00 and far cheaper than Textract FORMS at $50.00, for the first million pages.
No. We counted again on September 9, 2026: the word "confidence" appears zero times in Amazon's BDA document-output documentation, across 23,637 characters of text. Textract puts a Confidence float on every single Block, valid range 0 to 100. If your pipeline routes documents to a human on a confidence threshold, that threshold has nothing to read after you migrate.
No, and AWS says so in its own quota page. "DetectText can detect up to 100 words in an image", and "Amazon Rekognition supports the PNG and JPEG image formats". A single-page invoice runs well past 100 words and usually arrives as a PDF, which Rekognition will not accept. Rekognition is built for text in photographs: signs, labels, packaging, license plates.
Almost always because a feature flag was added to the call. Going from DetectDocumentText to AnalyzeDocument with FORMS multiplies the page rate by 33, from $1.50 to $50.00 per 1,000 pages, and nothing in the SDK warns you. TABLES and QUERIES are each a 10x step. The bill line item is per page, so the multiplier applies to every page you have ever sent.
Yes, but only past one million pages a month, and only on Textract. Text drops from $1.50 to $0.60 per 1,000, TABLES from $15.00 to $10.00, FORMS from $50.00 to $40.00. AnalyzeID is the exception and the only good one: its discount starts at 100,000 pages, dropping from $25.00 to $10.00 per 1,000. Bedrock Data Automation has no volume tier at all.
It is AWS's umbrella name for a pipeline rather than a product you buy. In practice it means Textract or Bedrock Data Automation for extraction, optionally Comprehend for classification and entity detection, S3 for storage, Lambda or Step Functions for orchestration, and a human review layer you build yourself. There is no single SKU called AWS IDP, which is why costing one is difficult.
Yes, and it is usually a one-line change. Textract prices feature combinations below the sum of the parts, so asking for QUERIES and TABLES in one AnalyzeDocument call costs $20.00 per 1,000 pages while asking in two separate calls costs $30.00. Same result, 50% more money. The same holds for FORMS with QUERIES, $55.00 together against $65.00 split.
Yes, and it labels it. Every Block carries a TextType field whose valid values are HANDWRITING and PRINTED, so you can tell which parts of a page were written by hand and treat them differently. It costs nothing extra: handwriting is covered by the same DetectDocumentText and AnalyzeDocument rates as printed text.
Textract turns pixels into text and structure. Comprehend takes text that already exists and finds entities, sentiment, language and PII in it. They are stages, not alternatives, which is why AWS documents a pattern for sending Textract output into Comprehend. If your document is a scan, Comprehend cannot start until Textract has finished.
Not for a production workload. AWS free tiers on these services are small and time-limited, and any real document pipeline exhausts them in the first week. Budget from the per-page rates instead. The number that matters is your monthly page count multiplied by the rate for the specific feature combination your code sends, not the trial allowance.
Take your monthly pages, decide the exact feature combination each document type needs, and price each type separately. A pipeline that sends every page through FORMS because one document type needs key-value pairs pays $50.00 per 1,000 on all of them. Splitting the routing so only the forms go through FORMS is usually the single largest saving available.
It returns JSON, but not fields. Textract gives you an array of Block objects with types like WORD, LINE, CELL, MERGED_CELL, KEY_VALUE_SET and QUERY_RESULT, plus relationships between them. Turning that into "invoice_total: 1420.00" is code you write and maintain. That reconstruction layer is a real cost that never appears in a per-page rate comparison.
Often yes, and we will say so. If your documents already live in S3, your team knows IAM, and you need plain text at very high volume, Textract at $0.60 per 1,000 pages above a million is hard to beat on price. The case for leaving is when you need typed fields with confidence per field, and you would rather not build and own the reconstruction layer.
DocuOCR has no feature flags to multiply your bill, no Block graph to reconstruct and no missing confidence score. You send a document and get back named fields with a score on each one, at a published per-page rate. Upload something real above and compare the output against what a Textract AnalyzeDocument response would need before it becomes a database row. If plain text at eight figures of volume is genuinely all you need, stay on Textract and spend the afternoon elsewhere.