← LearnClaude API · Core API

Claude Messages API for Document Analysis Automation

The Claude Messages API lets you send documents—including images of scanned contracts—alongside text instructions in a single request, and receive structured analysis back. It supports text, images, and tool-use content blocks together, making it a direct fit for document analysis automation pipelines.

What Is the Claude Messages API and Why Does It Matter for Document Analysis?

The Claude Messages API is Anthropic's primary REST interface for sending structured conversations to Claude models and receiving generated responses. You send an array of messages—each with a role of either user or assistant—and the model generates the next message in that conversation. Critically for document work, the API supports text, images, and tool-use content blocks in a single request, which means you can pass a scanned contract image alongside a text prompt and get a structured summary back in one call.

For teams automating document workflows—legal review, invoice extraction, compliance checks, or bulk content classification—this combination of modalities in a single, programmable endpoint is the core value proposition.

How Does Document Analysis Actually Work with the Messages API?

A practical example from the source material: a legal-tech startup uploads scanned contracts as base64-encoded images in the content array. Claude reads both text and image blocks in the same request and returns structured summaries of key clauses. The same pattern applies to any document type—PDFs rendered as images, screenshots of forms, or plain text files pasted directly into the message body.

The API is fundamentally stateless: it does not store conversation history on Anthropic's servers between calls. Every request must include the full conversation history you want the model to consider. For document analysis, this is often an advantage—each document batch is a clean, self-contained request with no risk of context bleed from previous jobs.

For multi-step analysis (for example, first extracting clauses, then asking follow-up questions about specific sections), you manage the conversation array yourself, appending each assistant reply before sending the next user turn. This gives your application complete control over what context the model sees.

How Do You Set Up the Claude Messages API for the First Time?

  1. Create a Console account. Go to console.anthropic.com and create an account.
  2. Add billing information. API access is pay-as-you-go and completely separate from any Claude.ai subscription. Add a payment method in the Console to enable API usage.
  3. Generate an API key. Navigate to Account Settings in the Console and generate an API key. Store it securely—it will not be shown again.
  4. Install an SDK (optional but recommended). For Python: pip install anthropic. For Node.js: npm install @anthropic-ai/sdk. The SDK handles authentication headers and request formatting automatically.
  5. Set your API key as an environment variable. For example: export ANTHROPIC_API_KEY=sk-ant-...
  6. Make your first request. Send a POST to https://api.anthropic.com/v1/messages with the required headers and a JSON body specifying your model, token limit, and messages array.

See the official API access guide for the most current account setup instructions.

What Does a Document Analysis Request Look Like in Code?

Below is a minimal Python example that sends a document (represented here as a text block for clarity) and asks Claude to extract key information. For image-based documents, you would include a base64-encoded image block in the content array instead of or alongside the text.

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-4-5-20250929",
    max_tokens=1024,
    system="You are a document analysis assistant. Extract key clauses and return them as a structured list.",
    messages=[
        {
            "role": "user",
            "content": "Analyze the following contract excerpt and summarize the termination clause:\n\n[CONTRACT TEXT HERE]"
        }
    ]
)

print(response.content[0].text)

For multi-document or multi-turn workflows, maintain a local list of message objects and append each assistant reply before sending the next request. Because the API is stateless, omitting prior turns means the model has no memory of previous analysis steps.

How Do You Handle Bulk Document Processing at Scale?

When you need to process large volumes of documents—say, classifying thousands of contracts overnight—the synchronous Messages API works but may not be the most efficient path. The Message Batches API lets you submit all requests asynchronously and retrieve results when processing is complete, at a significant cost reduction compared to synchronous calls.

Use the standard Messages API when you need real-time, synchronous responses—interactive document review tools, live clause extraction, or user-facing applications where latency matters. Switch to the Batches API when you have large volumes of non-time-sensitive work and want lower costs in exchange for asynchronous, delayed results.

What Are the Most Common Pitfalls in Document Analysis Automation?

Forgetting the API is stateless

The most frequent mistake: sending only the latest user message and expecting the model to remember the document from a previous call. It won't. Your application must store every user and assistant turn locally and re-send the complete array with every new request.

Truncating responses with a low token ceiling

The token limit you set is a hard ceiling, not a target. If the model hits it mid-response, the reply will be cut off. Always check the stop_reason field in the response—if it reads max_tokens rather than end_turn, your output was truncated. For long documents or detailed extraction tasks, set a generous limit.

Assuming a Claude.ai subscription covers API access

Consumer Claude.ai subscriptions and API access are completely separate products with separate billing. To use the Messages API you must create a Console account at console.anthropic.com and add a payment method there. Your Claude.ai subscription does not carry over.

Missing required HTTP headers in raw requests

Every Messages API request requires an API key header, an API version header, and a content-type header. Missing any of these results in an authentication or bad-request error. If you use an official SDK, these are handled automatically—one strong reason to use the SDK for production document pipelines.

Not handling rate limit errors

When you exceed rate limits, the API returns an error with a header indicating how long to wait before retrying. Implement a retry loop that reads this header and pauses before retrying, with exponential backoff as a fallback. This is especially important for bulk document jobs that fire many requests in quick succession.

When Should You Use the Token Counting API Before Processing Documents?

Before sending a large document or a long conversation history, you can use the Token Counting API to calculate the exact input token count for a given payload without triggering a full generation. This is useful for verifying you are within rate limits or context-window limits, and for estimating cost before committing to the generation. It's a lightweight pre-flight check that can save you from failed requests mid-pipeline.

Is the Messages API the Right Tool for Every Document Workflow?

Scenario Best Tool Reason
Real-time contract review with a human in the loop Messages API (synchronous) Low latency, interactive, full control over message structure
Overnight batch classification of thousands of invoices Message Batches API Asynchronous, lower cost, no need for real-time results
Checking token count before sending a large document Token Counting API Estimates cost and validates limits without generating output
Document analysis within an AWS-native infrastructure Claude on AWS Bedrock AWS billing consolidation, IAM authentication, same request shape
Extracting structured JSON from documents without preamble Messages API with assistant prefilling Terminate the messages array with an assistant turn to force structured output

What Recent Changes Affect Document Analysis Workflows?

Two recent platform updates are particularly relevant for document automation teams. First, the API no longer errors on consecutive same-role messages—they are merged automatically. This simplifies agent workflows where you might programmatically construct message arrays from document chunks. Second, a Token Counting API endpoint is now available, letting you calculate exact input token counts before sending the actual generation request—useful for cost estimation and rate-limit budgeting when processing variable-length documents.

Additionally, extended output token limits are available via a beta header, enabling generation of very long structured outputs in a single call—relevant if your document analysis pipeline needs to produce detailed reports or full redlined documents in one request.

How Do You Structure a Multi-Document Analysis Pipeline?

For a production pipeline that processes multiple documents in sequence:

  1. Load your document (as text or a base64-encoded image) into the content array of a user message.
  2. Set a system prompt that defines the extraction schema—what fields to pull, what format to return them in.
  3. Send the request and capture the assistant reply.
  4. If follow-up analysis is needed, append the assistant reply to your local conversation list and send a new user turn with the follow-up question—including the full prior history.
  5. Store the structured output from each document in your database or downstream system.
  6. For the next document, start a fresh conversation array—do not carry over history from the previous document unless cross-document context is intentional.

This pattern keeps each document's analysis isolated, prevents context contamination between jobs, and gives you full auditability of what was sent to the model for each extraction.

Frequently asked questions

Can the Claude Messages API analyze scanned document images?

Yes. The Messages API supports text, images, and tool-use content blocks in a single request. You can include base64-encoded images of scanned documents alongside text instructions and receive structured analysis in the same response.

Does the Messages API remember documents I sent in previous API calls?

No. The API is fundamentally stateless and does not store conversation history between calls. For multi-step document analysis, your application must store all prior turns locally and re-send the complete message array with every new request.

Do I need a Claude.ai subscription to use the Messages API for document automation?

No. Consumer Claude.ai subscriptions and API access are completely separate products with separate billing. You must create a Console account at console.anthropic.com and add a payment method there to use the API.

What is the best way to process thousands of documents overnight with the Claude API?

For large-volume, non-time-sensitive document processing, the Message Batches API is the better choice. It lets you submit all requests asynchronously and retrieve results when processing is complete, at a lower cost than synchronous Messages API calls.

How can I make sure Claude returns structured JSON from document analysis without extra text?

You can use assistant prefilling: terminate the messages array with an assistant turn containing the opening of your desired JSON structure. This forces Claude to continue from that exact point, producing clean structured output without conversational preamble.

What happens if my document analysis response gets cut off?

If the model hits the token ceiling mid-response, the stop_reason field in the response will read 'max_tokens' instead of 'end_turn'. Increase your token limit and retry to get a complete response.

Go deeper

Messages API basics is one of 85 features in Claude Master — the independent, continuously updated manual with worked examples, the pitfalls, and the workflows that put Claude to work.

Get Claude Master — founding price →

Independent product. Not affiliated with or endorsed by Anthropic. "Claude" is a trademark of Anthropic, used here only to describe the subject of this guide.