Tools
Reference of the nine tools exposed by the DocuTray MCP server, the scope each one requires, their annotations, and example prompts.
The server exposes nine tools. Your assistant picks them on its own based on what you ask; you don't need to name them. Every tool runs against the organization chosen when you connected.
Summary
| Tool | What it does | Scope | Read-only | Destructive | Open world |
|---|---|---|---|---|---|
list_document_types | Lists the document types available to the organization | docutray:read | ✅ | — | — |
get_document_type | Gets a document type and the JSON Schema it extracts | docutray:read | ✅ | — | — |
convert_document | Extracts structured data from a PDF or image | docutray:convert | — | — | ✅ |
get_conversion_status | Gets the status and result of a conversion | docutray:read | ✅ | — | — |
identify_document | Identifies which document type a file belongs to | docutray:convert | — | — | ✅ |
get_identification_status | Gets the result of an identification | docutray:read | ✅ | — | — |
validate_document_data | Validates data against a document type's schema and rules | docutray:read | ✅ | — | — |
search_knowledge_base | Semantic search over one of the organization's knowledge bases | docutray:read | ✅ | — | — |
get_usage | Reports processed pages and conversions for a month | docutray:read | ✅ | — | — |
The last three columns are the MCP tool annotations readOnlyHint,
destructiveHint and openWorldHint. Clients use them to decide when to ask
for your confirmation:
- No tool deletes or overwrites anything.
destructiveHintisfalseeverywhere. - Only two tools change state:
convert_documentandidentify_documentcreate a processing record and consume pages from your monthly quota, so they are not read-only and require thedocutray:convertscope. - Only those two tools are open world: they can download the file from a URL you provide, so they reach arbitrary sites on the internet. No other tool fetches anything from outside DocuTray.
- Three tools send data to an AI provider: open world is not the same as
third-party processing.
convert_documentandidentify_documentsend the document to the AI provider configured for the document type, andsearch_knowledge_basesends your query text to an embedding model (Google Gemini) to find similar passages. These providers process the data only for that request, as DocuTray's subprocessors; see AI providers. The other six tools only read or check your organization's data inside DocuTray.
Document types
A document type is an extraction schema: it defines which fields DocuTray returns for a kind of document (an invoice, a bank statement, a purchase order…). See Document types to learn more, or the guide to create your own.
list_document_types
| Argument | Type | Description |
|---|---|---|
search | string, optional | Filter by name, code or description |
page | integer, optional | Page number, starting at 1 |
limit | integer, optional | Results per page, up to 50 (default 20) |
Returns each type's id, code, name, description, status and whether
it is a public type.
get_document_type
| Argument | Type | Description |
|---|---|---|
document_type | string | Document type id or code |
Returns the type with its conversion_mode and the full json_schema of the
data it extracts.
Conversions
convert_document
| Argument | Type | Description |
|---|---|---|
document_type_code | string | Code of the document type to extract with |
file_url | string, optional | Public http(s) URL of the document |
file_base64 | string, optional | Document content in base64 (raw or data URI) |
mime_type | string, optional | MIME type; required with raw base64 |
filename | string, optional | Original file name, kept for auditing |
document_metadata | object, optional | Key/value pairs stored with the conversion |
Send exactly one of file_url or file_base64. Supported formats are PDF,
JPEG, PNG, GIF, BMP, TIFF and WebP, up to 5 MB.
The conversion runs asynchronously. The tool waits a few seconds: if the
conversion finishes in that time it returns the extracted data; otherwise it
returns the conversion_id with a non-final status, and the assistant can
follow up with get_conversion_status.
get_conversion_status
| Argument | Type | Description |
|---|---|---|
conversion_id | string | The id returned by convert_document |
Returns ENQUEUED, PROCESSING, SUCCESS (with the extracted data) or
ERROR (with the error message). Field bounding boxes are not included; very
large results are truncated with a notice — fetch the full result with the
REST API using the same id.
Identification
identify_document
| Argument | Type | Description |
|---|---|---|
document_type_codes | string[] | Candidate document type codes (1 to 50) |
file_url / file_base64 / mime_type / filename | Same as convert_document |
Starts an identification and returns its identification_id.
get_identification_status
| Argument | Type | Description |
|---|---|---|
identification_id | string | The id returned by identify_document |
When the status is SUCCESS, returns the identified document type and the
alternatives considered.
Validation and knowledge
validate_document_data
| Argument | Type | Description |
|---|---|---|
document_type | string | Document type id or code |
data | object | The data to validate |
Checks the data against the document type's JSON Schema and business rules and returns the errors and warnings found. Nothing is stored.
search_knowledge_base
| Argument | Type | Description |
|---|---|---|
knowledge_base_id | string | Id of the knowledge base |
query | string | Natural-language query (up to 1,000 characters) |
limit | integer, optional | Maximum results, up to 10 (default 5) |
similarity_threshold | number, optional | Minimum similarity from 0 to 1 (default 0.6) |
Returns the most similar entries with their similarity score.
Usage
get_usage
| Argument | Type | Description |
|---|---|---|
year | integer, optional | Year, e.g. 2026 |
month | integer, optional | Month from 1 to 12 |
Pass both year and month, or neither for the current month. Returns the
processed pages and successful conversions for that month. It keeps working
when the monthly quota is exhausted.
Example prompts
- "What document types do I have in DocuTray? Which one fits a supplier invoice?"
- "Extract the data from https://example.com/invoices/0042.pdf with my invoice document type and show the line items as a table."
- "I attached three files. Identify which ones are bank statements and which are invoices."
- "Check whether this JSON is valid for my purchase order type and tell me which fields fail."
- "How many pages has my organization processed this month, and how many in August?"
- "Search my product catalog knowledge base for 'stainless steel bolt M8' and list the closest matches."
Errors
Business errors come back as tool errors with a message the assistant can act on, for example:
- the monthly page quota is exhausted (with the quota, the pages used and the reset date);
- the per-minute rate limit was reached (with when to retry);
- the document type does not exist or is not accessible to the organization;
- the file is too large, in an unsupported format, password-protected, or its URL points to a private address;
- the conversion or identification id does not exist ("Conversion not found"), or belongs to another user or organization ("You do not have access to this conversion", or "…this identification").
See Security, limits and data for the limits.
Connect a client
Step-by-step instructions to add the DocuTray MCP server to Claude, Claude Code, ChatGPT, Cursor, VS Code and Copilot Studio.
Security, limits and data
How the DocuTray MCP server authorizes AI clients, how to revoke access, the limits that apply to each tool, and what happens to the documents you send.