DocuTray
MCP Server

Tools

Reference of the nine tools exposed by the DocuTray MCP server, the scope each one requires, their annotations, and example prompts.

The server exposes nine tools. Your assistant picks them on its own based on what you ask; you don't need to name them. Every tool runs against the organization chosen when you connected.

Summary

ToolWhat it doesScopeRead-onlyDestructiveOpen world
list_document_typesLists the document types available to the organizationdocutray:read✅——
get_document_typeGets a document type and the JSON Schema it extractsdocutray:read✅——
convert_documentExtracts structured data from a PDF or imagedocutray:convert——✅
get_conversion_statusGets the status and result of a conversiondocutray:read✅——
identify_documentIdentifies which document type a file belongs todocutray:convert——✅
get_identification_statusGets the result of an identificationdocutray:read✅——
validate_document_dataValidates data against a document type's schema and rulesdocutray:read✅——
search_knowledge_baseSemantic search over one of the organization's knowledge basesdocutray:read✅——
get_usageReports processed pages and conversions for a monthdocutray:read✅——

The last three columns are the MCP tool annotations readOnlyHint, destructiveHint and openWorldHint. Clients use them to decide when to ask for your confirmation:

  • No tool deletes or overwrites anything. destructiveHint is false everywhere.
  • Only two tools change state: convert_document and identify_document create a processing record and consume pages from your monthly quota, so they are not read-only and require the docutray:convert scope.
  • Only those two tools are open world: they can download the file from a URL you provide, so they reach arbitrary sites on the internet. No other tool fetches anything from outside DocuTray.
  • Three tools send data to an AI provider: open world is not the same as third-party processing. convert_document and identify_document send the document to the AI provider configured for the document type, and search_knowledge_base sends your query text to an embedding model (Google Gemini) to find similar passages. These providers process the data only for that request, as DocuTray's subprocessors; see AI providers. The other six tools only read or check your organization's data inside DocuTray.

Document types

A document type is an extraction schema: it defines which fields DocuTray returns for a kind of document (an invoice, a bank statement, a purchase order…). See Document types to learn more, or the guide to create your own.

list_document_types

ArgumentTypeDescription
searchstring, optionalFilter by name, code or description
pageinteger, optionalPage number, starting at 1
limitinteger, optionalResults per page, up to 50 (default 20)

Returns each type's id, code, name, description, status and whether it is a public type.

get_document_type

ArgumentTypeDescription
document_typestringDocument type id or code

Returns the type with its conversion_mode and the full json_schema of the data it extracts.

Conversions

convert_document

ArgumentTypeDescription
document_type_codestringCode of the document type to extract with
file_urlstring, optionalPublic http(s) URL of the document
file_base64string, optionalDocument content in base64 (raw or data URI)
mime_typestring, optionalMIME type; required with raw base64
filenamestring, optionalOriginal file name, kept for auditing
document_metadataobject, optionalKey/value pairs stored with the conversion

Send exactly one of file_url or file_base64. Supported formats are PDF, JPEG, PNG, GIF, BMP, TIFF and WebP, up to 5 MB.

The conversion runs asynchronously. The tool waits a few seconds: if the conversion finishes in that time it returns the extracted data; otherwise it returns the conversion_id with a non-final status, and the assistant can follow up with get_conversion_status.

get_conversion_status

ArgumentTypeDescription
conversion_idstringThe id returned by convert_document

Returns ENQUEUED, PROCESSING, SUCCESS (with the extracted data) or ERROR (with the error message). Field bounding boxes are not included; very large results are truncated with a notice — fetch the full result with the REST API using the same id.

Identification

identify_document

ArgumentTypeDescription
document_type_codesstring[]Candidate document type codes (1 to 50)
file_url / file_base64 / mime_type / filenameSame as convert_document

Starts an identification and returns its identification_id.

get_identification_status

ArgumentTypeDescription
identification_idstringThe id returned by identify_document

When the status is SUCCESS, returns the identified document type and the alternatives considered.

Validation and knowledge

validate_document_data

ArgumentTypeDescription
document_typestringDocument type id or code
dataobjectThe data to validate

Checks the data against the document type's JSON Schema and business rules and returns the errors and warnings found. Nothing is stored.

search_knowledge_base

ArgumentTypeDescription
knowledge_base_idstringId of the knowledge base
querystringNatural-language query (up to 1,000 characters)
limitinteger, optionalMaximum results, up to 10 (default 5)
similarity_thresholdnumber, optionalMinimum similarity from 0 to 1 (default 0.6)

Returns the most similar entries with their similarity score.

Usage

get_usage

ArgumentTypeDescription
yearinteger, optionalYear, e.g. 2026
monthinteger, optionalMonth from 1 to 12

Pass both year and month, or neither for the current month. Returns the processed pages and successful conversions for that month. It keeps working when the monthly quota is exhausted.

Example prompts

  • "What document types do I have in DocuTray? Which one fits a supplier invoice?"
  • "Extract the data from https://example.com/invoices/0042.pdf with my invoice document type and show the line items as a table."
  • "I attached three files. Identify which ones are bank statements and which are invoices."
  • "Check whether this JSON is valid for my purchase order type and tell me which fields fail."
  • "How many pages has my organization processed this month, and how many in August?"
  • "Search my product catalog knowledge base for 'stainless steel bolt M8' and list the closest matches."

Errors

Business errors come back as tool errors with a message the assistant can act on, for example:

  • the monthly page quota is exhausted (with the quota, the pages used and the reset date);
  • the per-minute rate limit was reached (with when to retry);
  • the document type does not exist or is not accessible to the organization;
  • the file is too large, in an unsupported format, password-protected, or its URL points to a private address;
  • the conversion or identification id does not exist ("Conversion not found"), or belongs to another user or organization ("You do not have access to this conversion", or "…this identification").

See Security, limits and data for the limits.

On this page