On This Page

A Madrid developer API that converts PDFs, spreadsheets, Word files, images, text and audio into JSON that follows a customer-defined schema, attaches a source location to every value, and keeps processed documents queryable for AI agents.

Claix homepage with the headline Turn Unstructured Data Into Structured Intelligence and six file-to-JSON endpoints

100Free successful calls at signup
€0.15Per call after that, any page count
€0.03Per context query, up to 5 questions
3Access routes: REST, MCP, A2A

Overview

Claix sells data extraction as a metered API for developers who build agents and automations rather than for operations teams who review documents. A file goes in through one of six format-specific endpoints, a JSON schema defines the output, and each field comes back as a value paired with where it was found: a page and fragment in a PDF, a column and row in a spreadsheet. Processed documents can be kept and queried later by ID, alone or grouped in a "Knowledge Space" for questions that span several files.

The business is very young and very small. The terms of service name the operator as Gael Anaya Carballo, trading as Claix, a sole trader in Madrid working under Spanish law. The GitHub organization was opened in August 2026, the MCP server repository in September 2026, and the blog, which runs back to July 2026, is written by the founder. There's no announced funding, no named customer and no published accuracy figure. What Claix does publish is unusually specific for its size: a flat price per call, a full subprocessor list and a vendor-written comparison that says plainly where Reducto is the better choice.

The buyer is a developer or small team wiring documents into an agent, an n8n or Make workflow, or a backend service who wants to avoid running OCR, chunking and a vector database. Comparable APIs in this IDP vendor directory include Reducto, LlamaParse, Unstructured, Parseur and the Madrid platform anyformat, which also returns field-level citations but targets enterprise workflows.

Source tracing instead of a confidence score

The main design choice is how Claix signals doubt. Many extraction APIs attach a confidence score to each field. Claix argues against that: a score of 0.98 says nothing about where the number came from. With source verification switched on, every field is returned as a { value, source } pair:

Input What the source points to
Spreadsheet Column name and row number
PDF, Word, image Page, paragraph, clause, table or quoted fragment
Knowledge Space query Document IDs, file names and the cross-document reasoning path

When the model can't locate a value, Claix doesn't invent a citation. The source is set to requires_human_revision and the value is left null, so a workflow can route that field to a person instead of passing on a plausible guess. That is a sound pattern for quality verification in agent pipelines, and it is easier to audit than a probability. The limits are stated by the vendor as well: a citation makes a value checkable, not correct, and Claix publishes no measurement of how often values or citations are wrong.

Six endpoints and three ways in

Each format has its own endpoint under https://claix.dev/api: excel-json for XLSX and CSV, pdf-json for native and scanned PDFs, doc-json for Word, img-json for PNG, JPEG and WebP, txt-json for text, HTML, Markdown and XML, and audio-json for MP3, WAV, M4A and OGG recordings. Requests are multipart uploads with a schema_id, authenticated with an API key in an x-api-key or Bearer header. The Excel documentation is candid about small constraints, such as only the first sheet of a workbook being processed, and lists which error codes are safe to retry. An OpenAPI file is published.

Beyond REST, Claix ships an MCP server so Claude Desktop, Cursor, Windsurf and other MCP clients can extract and query documents as tools, and an A2A endpoint with a public Agent Card at /.well-known/agent.json so other agents can discover and delegate to it. An "agent mode" accepts a natural-language instruction instead of a fixed schema. For low-code users, the blog walks through calling the API from n8n and Make. This puts Claix firmly in the agentic end of the market: the documentation assumes the caller is software.

Knowledge Spaces: document memory without a vector database

Documents can be processed and discarded, kept temporarily, or kept persistently. Persisted documents get a document_id and can be queried later with up to five questions per call, at €0.03, without being sent again. Grouping documents under a space_id lets a query compare, sum or reconcile across files, for example a supplier contract, its invoices and purchase orders. Documents can be added to, removed from or replaced in a space without changing their ID, and those management calls are free.

Claix pitches this as an alternative to stuffing whole PDFs into a prompt or building a retrieval stack. It doesn't document how retrieval works inside a space, how large a space can grow, or how answers are affected as the number of documents rises. Teams planning to put hundreds of files in one space should test that before relying on it.

Pricing: flat per call, or bring your own key

Pricing is the most transparent part of the offer. Signup includes 100 successful calls with no credit card. After that, each successful call costs €0.15 regardless of file size or page count, and a context query costs €0.03. Failed calls, whether authentication errors, validation errors or server errors, are not billed. Under bring-your-own-key (BYOK), the customer supplies an OpenAI, Gemini, Anthropic or xAI key, pays that provider's token costs directly, and Claix charges no processing fee.

A flat per-call price favors long documents: a 60-page contract costs the same as a one-page receipt. Buyers comparing against per-page APIs should check two things. The pricing figures in older Claix blog posts, such as €0.10 to €0.15 per document in the Reducto comparison, differ from the current homepage, so the live pricing section is the one to rely on. And the BYOK route, being free, has no stated service level; how Claix funds that tier over time isn't explained.

Data handling: EU hosting, US model providers

The homepage states data is hosted in Stockholm, with encryption in transit and at rest, API-key authentication and tenant isolation. Customers choose retention per document: instant deletion, temporary, or persistent. Claix says neither it nor its providers train on customer documents.

The data processing agreement fills in the rest. Managed AI runs on Google Gemini, with OpenAI and Anthropic listed as contingency providers; Supabase provides the database and authentication in the EEA, AWS provides cloud infrastructure depending on region, and internal workflows run on a self-hosted n8n instance in Frankfurt. Transfers outside the EEA rely on standard contractual clauses and the EU-US Data Privacy Framework. For security and compliance reviews, that means document content can reach US model providers even though storage sits in the EU, and it means the processor is a single individual rather than a company. No ISO 27001 or SOC 2 report is claimed, and there's no on-premises or VPC option. Claix's own Reducto comparison says so directly.

Use cases

Claix publishes no customer case studies, so the use cases below come from its documentation and blog examples rather than from named deployments.

Invoice and receipt extraction in automations

The homepage example is an invoice returned with number, supplier, taxable base and total, each with its source. The typical setup is an n8n or Make workflow that sends each incoming file to the matching endpoint and writes the JSON to a database, CRM or spreadsheet, with fields marked requires_human_revision routed to an approval step.

Spreadsheet normalization

The Excel endpoint maps columns from arbitrary supplier or customer workbooks onto a fixed schema, recognizing synonyms, abbreviations and translated headers. This is where Claix started: its first blog post, in July 2026, is about Excel-to-JSON mapping.

Agent memory over contracts and supporting documents

An agent processes a set of related files once, stores them in a Knowledge Space, and answers later questions, such as whether invoiced amounts exceed a contract ceiling, without resending the documents or holding them in its context window.

Technical specifications

Feature Specification
Input formats PDF (native and scanned), XLSX, CSV, DOC, DOCX, PNG, JPEG, WebP, TXT, HTML, Markdown, XML, MP3, WAV, M4A, OGG
Output JSON validated against a customer schema; optional value and source pair per field
Missing evidence Value null, source requires_human_revision
Document memory Per-document queries by document_id; cross-document queries by space_id
Access REST API with OpenAPI spec, MCP server, A2A endpoint with Agent Card
Authentication API key in x-api-key or Bearer header
Models Google Gemini by default; OpenAI and Anthropic as contingency; BYOK with OpenAI, Gemini, Anthropic or xAI
Hosting Stockholm per homepage; Supabase in the EEA, AWS, self-hosted n8n in Frankfurt
Retention Instant deletion, temporary or persistent, chosen by the customer
Deployment Managed cloud only; no on-premises or private cloud
Certifications None claimed; GDPR processor agreement
Pricing 100 successful calls free; €0.15 per call; €0.03 per context query; BYOK without Claix fee

Resources

Company information

Claix is operated by Gael Anaya Carballo as a sole trader
Madrid, Spain
info@claix.dev
claix.dev

Because the contracting party is an individual rather than a company, buyers should look closely at the liability, continuity and data-return terms before putting Claix into a production workflow, and keep their schemas and prompts portable. The BYOK option and the OpenAPI spec make switching away easier than with most closed APIs.