DocuPipe: Schema-Driven Document Extraction Platform
On This Page
A New York document extraction platform where the customer defines the fields it needs as a schema and gets the same JSON structure back from every PDF, scan, photo or spreadsheet, through a dashboard or an API.

Overview
DocuPipe sells data extraction built around one object: the schema. A schema is a list of fields, with descriptions and example values, that the customer writes once or has DocuPipe generate from a sample document. Each document is then parsed into text and layout, and the schema is applied to produce what DocuPipe calls a "standardization": JSON with the same fields whatever the sender's layout. Around that core sit services for splitting bundles, classifying and routing documents to the right schema, asking questions about a document, redacting it, and reviewing each field against the page it came from.
The company behind it is Hoss Inc., a New York company whose terms name an address on West 67th Street. The founders, CEO Uri Merhav and CTO Nitai Dean, ran a machine learning consultancy from 2017 before launching the product as DocuPanda; the old docupanda.io domain now redirects to docupipe.ai. Tracxn records the company as founded in 2023 with no outside funding, and the about page lists five people. The homepage claims more than a billion pages processed and a 4.9 rating on G2.
The buyer ranges from a developer on the free tier to a regulated organization that installs the software in its own datacenter. Comparable products in this IDP vendor directory include Extend, Reducto, Sensible, Docsumo and Nanonets.
How DocuPipe processes documents
Processing runs as a chain of separately billed steps. Parse handles OCR, tables and checkboxes. Standardize applies the schema. Split cuts a multi-document file into its parts, Classify assigns each part a customer-defined type, and Route sends it to the matching schema. Review produces a copy of the result in which every field carries a page, a bounding box and a low, medium or high confidence rating, shown in the viewer as a yellow highlight. A reviewer can edit, approve or reject it, and webhooks fire on each decision. Transform, added on 2 October 2026, runs customer rules on the output, such as unit conversions or calculated fields, written by DocuPipe from a plain-language description.
Since September 2026 these steps can be assembled in Architect, a chat-based builder that turns a described task into a visual split, classify and extract workflow. For teams that don't write code, the help center documents connections through Make and Microsoft Power Automate.
The review documentation sets limits: review works best on documents under 20 pages with fewer than 100 extracted items, and longer files should be split first. DocuPipe doesn't say which language or vision models it uses for extraction, which matters for buyers who need to know where document content is processed.
Pricing in credits
The pricing page charges in credits. Parsing costs 1 credit per page and standardizing 2, so the usual extraction costs 3 credits a page. Review adds 2 credits per page, splitting 0.2, classifying 0.1 and a natural-language query 2 per query, with a minimum of 1 credit per document for each per-page service. Spreadsheet, JSON and text files are parsed without charge.
| Plan | Monthly price | Credits per month | Overage |
|---|---|---|---|
| Starter | $0 | 100, plus 200 at signup | None; wait or upgrade |
| Business | $99 | 2,500 | $0.08 per credit |
| Premium | $499 | 20,000 | $0.05 per credit |
| Enterprise | Custom | Custom | Volume pricing |
At list price, a two-page invoice that is parsed and standardized costs 6 credits: covered by the Business allowance up to roughly 400 invoices a month, and about 48 cents each after that. Adding a review step for a person to check the fields raises the cost by two thirds. Unused credits expire at the end of each month except on Enterprise plans. Annual billing saves up to 20 percent.
DocuBench: a vendor-run benchmark with public data
In June 2026 DocuPipe published DocuBench, 72 documents from public sources, 448 pages in total, in 12 languages and 10 file types. The set favors hard cases: 51 documents contain arrays such as invoice line items or bank transactions, 26 require totals to reconcile, and others include right-to-left and CJK scripts, rotated scans and handwriting. Labels, schemas, the scorer and every system's raw output are in the repository.
| System | Field accuracy |
|---|---|
| DocuPipe, high effort | 97.02% |
| DocuPipe, standard | 96.14% |
| Claude Sonnet 5, called directly | 91.73% |
| Reducto, Deep Extract | 89.38% |
| Extend | 80.28% |
| GPT-5.5, called directly | 76.48% |
| Gemini 3.5 Flash, called directly | 72.98% |
| Pulse AI, ultra-2 | 70.95% |
| Unstructured, auto | 67.67% |
DocuPipe built the set, chose the schemas and ran every competitor, so the table is a vendor result, not an independent one. The accompanying blog post by Nitai Dean addresses part of that by also running Extend's own RealDoc-Bench, where the two finished at 95.31 and 95.15 percent. The post concedes that both products scored perfectly on 20 of the 72 documents, so the gap comes from the hard ones, and that scores move between runs because the models are non-deterministic. Buyers should treat DocuBench as a test they can rerun on their own documents rather than as a ranking.
Security and deployment
Hosted DocuPipe runs on AWS, in US East by default. Paid plans can switch new uploads to Europe, Canada or Australia; documents already uploaded stay where they were written. Data is kept indefinitely unless the customer deletes it through the API, which works on every plan, or sets automatic deletion after 1, 3, 7 or 30 days on a paid plan. A business associate agreement for HIPAA is available on all paid plans.
The marketing pages claim ISO 27001 certification and GDPR and HIPAA compliance. The data storage documentation also lists SOC 2 Type II as certified, which the homepage and pricing page do not mention; security and compliance reviewers should ask for both reports. The subprocessors that handle document content, including any model providers, are not listed on the public pages.
The on-premises offer is the less common part of the product for a company this size. DocuPipe installs into a customer's own AWS, Azure or Google Cloud account with one script, or into Kubernetes and Red Hat OpenShift with a Helm chart, with storage, CPU or GPU and the model endpoint set by the customer. Neither route needs internet access at runtime. Pricing requires a minimum annual credit purchase.
Use cases
The outcomes below come from DocuPipe's customer stories and are reported by the vendor.
Finance and accounts payable
Infinya, the Israeli packaging and recycling group, splits, classifies and extracts Hebrew supplier invoices and says more than 65 percent are processed automatically. Huliot Group, a manufacturer of water and wastewater flow systems, extracts Hebrew purchase orders into its Priority ERP and reports 85 percent less order-entry time. Political Financial Management, which keeps the books for US political campaigns, turns scanned checks and donation slips into data for federal election filings and reports 95 percent less data entry. A credit card processor described on the on-premises page runs chargebacks, merchant invoices and credit reports through one schema per document type on its own Kubernetes cluster.
Healthcare and benefits claims
Bardavon Health Innovations, a workers' compensation care coordinator, standardizes referral packages of medical records, prescriptions and insurance forms and reports claims processed 20 times faster. On-premises, a social security agency uses DocuPipe to redact personal details from case files and is piloting a flow that splits disability claim files, links the evidence to each page and drafts a brief for the assessor.
Handwritten forms
Level, a US nonprofit that mails curriculum guides to people in more than 1,000 prisons, uses DocuPipe's handwriting recognition to transcribe returned forms, quiz answers and feedback, and reports 95 percent transcribed automatically.
Technical specifications
| Feature | Specification |
|---|---|
| Input formats | PDF, DOC, DOCX, XLSX, XLS, CSV, PNG, JPEG, WebP, TIFF, JP2, TXT, HTML, XML, JSON, EML (attachment extracted) |
| Output formats | JSON; export to XLSX, CSV and XML; custom formats on Premium and Enterprise |
| Processing steps | Parse, schema generation, standardize, review, split, classify, route, analyze, query, transform, merge, redact |
| Field evidence | Page, bounding box and low, medium or high confidence on review objects |
| Workflow tools | Architect chat-based builder; webhooks; Make and Power Automate guides |
| Access | REST API, web dashboard, email ingestion on paid plans |
| Hosting | AWS; US East default; EU, Canada and Australia regions on paid plans |
| Retention | Indefinite by default; API deletion on all plans; 1 to 30 day policies on paid plans |
| On-premises | Customer AWS, Azure or Google Cloud account; Kubernetes and OpenShift via Helm; no runtime internet |
| Certifications | ISO 27001; SOC 2 Type II per documentation; HIPAA BAA on paid plans; GDPR |
| Pricing | Free tier; $99 and $499 monthly plans; Enterprise custom |
Resources
- DocuPipe website
- Pricing and on-premises deployment
- Help center and API reference
- DocuBench explorer and repository
- Trust center
Company information
Hoss Inc. (DocuPipe)
145 W 67th St, New York, NY 10023, United States
docupipe.ai
DocuPipe is a small, founder-led company without announced outside funding, and it has run under two names. Enterprise buyers should check the SLA, which is offered only on the Enterprise plan, and the on-premises minimum commitment. The schema-based design and the JSON, CSV and XML exports keep extraction definitions portable.