On This Page

The fake receipts got cheaper

In March 2025, AppZen's expense-audit platform flagged no AI-generated receipts at all. By mid-May 2026 they made up 70.8% of the fraudulent receipts it caught: 1,471 of them, submitted by 745 employees across 174 companies, claiming $148,143 between them.

The share is not the interesting number. The price is. Those AI-generated receipts averaged about $100. The template-based fakes they replaced averaged $182.

A powerful new tool did not push the fraud upmarket. It pushed it down. This is a sample of caught fraud rather than of all expenses, and a detection vendor counted it, so the level is soft. The direction of the price is the finding.

Why the price fell is a reading rather than a measurement. Amounts get recorded and intent does not, and cheap generation would make small-dollar fraud worth attempting even if no approval threshold existed anywhere. Both explanations arrive at the same place. A $100 claim is one a company is far more likely to process without a person opening the file, and whoever submits it is working to stay under a fixed number rather than to convince a reviewer.

The score that opens the gate says nothing about authenticity

Automated document pipelines gate on a confidence score. Documentation from eleven major platforms agrees on what that score measures, and none of it is authenticity.

AWS Textract defines it as "a number between 0 and 100 that indicates the probability that a given prediction is correct". Azure AI Document Intelligence gives "an estimated probability between 0 and 1 that the prediction is correct", adding that a value of 0.95 means the prediction is "likely correct 19 out of 20 times". Google Document AI describes "how strongly your model associates each entity with the predicted value". Hyperscience is blunter, calling its machine confidence "an internal non-configurable number" and warning that "if a field is transcribed with 0.96 confidence, this does not mean that the model is 96% sure that the data was transcribed correctly". Rossum goes furthest toward a real guarantee, stating it calibrates scores "to correspond to the probability of a given value being right", while conceding calibration is "still an unsolved problem in Artificial Intelligence".

Every one of those sentences is about the same question: did the machine read the characters correctly. Across the documentation for Textract, Document AI, Azure, ABBYY Vantage, Hyperscience, Rossum, Nanonets, Instabase, Mistral OCR, Reducto and Docling, the words fraud, tampering and authenticity do not appear in the definition of confidence.

A recent framework paper, RaV-IDP, states the mechanism in one line: "Model-internal confidence scores measure inference certainty, not correspondence to the document." Correctly reading a document and the document being genuine are separate properties, and the pipeline only measures the first. A forged bank statement generated cleanly in a PDF library has nothing wrong with its characters, so data extraction returns a high score, and the high score is what routes it past review.

Nobody has run the experiment that would settle it

The obvious next claim is that the forgery does not merely pass but scores better than the real thing, because a generated PDF is clean by construction while a genuine document has been folded, faxed, photographed on a kitchen table and scanned at an angle.

That claim is unproven. Searching the benchmark literature and the vendor documentation turns up no published comparison of extraction confidence on AI-generated forgeries against genuine degraded scans. The industry benchmarks reading fidelity in enormous detail, as the IDP accuracy work covered here in June shows, and has never run the one test that would establish whether its automation gate is biased toward fakes.

The mechanism is sound and the measurement is missing. Anyone can reason that clean input extracts more cleanly than dirty input, and that a forger controls input quality completely while an honest applicant does not. The size of the effect is the open quantity, and it is what separates a rounding error from the dominant term. These pipelines approve credit applications and expense claims.

Detection accuracy is vendor-published and nobody audits it

Generation has been tested under controlled conditions. The detection firm AI or Not tested 16 image models across 14 vendors in 75 attempts and reported a 92% success rate at producing realistic synthetic government identity documents in June 2026. The more useful finding in that study is about refusals. Every one of the 16 models produced synthetic IDs when the request was reframed as a Know Your Customer review, a compliance evaluation or a security audit, including models that had declined the same request asked directly.

On the detection side the numbers are self-published. Resistant AI states 99.2% accuracy on its product page with no test set, document mix or definition of correctness. Veryfi claims 99.7% fraud detection accuracy alongside a false-positive rate below 0.3%, both set in a comparison table against an unsourced industry average. Ocrolus heads a callout in its Detect documentation "Watch out for false positives!", telling customers that "some flagged documents may have legitimate explanations" and to review flagged activity manually, without attaching a number to it.

The false-positive rate is what prices the product, because every genuine document wrongly flagged becomes manual review work. Only Veryfi publishes one, inside its own marketing comparison. No independent benchmark of financial-document fraud detection turned up for any of the three.

Flattening the file beats the forensics

The forensic checks that do work mostly ignore the content. PDF incremental-update history records every save the file has been through and is hard to strip without rebuilding the document. Font subsetting is predictable for a given issuer's generator, so overlaid text tends to break the pattern. Metadata and producer strings betray the wrong editing tool. Error Level Analysis finds re-compression seams where a region was pasted in.

Each has a known defeat, and the defeats are cheaper than the checks. Flattening the document to an image destroys the update history and the font structure. Screenshotting an AI-generated document creates a new file carrying no trace of what produced it. Error Level Analysis is a JPEG technique that degrades on other formats and finds nothing in a file that was never a photograph of anything. A generated document has no compression history to be discontinuous, because it was never a scan.

Arithmetic reconciliation looks like the durable exception, since a forger who edits one number breaks the running balance. It is not. Resistant AI's own blog post on AI bank statement generators, published 18 May 2026 and updated a week later, describes a generated statement plainly: "It can include the required fields. It can use plausible dates. It can make the totals add up." A generator that starts from the intended answer satisfies the consistency check on the way out.

Substitution fits the data better than a surge

The prevalence data does not support a panic.

Veriff's 2026 identity fraud report found document fraud down 13% year over year, even as digitally presented media was "300% more likely to be either entirely AI-generated or otherwise altered" than the year before, and it describes AI-driven fraud as still a small fraction of the total. Read together, those two figures describe criminals moving between channels rather than a new wave arriving on top of the old one. Much of the alarming material in circulation is percentage growth measured from a near-zero base, which produces impressive multiples and says little about exposure. Most of it is published by firms selling detection. That includes the Resistant AI profile published on this site alongside this piece, where a 9.08% high-risk rate is measured on the traffic arriving at companies that had already bought fraud detection.

Substitution is not reassuring, though, because of where the traffic is moving. The forged and doctored documents that fell 13% are the class human reviewers were trained to catch. The digitally generated files that rose are the class an extraction confidence score does not look at, and they are arriving as those scores take over the reviewing. Veriff's figure counts fraud attempts. It does not count which of them cleared a gate that was never checking.

What this means if you are buying

Ask for the false-positive rate in writing. An accuracy figure without it cannot be priced. Where a vendor does publish one, the follow-up questions are which test set produced it and what document mix that set contained.

Measure it on your own traffic. Category benchmarks do not exist, and vendor accuracy figures are drawn from customer populations that are not yours. A pilot on real intake is the only evidence available.

Treat authenticity as a separate purchase from extraction. Every authenticity product found in this research is a distinct SKU, add-on or partner integration layered on an IDP platform. None of the eleven confidence scores surveyed includes it, and no amount of extraction accuracy substitutes for it.

Check what your straight-through rate is actually straight through. The automation rate a platform reports is a statement about reading confidence. It is not a statement about how many unverified documents reached a decision.

Prefer not reading the document at all. Where the data can come from the issuer instead of the applicant, through a bank API or a signed feed, the authenticity question disappears rather than getting solved. C2PA content credentials, the provenance standard built for synthetic media, have essentially no footprint on bank statements, payslips and invoices, and a screenshot of a signed document is an unsigned document.

Note the compliance clock, and note that it moved. The EU AI Act classes credit scoring as high risk and requires human oversight able to override the automated output, which turns what the automation actually verified into a compliance question. That obligation was due to apply from 2 August 2026. Under the Commission's digital omnibus it is being deferred to 2 December 2027 for standalone Annex III systems, credit scoring among them, and the deferral binds only once published in the Official Journal, which had not happened at the time of writing.

Caveats

The AppZen figures come from a vendor measuring fraud it caught, so they describe detected fraud in that customer base rather than all expense fraud, and the 70.8% share is a share of flagged fraudulent receipts rather than of all receipts. The documentation survey establishes that eleven platforms do not claim to check authenticity, which is weaker than establishing that none of them does anything about it internally. Why the average claim value fell is inferred from the amounts, since no published figure records what approval thresholds those 174 companies actually used. The central claim that forgeries extract more confidently than genuine degraded documents is argued from mechanism here and has not been measured by anyone, including this site. Veriff, Resistant AI, AppZen, Ocrolus and Veryfi all sell into this problem, and their numbers are cited as vendor statements throughout rather than as independent findings.