Skip to content
Guide11 min read

AI-Generated Document Detection: Spotting Visual Artifacts

Learn to spot AI-generated fake documents through texture, lighting and MRZ mismatch artifacts โ€” a forensic technique compliance and fraud teams can act on.

CheckFile Team
CheckFile Teamยท
Illustration for AI-Generated Document Detection: Spotting Visual Artifacts โ€” Guide

Summarize this article with

AI-generated document detection relies on spotting visual artifacts left by the generation process itself, not by manual editing. Three families of signal matter: texture and noise patterns too statistically regular for a real scan or photograph; lighting and shadow physics that break down across an image, especially on stamps and embossed seals; and mismatches between what a field shows visually and what it encodes in a machine-readable zone, barcode or QR code. None require the document to "look wrong" โ€” a fake can be visually flawless and still fail on all three.

This article is provided for informational purposes and does not constitute legal or regulatory advice.

What counts as a visual artifact in an AI-generated document

Fraudulent documents produced with generative AI fall into two categories. Fully synthetic documents are built from nothing by a generative model โ€” a Generative Adversarial Network (GAN) or a diffusion model โ€” trained to reproduce a document type (payslip, bank statement, ID card, utility bill) without ever starting from a genuine original. Partially edited documents start from a real scan and use generative AI to replace specific fields โ€” a name, an amount, a photograph โ€” while keeping the original's genuine security features intact elsewhere. The second category is harder to catch precisely because most of the document is real.

A visual artifact, in the forensic sense used here, is a trace left by the generation mechanism rather than by human retouching. A diffusion model builds an image through iterative denoising, starting from random noise and refining it step by step until it resembles the target document type. That statistical process produces pixel-level texture distributions that differ measurably from an optically captured photograph or scanner output, even when the result looks convincing to a reviewer. Recent academic forensics work documents this gap between generated and optically captured pixels in document images specifically, as distinct from the natural photography most AI-detection research has historically focused on (arXiv, Beyond Natural Images: Rethinking AI-Generated Image Detection in Documents).

Detection of synthetic content is deployed as an additional layer of AI-generation signals according to client configuration and sector risk profile, complementing existing structural checks โ€” an approach that sits alongside the broader family of techniques covered in our deepfake document detection guide, which looks at synthetic identity documents more generally, and in our companion piece on Error Level Analysis for document fraud, which targets compression artifacts from conventional photo editing rather than AI generation.

Why this vector is growing, and who it targets

Generative tools capable of producing a convincing payslip, bank statement or ID photograph in under a minute became widely accessible from 2024 onward, while human review capacity has not scaled at the same pace. Manual fraud detection alone catches only 37% of cases, with an average detection delay of 87 days, according to the Association of Certified Fraud Examiners (ACFE, Report to the Nations 2024). That gap predates generative AI, and a document a model produces in seconds does nothing to close it.

Cifas, the UK's fraud prevention community and a body regularly cited by government and financial institutions, documents in its Fraudscape 2026 report that criminals are increasingly using AI and generative tools to create convincing impersonations, fake documents and synthetic identities at speed and scale (Cifas, Fraudscape 2026). ENISA's threat landscape work independently flags AI-generated synthetic content as a growing document-fraud vector since 2024, a trend that increasingly informs the supervisory priorities of financial regulators across Europe (ENISA, Threat Landscape 2024).

The sectors most exposed process high volumes of supporting documents remotely, without a physical encounter with the applicant: consumer lending and asset finance (fabricated payslips and bank statements inflating declared income), insurance (fabricated claims evidence), residential rental (fabricated proof-of-income letters), and recruitment (fabricated qualifications and references). A reviewer has seconds to validate a document a model took roughly as long to produce. In the US, FinCEN Alert FIN-2024-Alert004 (November 2024) similarly warned institutions about generative-AI deepfake fraud circumventing identity verification โ€” a sign of how seriously regulators outside the UK treat this vector (FinCEN Alert FIN-2024-Alert004).

The technical breakdown: texture, lighting, and field-encoding mismatch

Three categories of signal recur in a serious forensic review, each targeting a different weakness in how generative models work.

Texture and statistical noise. A genuine scan or photograph carries sensor noise and print grain physically tied to the capture device and paper stock โ€” irregular by nature. A diffusion-generated document instead tends to produce background textures (security guillochรฉ patterns, watermark-style backgrounds, paper grain) that are statistically too regular: patterns repeat with a periodicity invisible to the eye but immediately visible to a frequency-domain analysis. Some newer models overcorrect and inject uniform artificial noise across the entire image instead โ€” noise that, unlike a real scan, does not vary by region.

Lighting and shadow physics. An authentic photograph of a document resting on a surface respects a single, coherent light source: shadows cast by a fold, a staple or an embossed seal all follow the same direction and intensity. Generative models still struggle to maintain that physical coherence across a full image, particularly around raised or textured elements โ€” stamps, simulated holograms, wet-ink signatures โ€” where contradictory shadow directions, or the complete absence of a cast shadow where one is physically required, give the generation away.

Visible field versus encoded field mismatch. On structured documents โ€” an ID card with a Machine-Readable Zone (MRZ), an invoice with a QR code, a certificate with a barcode โ€” the information must match exactly between what a person reads visually and what is encoded in the machine-readable zone. A model trained on visual appearance routinely fails to generate an MRZ whose check digit is mathematically valid, because the check digit is a deterministic function of the surrounding characters rather than a visual pattern to imitate. The same failure shows up as a visible date of birth that does not match the encoded date, or a document number that differs by a digit from the barcode payload โ€” a mismatch close to impossible to fabricate consistently by hand without knowing the checksum algorithm.

Ready to automate your checks?

Free pilot with your own documents. Results in 48h.

Request a free pilot

Comparison: authentic document markers versus AI-generation signals

Element checked Authentic document (genuine scan or photo) AI-generation artifact signal
Background texture Irregular sensor noise and print grain tied to the physical capture Repetitive patterns with measurable periodicity, or uniform artificial noise
Lighting and shadows Single coherent light source, consistent shadows across the whole image Contradictory shadow directions, missing cast shadows on stamps or staples
MRZ / barcode / QR code Visible field matches encoded field exactly; check digit is valid Visible field diverges from encoded data; check digit fails validation
File metadata Consistent with the declared capture device or editing software Missing, generic, or produced by an image-generation tool
Fine print / microtext Sharp and regular under magnification, consistent with the original print method Blurred, garbled or approximate on the smallest characters under magnification
Cross-document consistency Amounts, dates and identity details match across the rest of the file Discrepancies between the suspect document and the rest of the case file

No single row is reliable alone โ€” a genuine document can have poor metadata, and a well-executed fake can pass one texture check. Cross-referencing several independent signals is what makes the approach actionable, the same reasoning behind Error Level Analysis as a complementary compression-based check.

What compliance and fraud teams are asking

Practitioners on compliance and fraud-prevention forums often ask a version of the same three questions once they start dealing with AI-generated submissions in volume.

The first is whether a document that "looks perfect" can still be AI-generated. It can, and that is the core problem: visual perfection used to correlate with authenticity, and no longer does. A diffusion model's output can be indistinguishable from a genuine scan to a reviewer while still failing a statistical texture check or an MRZ checksum.

The second recurring question is what separates a retouched document from a fully generated one, and whether the two need different detection methods. They do. A retouched document starts from a genuine scan with fields swapped using a conventional editor, leaving localized compression artifacts confined to the edited region โ€” the signal Error Level Analysis is built to surface. A fully generated document never existed physically, so its artifacts are distributed across the entire image as a statistical property of generation, not localized to one patch. Applying a compression check to a fully synthetic document, or a texture check to a localized edit, misses the signal each method actually targets.

The third question is operational: how do you apply this at volume without a human examining every submission under magnification? These checks run as an automated first pass โ€” flagging texture anomalies, shadow inconsistencies and MRZ mismatches across every incoming document โ€” so review time concentrates on the cases the automated layer actually surfaces, rather than spreading thinly across a queue where most documents are genuine.

UK regulatory framework

In the UK, submitting a fabricated payslip, bank statement or identity document to obtain a financial benefit โ€” a loan, an insurance payout, a tenancy, a job offer โ€” falls within fraud by false representation under section 2 of the Fraud Act 2006. Where the document itself has been manufactured to appear genuine, the Forgery and Counterfeiting Act 1981 provides an additional basis for prosecution. Neither statute distinguishes between a document forged manually and one produced end-to-end by a generative model; the legal test turns on the false representation and intent to gain, not the method.

The UK is not subject to the EU AI Act post-Brexit. Instead, the government's pro-innovation approach relies on existing sectoral regulators applying principles-based frameworks โ€” the FCA for conduct in regulated financial services, the ICO for data protection where verification processes personal data. Suspected fraud should be reported through Action Fraud or, for organised-crime cases, the National Crime Agency.

Within that framework, a verification platform such as CheckFile gives compliance teams AI-generation signals as a complement to your existing controls, without claiming to catch every possible forgery on its own. Our dedicated page on AI-generated document and deepfake detection sets out how this layer fits into a broader verification workflow, applicable to banking KYC as much as to lending and leasing onboarding, with pricing by volume on our pricing page.

Frequently Asked Questions

Can a visually perfect document still be AI-generated?

Yes โ€” this is the central challenge here. Diffusion models can produce images indistinguishable from a genuine scan to the human eye. The artifacts that give them away are usually statistical (texture distribution, MRZ checksum validity) rather than something a reviewer can spot by looking harder.

What's the difference between a retouched document and a fully AI-generated one?

A retouched document starts from a genuine scan with specific fields edited, leaving compression artifacts localized to the edited area. A fully generated document was never a real scan, so its artifacts are distributed statistically across the whole image โ€” which is why texture analysis and Error Level Analysis are complementary, not interchangeable.

Does the MRZ or barcode really reveal an AI-generated document?

It is one of the more reliable signals available. A generative model trained on a document's visual appearance frequently fails to produce a Machine-Readable Zone with a mathematically valid check digit, or encodes data that does not match what is printed โ€” a discrepancy hard to fabricate manually without knowing the checksum algorithm.

Which industries are most exposed to AI-generated document fraud?

Lending and asset finance, insurance claims, residential rental and recruitment are the most exposed, because each processes high volumes of supporting documents submitted remotely with little or no physical contact with the applicant, leaving a narrow review window per case.

What UK law applies to submitting an AI-generated fake document?

Submitting a fabricated document to obtain a financial or material benefit falls under fraud by false representation, section 2 of the Fraud Act 2006, with the Forgery and Counterfeiting Act 1981 potentially applying to the manufacture of the document itself. Neither law depends on whether the document was forged manually or generated by AI.

In summary

Texture and noise irregularities, shadow and lighting inconsistencies, and mismatches between visible and machine-readable fields are three distinct, complementary signals for catching AI-generated documents โ€” each catching failure modes the others miss, none reliable alone at operational volume. To see how this layer fits alongside metadata checks, Error Level Analysis and cross-document review, visit AI-generated document and deepfake detection, compare plans on the pricing page, or start from our broader document verification guide.

Stay informed

Get our compliance insights and practical guides delivered to your inbox.

Ready to automate your checks?

Free pilot with your own documents. Results in 48h.