AI-Generated Document Detection: Spotting Visual Artifacts
Learn to spot AI-generated fake documents through texture, lighting and MRZ mismatch artifacts โ a forensic technique Australian compliance and fraud teams can act on.

Summarize this article with
AI-generated document detection relies on spotting visual artifacts left by the generation process itself, not by manual editing. Three families of signal matter: texture and noise patterns too statistically regular for a real scan or photograph; lighting and shadow physics that break down across an image, especially on stamps and embossed seals; and mismatches between what a field shows visually and what it encodes in a machine-readable zone, barcode or QR code. None require the document to "look wrong" โ a fake can be flawless and still fail on all three.
This article is provided for informational purposes and does not constitute legal or regulatory advice.
What counts as a visual artifact in an AI-generated document
Fraudulent documents produced with generative AI fall into two categories. Fully synthetic documents are built from nothing by a generative model โ a Generative Adversarial Network (GAN) or a diffusion model โ trained to reproduce a document type (payslip, bank statement, driver licence, passport) without starting from a genuine original. Partially edited documents start from a real scan and use generative AI to replace specific fields โ a name, an amount, a photograph โ while keeping the original's genuine security features intact elsewhere. The second category is harder to catch precisely because most of the document is real.
A visual artifact, in the forensic sense used here, is a trace left by the generation mechanism rather than by human retouching. A diffusion model builds an image through iterative denoising, starting from random noise and refining it step by step until it resembles the target document type. That statistical process produces pixel-level texture distributions that differ measurably from an optically captured photograph or scanner output, even when the result looks convincing to a reviewer. Recent academic forensics work documents this gap between generated and optically captured pixels in document images specifically, distinct from the natural photography most AI-detection research has historically focused on (arXiv, Beyond Natural Images: Rethinking AI-Generated Image Detection in Documents).
Detection of synthetic content is deployed as an additional layer of AI-generation signals according to client configuration and sector risk profile, complementing existing structural checks โ an approach that sits alongside the broader techniques in our deepfake document detection guide, which covers synthetic identity documents more generally, and our companion piece on Error Level Analysis for document fraud, which targets compression artifacts from conventional editing rather than AI generation.
Why this vector is growing, and who it targets
Generative tools capable of producing a convincing payslip, bank statement or ID photograph in under a minute became widely accessible from 2024 onward, while human review capacity has not scaled at the same pace. Manual fraud detection alone catches only 37% of cases, with an average detection delay of 87 days, according to the Association of Certified Fraud Examiners (ACFE, Report to the Nations 2024). That gap predates generative AI, and a document produced in seconds does nothing to close it.
Australia's markets regulator has already flagged the broader trend. In August 2026, ASIC warned that scammers are using generative AI to "spin vast webs of deception" โ deepfake videos and fabricated endorsements from public figures used to drive investment scams (ASIC, ASIC warns scammers are using AI to spin vast webs of deception, media release 26-195MR). Its own figures back this up: in FY26, ASIC coordinated the removal of more than 19,400 online scams, up 182% on the previous year. That release focuses on deepfake video and investment scams rather than forged documents specifically, but it describes the same underlying shift: generative tools lowering the cost of convincing fake material at scale, of which fabricated payslips and identity documents are an adjacent, fast-growing vector.
The sectors most exposed process high volumes of supporting documents remotely, without a physical encounter with the applicant: consumer lending and asset finance (inflated payslips and bank statements), insurance (fabricated claims evidence), residential rental (fabricated proof-of-income letters), and recruitment (fabricated qualifications and references). A reviewer has seconds to validate a document a model took roughly as long to produce.
The technical breakdown: texture, lighting, and field-encoding mismatch
Three categories of signal recur in a serious forensic review, each targeting a different weakness in how generative models work.
Texture and statistical noise. A genuine scan or photograph carries sensor noise and print grain physically tied to the capture device and paper stock โ irregular by nature. A diffusion-generated document instead tends to produce background textures (security guillochรฉ patterns, watermark-style backgrounds, paper grain) that are statistically too regular: patterns repeat with a periodicity invisible to the eye but visible to a frequency-domain analysis. Some newer models overcorrect and inject uniform artificial noise across the entire image instead โ noise that, unlike a real scan, does not vary by region.
Lighting and shadow physics. An authentic photograph of a document resting on a surface respects a single, coherent light source: shadows cast by a fold, a staple or an embossed seal all follow the same direction and intensity. Generative models still struggle to maintain that physical coherence across a full image, particularly around raised or textured elements โ stamps, simulated holograms, wet-ink signatures โ where contradictory shadow directions, or the complete absence of a cast shadow where one is physically required, give the generation away.
Visible field versus encoded field mismatch. On structured documents โ a passport or driver licence with a Machine-Readable Zone (MRZ), an invoice with a QR code, a certificate with a barcode โ the information must match exactly between what a person reads visually and what is encoded in the machine-readable zone. A model trained on visual appearance routinely fails to generate an MRZ whose check digit is mathematically valid, because the check digit is a deterministic function of the surrounding characters rather than a visual pattern to imitate. The same failure shows up as a document number that differs by a digit from the barcode payload โ a mismatch close to impossible to fabricate consistently by hand without knowing the checksum algorithm.
Ready to automate your checks?
Free pilot with your own documents. Results in 48h.
Request a free pilotComparison: authentic document markers versus AI-generation signals
| Element checked | Authentic document (genuine scan or photo) | AI-generation artifact signal |
|---|---|---|
| Background texture | Irregular sensor noise and print grain tied to the physical capture | Repetitive patterns with measurable periodicity, or uniform artificial noise |
| Lighting and shadows | Single coherent light source, consistent shadows across the whole image | Contradictory shadow directions, missing cast shadows on stamps or staples |
| MRZ / barcode / QR code | Visible field matches encoded field exactly; check digit is valid | Visible field diverges from encoded data; check digit fails validation |
| File metadata | Consistent with the declared capture device or editing software | Missing, generic, or produced by an image-generation tool |
| Fine print / microtext | Sharp and regular under magnification, consistent with the original print method | Blurred, garbled or approximate on the smallest characters under magnification |
| Cross-document consistency | Amounts, dates and identity details match across the rest of the file | Discrepancies between the suspect document and the rest of the case file |
No single row is reliable alone โ a genuine document can have poor metadata, and a well-executed fake can pass one texture check. Cross-referencing several signals is what makes the approach actionable, the same reasoning behind Error Level Analysis as a complementary check.
What compliance and fraud teams are asking
Practitioners on compliance and fraud-prevention forums often ask a version of the same three questions.
The first is whether a document that "looks perfect" can still be AI-generated. It can, and that is the core problem: visual perfection used to correlate with authenticity, and no longer does. A diffusion model's output can be indistinguishable from a genuine scan while still failing a statistical texture check or an MRZ checksum.
The second is what separates a retouched document from a fully generated one. A retouched document starts from a genuine scan with fields swapped using a conventional editor, leaving compression artifacts localized to the edited region โ the signal Error Level Analysis is built to surface. A fully generated document never existed physically, so its artifacts are distributed across the entire image as a statistical property of generation. Applying a compression check to a fully synthetic document, or a texture check to a localized edit, misses the signal each method actually targets.
The third is operational: how do you apply this at volume without a human examining every submission under magnification? These checks run as an automated first pass, flagging texture anomalies, shadow inconsistencies and MRZ mismatches so review time concentrates on the cases the automated layer surfaces.
Australian regulatory framework
In Australia, submitting a fabricated payslip, bank statement or identity document to obtain a benefit โ a loan, an insurance payout, a tenancy, a job offer โ falls within the general dishonesty and fraud offences of the Criminal Code Act 1995 (Cth), covering a financial advantage obtained by deception, alongside equivalent fraud and forgery provisions in state and territory Crimes Acts. None distinguishes between a document forged manually and one produced end-to-end by a generative model; the legal test turns on the false representation and intent to gain, not the method.
For AML/CTF-regulated entities โ banks, lenders, remittance providers and other reporting entities โ the relevant framework is the Anti-Money Laundering and Counter-Terrorism Financing Act 2006 (AML/CTF Act), which sets out customer due diligence and document-verification obligations, overseen by AUSTRAC (the Australian Transaction Reports and Analysis Centre; obligations are set out on AUSTRAC's website). Its guidance on customer identification increasingly assumes a submitted identity document cannot be taken at face value โ exactly the gap AI-generation detection is built to close. ASIC, which regulates conduct in financial services more broadly, has taken an active public posture on AI-enabled fraud, and a company applying for finance can be checked against an ASIC company extract rather than a submitted certificate alone.
Where a verification process handles personal information โ a name, date of birth, Tax File Number, or scanned identity document such as a passport, driver licence or ImmiCard โ the Privacy Act 1988 (Cth) and the Australian Privacy Principles (APPs) govern how that data can be collected, stored and used, overseen by the Office of the Australian Information Commissioner (OAIC). Suspected organised or large-scale fraud should be reported to the Australian Federal Police (AFP).
Within that framework, a verification platform such as CheckFile gives compliance teams AI-generation signals as a complement to your existing controls, without claiming to catch every possible forgery on its own. Our dedicated page on AI-generated document and deepfake detection sets out how this layer fits into a broader verification workflow, applicable to banking KYC as much as to lending and leasing onboarding, with details on how submitted data is handled on our security page and pricing by volume on our pricing page.
Frequently Asked Questions
Can a visually perfect document still be AI-generated?
Yes. Diffusion models can produce images indistinguishable from a genuine scan to the human eye. The artifacts that give them away are usually statistical โ texture distribution, MRZ checksum validity โ rather than something a reviewer can spot by looking harder.
What's the difference between a retouched document and a fully AI-generated one?
A retouched document starts from a genuine scan with specific fields edited, leaving compression artifacts localized to the edited area. A fully generated document was never a real scan, so its artifacts are distributed statistically across the whole image โ which is why texture analysis and Error Level Analysis are complementary, not interchangeable.
Does the MRZ or barcode really reveal an AI-generated document?
It is one of the more reliable signals available. A generative model trained on visual appearance frequently fails to produce a Machine-Readable Zone with a mathematically valid check digit, or encodes data that does not match what is printed โ hard to fabricate manually without knowing the checksum algorithm.
Which industries are most exposed to AI-generated document fraud?
Lending and asset finance, insurance claims, residential rental and recruitment are the most exposed, because each processes high volumes of supporting documents submitted remotely with little or no physical contact with the applicant.
What Australian law applies to submitting an AI-generated fake document?
Submitting a fabricated document to obtain a financial or material benefit can fall under the fraud and general dishonesty offences in the Criminal Code Act 1995 (Cth), or equivalent state and territory Crimes Act provisions, with forgery offences potentially applying to the document's manufacture. For AML/CTF-regulated entities, failing to catch a fabricated identity document during due diligence can also raise compliance issues under the AML/CTF Act 2006, overseen by AUSTRAC. None of this depends on whether the document was forged manually or generated by AI.
In summary
Texture and noise irregularities, shadow and lighting inconsistencies, and visible-versus-encoded field mismatches are three distinct, complementary signals for catching AI-generated documents โ none reliable alone at operational volume. To see how this layer fits alongside metadata checks, Error Level Analysis and cross-document review, visit AI-generated document and deepfake detection, compare plans on the pricing page, or start from our broader document verification guide.
Stay informed
Get our compliance insights and practical guides delivered to your inbox.