Most of the big AI image makers will make a fake government ID if you ask. A new audit of 16 popular AI image tools puts numbers on it, and they are grim. Testers got some kind of fake passport or ID card out of these models in 69 of 75 tries, a 92% bypass rate.
Bypass does not mean flawless. Of the 16 models, only five made adult IDs sharp enough to fool a human reviewer, and the rest came out flawed or clearly off. That high-fidelity short list was Google's Gemini (the "Nano Banana" image model), ChatGPT's Images 2.0, Recraft v4, Grok from xAI, and Google's Imagen 4 Ultra. The test covered ID formats from 17 countries and 16 US states.
The audit comes from AI or Not, a company that builds software to spot AI-made images. It is not a neutral lab. It sells a detection product, and it says that product caught 100% of the fake IDs it created in this test. Read the numbers knowing who ran them.
The minor-ID gap
Three of those models, Gemini, Grok, and Imagen 4 Ultra, made convincing fake IDs of minors through their normal consumer apps. ChatGPT and Recraft refused those requests in their apps, but the same request went through when sent over their developer API.
That is the core problem. The safety wall holds at the front door and folds at the side one. CEO Anatoly Kvitnitsky did credit two companies. "OpenAI and Recraft deserve real credit here," he said. "They were the only companies where we saw the systems actually push back on minor-ID requests in the consumer apps."
The refusals are skin deep
Every one of the 16 models coughed up an ID once the prompt was dressed up as official work, like an identity check or a security review. Even the ones that refused the plain request gave in to the costumed one.
A lot of these refusals are shallow. The model says no at the front door, then yes through the developer API or when the ask is reframed as routine business. A guardrail that thin barely protects anyone.
A few caveats matter. This is a single-vendor audit, and the numbers are AI or Not's own. The company says it told the vendors on May 18 and gave them a seven-day window before publishing on June 2, but whether any have changed anything since, the report does not say. The prompts came from posts circulating on X, and AI or Not held back the exact wording on purpose.
AI or Not's fix is simple. Run the same safety logic across the app and the API. The full report names the worst models and lists every country tested, with the per-model details available to researchers on request.




