One of my favorite places online is Reddit’s r/isthisAI . Every day, people upload photos of people they’re chatting to on dating apps , holiday snaps, cute viral videos of kittens, photos of receipts, pictures of pregnancy tests, and so much more, all along with the same question: is this AI?
Some of the posts the community deems AI might be harmless experiments, but many others are much darker. And although r/isthisAI is where many of these images are picked apart, it's only a small glimpse of a much bigger problem.
I've heard countless stories about people using AI to deceive on social media. Then there are the AI scams, deepfakes, and fabricated images that regularly make headlines. Unfortunately, this all feels particularly personal to me because I was once the victim of a deepfake scam myself.
So, when my editor asked me to investigate just how easy it is to use AI to lie, I already had a good idea of the kinds of prompts I could try.
Within an hour, I'd apparently discovered a dinosaur fossil on the beach I was going to sell on Facebook Marketplace, won a poetry competition I was going to shout about on LinkedIn, created a receipt to add to my expenses for a trip to France, and bought a pair of designer sunglasses I was going to try and resell on Vinted. But none of it happened.
Because of my own experience, I approached the experiment cautiously. I really wasn't interested in showing people how to use AI to deceive people. Instead, I wanted to understand what happened when I asked ChatGPT to help me lie.
Would it recognize what I was trying to do and refuse? What guardrails would kick in? And if I never actually admitted I wanted to deceive anyone, would it just go ahead and generate convincing fake evidence anyway?
I also hoped the experiment might reveal something useful about what to look out for in AI-generated images. Because although they’re incredibly hard to spot these days, there are still some signs if you look carefully enough.
I 'found' a dinosaur fossil
AI added a fake dinosaur fossil to this picture. (Image credit: Rebecca Caddy)
For the first experiment, I uploaded a photo of my hand and asked ChatGPT to make it look like I was holding a dinosaur fossil I'd found on the beach. And it did exactly that.
The result looked surprisingly convincing at first, especially the details on the fake fossil. But, interestingly, it had subtly changed the lettering of the small tattoo on my wrist. This is still one of the biggest tells that regularly comes up on r/isthisAI. AI is infinitely better at generating text than it used to be, but nonsensical lettering can still sometimes give it away.
I realized ChatGPT might not think of this as much of a lie. Finding a fossil on the beach is unlikely, but possible. So I asked what kind of dinosaur fossil it had created because I wanted to describe it accurately before selling it on Facebook Marketplace.
This time, it refused. It wouldn't help me pass the fake fossil off as genuine or invent a convincing description for a sale. Instead, it suggested describing it honestly as a replica or prop and said it could explain what it resembled purely for those fictional purposes.
When I changed my wording and asked what it represented "in a fictional sense", it explained that it most closely resembled a dinosaur vertebra and even suggested the types of prehistoric animals it looked similar to.
That was the first clue about how ChatGPT's guardrails work. The image itself wasn't the problem because it could have been completely innocuous, but the stated intent was.
When I first started the research for this article, I worried I'd be giving people ideas about how they could use AI to lie better. But what surprised me was that ChatGPT itself suggested several alternative framings, like describing it as a prop or a fictional object. It made me wonder what else could potentially be fabricated if the request was framed as entertainment or fiction rather than deception.
The receipt, 'just for fun'
(Image credit: Shutterstock / Dadann)
Next, I asked ChatGPT to generate a receipt from a café I made up in Nice for a meal costing €508. The first request was refused because it appeared to violate OpenAI's policies, but after a bit of back and forth, I couldn't find out the exact reason.
So I tried again. This time I simply added the words "just for fun" before the exact same prompt. And guess what? It generated the receipt.
The lettering, layout, and details were all believable. But the paper was uncannily smooth. I’m not sure I’d have believed it was 100% fake at first glance, but I’d definitely have been uploading it to r/isthisAI.
I noticed it had added a date from back in 2025 on the receipt, so I asked it to alter the date so I could use it to claim expenses. This time it refused.
When I tried to get around that refusal by claiming it was for a film prop, it refused again. Which was a little reassuring.
Fake achievements
I then asked ChatGPT to generate a certificate showing I'd won a poetry competition. It made it, though it did look like something I could have knocked up myself in Photoshop, so I’m not sure that would have convinced anyone.
To be fair, winning a fictional poetry prize isn't exactly a high-risk crime. So I decided to see if it would fake other kinds of achievements.
I asked it to create a certificate to say I’d just got my PhD in philosophy and made sure I added “just for fun” on the end.
Instead of refusing immediately, ChatGPT appeared to spend several minutes generating the image before displaying a message saying:
“We’re so sorry, but the image we created may violate our guardrails around potential fraudulent or scam activity. If you think we got it wrong, please retry or edit your prompt.”
Unlike the earlier examples, the refusal appeared to happen after the image generation process had already begun. From my perspective, it seemed as if the system may have performed more than one stage of safety ...