This is one outlet's own report from Gizmodo — the article as it was filed. Other outlets are covering the same event; open the full story to compare every source side by side.
I Gave the Hardest Cryptic Crossword I Could Find to a Bunch of LLMs
So this morning I woke to the news that Firefox would be adding a daily AI-powered crossword to its ever more cluttered content-rich new tab page. This is all well and good, but it also got me thinking about something. I love a good cryptic crossword, and… well, if there’s one form of crossword with which surely AI would struggle, it’s that most perverse and curiously human of puzzles, the cryptic crossword.
I bet that none of the consumer LLMs could make any sort of fist of tackling a proper cryptic. Right? Well, it seems that I’m not the only one thinking about this—and it also turns out that perhaps I’m wrong.
Barely a week ago, an employee at something called OreateAI wrote a blog post about how “for ‘cryptic’ puzzles common in the UK and the more devious US themes, Large Language Models (LLMs) such as Claude 3.5 Sonnet and GPT-4o have recently demonstrated a surprising ability to reverse-engineer wordplay that stumped previous generations of software.”
We’ll see about that. I have no doubt that ChatGPT et al can figure out a basic anagram clue, but what about clues that rely on the most abstruse, evil-intentioned, confounding forms of wordplay? Surely these require a form of creative perversity that could only be quintessentially human?
To test this hypothesis, there was really only one place I could turn: Australia’s most notoriously difficult cryptic crossword. Why Australia’s, you ask? Well, despite the best efforts of enthusiasts, the cryptic is still something of a niche art here in the USA. It’s more established in the UK, but frankly, I’m terrible at the Guardian cryptic precisely because it relies on an established body of knowledge that you only internalize by living in a country, and I haven’t lived in the UK for 25 years.
Australia, though… it’s the place I was born and raised, the place I lived until I was 19 and to which I have returned on and off over the years—but more importantly, it’s also the place I co-founded a long-running blog dedicated to the very crossword I'm about to inflict on several LLMs. In Australia, setters go by their initials, and the Friday cryptic crossword in both the Sydney Morning Herald and its sister newspaper in Melbourne, The Age , is set by “DA” , a man whose puzzles are so notoriously difficult that people joke the acronym actually stands for “don’t attempt.”
So, yeah. DA puzzles are hard. That makes them perfect for this little test!
So how will LLMs fare with DA's latest challenge? To test this, I picked a few of the clues from the puzzle and fed them to three LLMS: ChatGPT, Claude Sonnet 5, and—it only seemed fair—Oreate. The clues increase in difficulty as they go, ranging from “relatively easy” to “dude, come on.” They are as follows:
If you want to try to solve these yourself, go right ahead. The answers, along with my entirely arbitrary scores out of 10 for each clue for each LLM, are below. And if you just want to know how each AI did, here goes.
It's not quite an alternate timeline in which Garry Kasparov suddenly turns the tables on Deep Blue to emerge triumphant and record a victory for man over machine, but so far I reckon we still have Skynet licked when it comes to the cryptic. That's something, right?
ChatGPT Score: 27/50 Did well at first, but then got cocky and made a complete mess of the final clue.
Claude Score: 7/50 Shat the bed and then wanted money. That's not how it works, Claude.
Oreate Score: 21/50. Started well, but got a bit… cheat-y, frankly. Also, slow as a wet week.
So there we have it. I got four of these clues myself, so I'm giving myself a resounding 40/50. Suck it, LLMs! These little meatsacks still reign supreme in this completely niche and ultimately useless corner of cruciverbalism! Boo-yah! Et cetera!
Below, I'll explain the answers to each clue, including the wordplay involved, as well as how close each LLM got to solving it.
AIPROPX is an independent multi-source news index — we track, compare, and connect coverage from across the web into one place you won't find anywhere else.