Why AI Struggles With Puns (and What That Reveals About How Language Models Work)
Language models read text as tokens — chunks of words — rather than as sounds or individual letters, so wordplay that depends on how a word sounds or is spelled is harder for them than wordplay based on meaning. They're good at reproducing familiar puns and double-meaning jokes, and less reliable at inventing new sound-based puns or judging which ones actually land.
We write every pun on this site by hand, which means people regularly ask why we don't just get an AI to do it. It's a fair question — language models are fluent, fast and have read more joke books than any human alive. The honest answer is that puns turn out to be a surprisingly sharp test of what these systems can and can't do.
Understanding why is a useful shortcut to understanding how language models work in general.
A pun is four skills pretending to be one
A working pun needs a word with two meanings or a sound-alike, a familiar phrase for it to hide inside, a situation where both readings make sense, and — the hard part — the judgment to know it's actually funny. Humans do all four at once without noticing. For a language model, each one draws on a different kind of knowledge, and some of that knowledge is surprisingly hard for it to get at.
The model doesn't hear the words
Language models don't read letters, and they certainly don't hear sounds. They read tokens: chunks of text that are often a whole common word, and sometimes a fragment of a longer one. The word "neigh" arrives as a token, and so does "nay". Nothing in those tokens says they sound identical. The model only knows they're related if it has seen enough text where people pointed that out.
You may have seen the same blind spot elsewhere. Asking a model how many times a letter appears in a word has tripped up many systems, because the word arrives as one or two chunks rather than a row of letters. Sound-based and spelling-based wordplay lives in exactly the layer the model can't directly see.
Meaning-based puns are the easy ones
Not all puns depend on sound. "Horses have very stable relationships" works because "stable" has two meanings, and learning the different senses of words is precisely what language models are trained on. Double-meaning puns like this are where AI does best, and the results can be perfectly respectable.
Familiar puns versus new ones
Models have absorbed enormous numbers of existing jokes, so they're excellent at retrieving and remixing familiar ones. That's also the trap. A 2023 study that asked ChatGPT for over a thousand jokes found most of them were repeats of a small handful — largely puns it had clearly met before.
Inventing a genuinely new pun means finding a sound link nobody has used, a host phrase it fits, and a reason for it to exist. That's a search through possibilities the model has little direct evidence for, so it tends to fall back on what it has already seen.
The hardest part is knowing which one is funny
Humour depends on surprise, timing and audience — a pun that kills in a birthday card dies in a condolence message. A model can generate fifty candidate puns in seconds. Choosing the three that actually land, and cutting the rest, is the step that makes a collection worth reading. It's also the step in our own five-step method that takes the longest.
What this tells you about using AI for anything
Puns are a small, fun example of a pattern that shows up across real work with AI tools.
- AI is excellent at volume — brainstorming, first drafts, variations
- It is weaker at judgment — deciding what's good, true or appropriate
- Anything that depends on exact spelling, letters or sounds deserves a double-check
- The human edit is where quality comes from, not an optional extra
Keep reading
Frequently Asked Questions
Can AI write good puns?
It can write passable ones, especially puns built on a word's double meaning and puns it has seen before. It's less reliable at inventing original sound-based puns and at judging which of its ideas are actually funny.
What is a token in AI?
A token is a chunk of text that a language model processes as a single unit — often a whole common word, sometimes part of a word or a punctuation mark. Models read and write in tokens, not in letters or sounds.
Why do AI models sometimes miscount letters in a word?
Because they see the word as one or a few tokens rather than as individual letters, so questions about spelling depend on information the model doesn't directly observe.