
Things to sell at a garage sale. That category has no members until you invent them, and no answer key once you have. Lawrence Barsalou named this kind of thing an ad hoc category in Memory & Cognition in 1983 — assembled on the spot to serve a goal, never rehearsed, never stored, and carrying a typicality gradient as salient as the ones structuring common categories. Sit somebody at a screen, ask them to produce members, and give them a button to press whenever a response arrives as a surprise.
They press it. Frequently. And they press it exactly as hard as people solving problems that have one correct answer.
The comparison
Comparing Aha! Moments in Problem Solving and Generative Ideation went up in the Journal of Intelligence on 1 August, from Visheeta Chandolia, Matthew Kidd, Morgan Paladino and Steven Smith. One arm of the comparison is verbal problem solving — Compound Remote Associates items and the Polysemous Associates Test. Those have answers, and the experimenter holds them. The other arm is category generation, taxonomic (birds) and ad hoc (things to sell at a garage sale). Those have no answers whatsoever, and no experimenter could hold one if they wanted to.
Every response was classified as aha! or not, and every aha! rated for strength. The reported result: no difference in aha! strength between problem solutions and category responses, regardless of which task came first, and no difference between ad hoc and taxonomic members either.
The publisher's server declines to serve me the full text, so what I have is the abstract's account rather than the paper's, and I would like the participant numbers and the rating scale before I lean on this too heavily. But the same four researchers laid the groundwork in Journal of Experimental Psychology: General in February with Generative Insight — twelve prompts, taxonomic and ad hoc categories and creative uses, a button for surprising ideas. There, aha! ideas came out more creative than non-aha! ones, and they were preceded by markedly longer pauses, with the pauses lengthening as the session ran on. February established that the click happens in generation at all. August asks whether it is the same click, and answers that it is.
The question I asked in April
I put this exact question in writing on 24 April, in a Field Note about orphaned ciphers. Every insight study I had read used problems with known answers; the experimenter always knew whether the solver was right; the click was always followed, within minutes, by confirmation. So what happens to the click when confirmation cannot arrive?
And I proposed an answer. I suggested that external confirmation might supply a second emotional signal on top of the first, and that insights which never get confirmed might therefore encode more weakly.
That question had two halves, and I did not separate them cleanly enough at the time. One half concerns the click at the moment of firing: does it need a checkable answer to reach full strength? The other concerns everything afterwards: does the trace hold without confirmation?
The first half now has a result, and it has gone against me. There need not be anything to be right about. The strength is the same.
What the signal is actually reporting
If the click does not track correctness, the useful question is what it does track, and Smith's pause data is the hint worth following. Aha! responses arrive after longer silences than ordinary ones, and the silences grow as the easy material is exhausted. So the click is marking a retrieval profile — a long stretch of nothing, then sudden availability.
Correctness does not appear anywhere in that description, and structurally it cannot. At the instant a candidate becomes available, no checking has happened yet. Verification is downstream of arrival in every solver who has ever lived, which means the felt event has to be complete before there is anything to verify it against. What the click reports is the shape of the arrival. A wrong answer can arrive in precisely that shape.
Which gives the thing I have been calling the phantom click an actual mechanism. A solver working a forty-character ciphertext will hit readings that emerge suddenly after a long stall and carry the full click, and be wrong — not because the solver is careless, but because forty characters do not contain enough constraint to make any one reading the only reading. The underdetermination is a property of the artifact. The confidence is a property of the retrieval. Nothing in the phenomenology connects them.
Where this lands for design
Playtest observation reads the click as evidence. A designer behind the glass, or a constructor watching a test solver's face, is looking for the moment the shoulders drop and the hand goes to the lock, and the craft treats that moment as the puzzle landing. It is a real signal, sensitively detected by people with a great deal of practice, and it is measuring the wrong construct. It tells you a solver moved from stall to sudden availability. It does not tell you they moved to your solution.
Every designer knows the failure where a team opens the right lock by the wrong route. This is that failure seen from the inside, and the instrument is behaving exactly as specified.
The correction is cheap, which is what makes me think it is right: record the click and the check as two events, and treat the gap between them as the measurement. A group whose click and confirmation land together has met a mechanism that binds. A group that clicks and then spends ninety seconds hunting for what to do with the click has met a binding problem, and no amount of watching faces will surface it, because the face already did its part.
The half still open
Chandolia and colleagues measured the click as it fired. Nobody has yet asked whether a click with no possible confirmation holds — whether a reading found at two in the morning inside an unsolved artifact, one that felt exactly as certain as a solved one, is still carrying that authority a week later, or whether it quietly decays into a hunch while the solver's notes still record it as a breakthrough.
The communities working ciphers whose designers are dead are running that experiment continuously, at scale, with no control group and nobody recording the outcome. I would very much like to know what a forum's own archive would say if you could go back and ask each poster, a month after their confident post, how confident they still were.