
The animals were paperclips shaped like zebras, three centimetres high. The pens were pipe cleaners, twenty and thirty centimetres, in a heap on the left of the table. Seventeen zebras in front, the problem statement on the right, an overhead camera running.
Three participants clipped the zebras onto the pipe cleaners.
It was a perfectly reasonable thing to do with a paperclip, and it was fatal. The solution to the problem requires that a single animal stand inside two overlapping pens at once, and a zebra fastened to a wire is in exactly one place forever. Those three could work for the full ten minutes, in good faith, with correct instincts, and never arrive — because the object in their hand had quietly ruled the answer out. The authors had piloted the material and not foreseen it. They removed the three from the analysis and said so.
I keep coming back to that footnote, because it is a cleaner statement of the paper's thesis than the thesis is.
The problem, and the zero
The task is the 17 animals problem, adapted from Metcalfe and Wiebe (1987): put seventeen animals into four pens so that each pen holds an odd number. It arrives dressed as arithmetic, and as arithmetic it is impossible — four odd numbers cannot sum to seventeen. It becomes possible the moment the pens stop being containers and become sets, free to overlap, so that an animal in the intersection is counted twice.
Vallée-Tourangeau, Steffensen, Vallée-Tourangeau and Sirota ran it two ways in Acta Psychologica 170 in 2016. Everyone got three minutes with pen and paper first, and everyone failed — no participant in either group sketched an overlapping pen in that window, and every sketch showed them reading it as arithmetic. Then a twenty-five-minute working-memory battery as an interval, and then ten more minutes in one of two conditions. Half were given a stylus and an electronic tablet to draw on. Half were given the pipe cleaners and the zebras and told to build the solution, with no pen and no paper anywhere.
Of the twenty-four participants who had the stylus, none solved it. Not one, in ten minutes, after already knowing the arithmetic was going nowhere. Two of them drew something resembling overlapping sets in passing and then went back to dividing seventeen.
Of the twenty-three left in the model condition, ten produced full or partial overlapping-set solutions. Six had it outright. Barnard's exact test puts that at p < .001, which is the sort of number you get when one cell is empty.
The second experiment replaced the pipe cleaners with hoops and the clip-on zebras with figurines that could not clip, and showed the tablet group a photograph of the hoops first so the two groups started from the same picture. Solution rates: four of twenty-three with the stylus, thirteen of twenty-four with the hoops. Seventeen per cent against fifty-four. The effect shrank and held.
What did not predict anything
The part I find hardest to argue away is the null.
Both groups were profiled on operation span and symmetry span, on actively open-minded thinking, on need for cognition. No group differences, which is what randomisation buys you. But also no difference between the solvers and the non-solvers inside the model condition. Whatever separated the six who got it from the thirteen who did not, it was not working memory. The largest gap on any of those measures failed to reach significance.
Experiment 2 is more complicated and the authors do not hide it. There, independent measures of creativity did correlate with success — fluency on the alternative uses task in both conditions, and the creative-behaviours inventory in the model condition. So something about the person registers. But the variable that carried the solve rate from zero to forty-three per cent was sitting on the table, and it had been issued by the experimenter.
Done, not had
There is a chapter out this year called How Insight Is Done: A Cognitive Ethnography of Problem-Solving in Action, by Frédéric Vallée-Tourangeau, in the Palgrave volume The Praxis of Insight, pages 91 to 128. I should say plainly that Springer's paywall stopped me at the title page — I have the citation and I have a decade of the same author's published work, and I have not read the chapter. What I can vouch for is the title, and the title is the argument. Insight done, in the sense that laundry is done and a repair is done: accomplished somewhere, with objects, over a stretch of recorded time, by somebody whose hands were busy.
That framing collides usefully with the standard picture. The classical accounts put restructuring in the head and treat the table as a display surface. On this reading the sequence runs the other way: the hand rebuilds a pen, the perceptual configuration changes, the change affords an action nobody premeditated, and the reinterpretation is the residue of that loop rather than its cause.
Which is why the zebras matter. If insight were happening behind the eyes, a paperclip's fastening behaviour would be an irrelevance — a fact about stationery, invisible to the psychology. Instead it functioned as a wall, and three people spent ten minutes with their backs to it.
The record and the reasoning surface
Yesterday I wrote about Nasvytis and Fan's think-aloud study, where the moment before a solve stayed stubbornly inarticulate and what arrived afterwards was a name for the trick. Set the two beside each other and the awkward result gets less awkward. If the restructuring is being carried out by the hands against the material, there may be very little in the head to narrate at that instant, because the work is not happening where the narration is looking. The naming comes after because the naming is a second act — the person catching up to what the table already did.
I want to be careful. These are different labs and different problems, and the chapter that gave this post its framing is one I have only the citation for. But both are studies whose real instrument is a recording, and both find the interesting thing outside the participant's report of themselves.
The design consequence is uncomfortable for anyone who thinks of props as dressing. On this account the physical objects in an escape room are the reasoning surface itself, which makes their incidental properties load-bearing — the weight, the stiffness, the way a thing fastens. A prop that cannot be lifted, a card stock too stiff to fold, a mechanism bolted where it looks moveable — each of these is a stylus condition wearing a set piece. And somewhere in every room there is a zebra that clips on: an affordance nobody designed, nobody piloted for, that silently forecloses the path while the team keeps working in good faith and blames themselves.
Playtesting catches the puzzles. What catches the affordances?