Picture a hand hovering over a matchstick, and the video of that hand coded in ELAN, frame by frame, so the hover carries a start time and an end time and sits in a timeline beside every other movement the person made. Frédéric Vallée-Tourangeau's argument is that if you want to know where a solution came from, that timeline is the evidence, and what it shows is a solution arriving on the table rather than in the head.

His book The Praxis of Insight: Interactivity and Prototyping in Creative Problem Solving came out from Palgrave on 14 March — 140 pages, in the Studies in Creativity and Culture series. The fifth chapter, "How Insight Is Done: A Cognitive Ethnography of Problem-Solving in Action," is the empirical one: participants working matchstick arithmetic, a triangle of coins, and anagrams with real objects in their hands, every interaction timestamped. The publisher's summary of it makes a strong claim, that solutions emerge from physical restructuring producing new perceptual affordances rather than from mental restructuring occurring inside individual minds.

I should say at the outset that the book is paywalled and I am reading its abstracts, not its pages. What follows concerns what has happened to the quantitative version of the same claim while the qualitative version was being written, rather than the evidence in those chapters, which I have not seen.

The bet, and where it came from

The materials are old and clean. Knoblich, Ohlsson, Haider and Rhenius introduced matchstick algebra in 1999 — false equations built from Roman numerals and arithmetic operators, made true by moving exactly one stick. Their theory named two separate sources of difficulty, and the names have held up for a quarter century. Chunk decomposition is the operation of breaking a perceptual unit apart: seeing that the V in a numeral is two sticks and can be dismantled. Constraint relaxation is the operation of dropping an assumption about what may legally be altered: realising that the equals sign, or the operator, was never protected.

The interactivity claim, in its simplest form, is that people do better on these when they can pick the sticks up. The theoretical machinery behind it — the Systemic Thinking Model, distributed cognition, the whole ecological turn — holds that information processing changes shape when it is spread across mental and material resources. Which is a beautiful idea, and one I have been leaning on without saying so, because it is the load-bearing assumption under everything I have written about hands and rooms and handwriting.

Three results in the last five years have taken runs at it.

Three failures to move the number

Adam Chuderski, Jan Jastrzębski and Hanna Kucwaj ran the largest of them for the British Journal of Psychology in 2021, under the title How physical interaction with insight problems affects solution rates, hint use, and cognitive load. Two hundred and forty-eight young adults, half of them able to physically manipulate the problem elements and half working on paper, across nine established insight problems. They found no general facilitating effect. One problem out of nine clearly benefited. Task load, working memory correlations, and the subjective sense of having had an insight came out substantively the same in both formats.

The retreat position after Chuderski was that manipulation helps spatial insight problems specifically — the ones where the answer is an arrangement, like the eight coins. In May a paper in the Journal of Intelligence went at that too. Does Physical Interaction with Insight Problems Really Affect the Solution Rate? put a new spatial problem, the Pencil problem, to two groups in paper and physical-interactive conditions. No significant difference in success rates. The authors' reading is that manipulating the materials does not appear to facilitate spatial insight problems as such, which removes the last protected category.

And in January, Vladimir Spiridonov, Maria Erofeeva, Nils Klowait, Maxim Morozov and Vladlen Ardislamov published The modulating role of sources of difficulty in interactive matchstick algebra in Frontiers in Psychology — three experiments, 144 participants in final analysis, on the original matchstick materials rather than a substitute. They designed four problem types to cross the two difficulty sources deliberately, varying how tightly the chunks were bound against which constraint had to be dropped — value, operator, tautology. They found no significant main effect for the interactive conditions, with Bayesian evidence running weak to moderate in favour of the null.

That is a normal replication story, and on its own it would be a footnote. The interesting part is the analysis they ran next.

The result that runs backward

Spiridonov and colleagues did not stop at the group means. They looked at how much each solver actually moved, and asked what that motor activity predicted.

For the chunk-decomposition problems, nothing. Movement was unrelated to success. For the two problems requiring the higher-order constraints — the operator constraint and the tautology constraint — movement predicted success negatively. Their conclusion is that solvers' movements do not affect the decomposition of chunks but "significantly interfere with overcoming two higher-level types of constraints."

Sit with which operation that is. Constraint relaxation is dropping an assumption you did not know you were holding. It is, as nearly as the vocabularies map onto each other, the laboratory name for un-fixation — the operation I have called the load-bearing solving skill across the anatomical-shadow, sleep-cheap-route and cryptodiagnosis lines. The mapping is mine and it is not exact; a matchstick constraint is a small, well-specified assumption, and a cipher solver's fixation is a sprawling one. But if the correspondence holds even loosely, the finding says the hands are neutral on searching and costly on the part where the click lives.

It also sharpens something in Chuderski that I glossed on first reading. In that study, the hints worked less well for participants who could touch the materials. The 2026-07-31 post argued that hint design aims at a computation that is not running — alternative-naming hints presuming a comparison the solver never started. This is a second, independent problem with hints, and it is worse, because it says the delivery channel degrades in exactly the condition escape rooms operate in. Hints land on people whose hands are full.

What this costs the room

The escape room is the maximally interactive puzzle format. That is the entire commercial proposition. Every prop is manipulable, every solver is standing up, and the industry's stated advantage over a puzzle book is that you are in it, touching things. Nobody in the field has ever needed to defend the assumption that the hands help, because the format is unimaginable without them.

Three studies now say that hands do not raise the solve rate on classical insight problems, and one says hand movement specifically taxes the operation that produces the insight. If those transfer — and I want to be careful about how large a gap that is — then the room's central affordance is spending in the currency the click needs.

The gap is genuinely large. A matchstick algebra problem is a sixty-second desk task with a single legal move. An escape room runs an hour, distributes across a team, mixes search with transcription with lock manipulation, and only occasionally contains a constraint-relaxation moment at all. Much of what happens in a room is chunk decomposition and plain search, and on those the movement data says the hands cost nothing. A room could be almost entirely composed of the operations interactivity does not damage, and the good ones probably are.

But the moment a room stakes its whole experience on one designed epiphany — the moment somebody has to realise the map is the cipher, or the equals sign was never protected — that is a constraint-relaxation moment, staged in the most hands-busy environment anyone has built.

Which instrument was wrong

Here is where I think the dispute is actually a category error, and where my own framework has an opinion.

The replication studies measure solution rate. That is a measure of the destination — did the plaintext come out, was the equation made true, count the successes and compare the proportions. The cognitive ethnography measures something else entirely: the sequence of interactions by which a person arrived somewhere, coded to the frame, which is a claim about the journey.

In the Kryptos K4 piece I argued that solving is a claim about the journey and not the destination — that a plaintext recovered from an archive without a method is not a solution, and that everyone involved agreed on this without needing persuasion. If I meant that, I owe it here. A solution rate cannot adjudicate a claim about how solving is done. It can only report how often it finished.

Which cuts both ways, and I would rather be precise about the direction of each cut.

Against the interactivity tradition: a claim about the journey does not get to imply a benefit at the destination and then decline to be measured there. If the ecological account has been read as promising better performance, the three studies have collected on that promise and the account has to stop making it.

Against the replications: showing that a rate does not move is not the same as showing that nothing happened. Two people can reach the same equation, one of them having built the answer on the table and the other in the head, and the rate is deaf to the difference between them. Wherever the difference matters — for what gets remembered, for what gets described afterward, for whether the solver can do it again — the rate was the wrong thing to count.

And that is precisely the register the escape room sells in. Nobody books a room to maximise the probability that the equation is made true. The industry's own success metric is escape rate, which it has always known is a poor proxy for whether the room was good, and it knows this because rooms with high escape rates get reviewed badly all the time. The format's product is the journey. Which means the literature that finds against interactivity is, mostly, measuring something escape rooms are not selling.

The uncomfortable residue is that the format selling the journey may also be the format most reliably taxing the journey's best moment.

Levers, held loosely

Three things a designer could actually reach for, and I hold the first more firmly than the third.

Give the constraint-relaxation beat a place to happen with the hands still. Not a phase-scheduled clock, which is a different argument, but a physical one: a chair, a window, a wall of text that has to be read rather than handled. The rooms with a moment where everyone stops moving and looks at one thing may have found this by feel.

Do not deliver a hint into busy hands. If Chuderski's result transfers at all, the hint that arrives while a team is prying at a drawer is being spent at a discount. A hint that first asks everyone to put things down costs ten seconds and might buy back the whole intervention.

Watch where the movement is when the room stalls. The Spiridonov design measured motor activity as a continuous predictor, and escape rooms already run cameras. A room that stalls in a burst of handling is a different failure from a room that stalls in stillness, and the craft vocabulary I have found does not yet distinguish them.

What I cannot tell from any of this is whether the interference is about the hands specifically or about attention having somewhere to go. A solver moving matchsticks is doing something, and doing something is famously how you avoid the uncomfortable stillness in which an assumption might get dropped. If that is the mechanism, the room's failing is that it never gives anyone a reason to put the objects down.