There is a puzzle you have almost certainly seen. Nine dots in a three-by-three grid, and the task is to connect them all with four straight lines without lifting your pen. Most people fail, because they draw as if the dots form a box and the lines must stay inside it. Nobody told them that. They imposed the boundary themselves, and the solution requires letting the lines run past the imagined edge — the literal origin of "thinking outside the box." The nine-dot problem has been a workhorse of insight research for the better part of a century.

I have been reading about a framework that made me look at that workhorse differently — not at the puzzle, but at what we have been able to see while people solve it. Because the uncomfortable answer is: almost nothing.

What the classic experiments actually record

The canonical laboratory tools for studying insight are a small stable of set-piece problems — the nine-dot, the Remote Associates Test (three words, find the fourth that links them), a handful of matchstick arithmetic and riddle problems. A participant sits with one, works, and either solves it or does not. Sometimes they narrate as they go; more often they report afterward what it felt like — the blankness, the stall, the sudden click.

From that thin record, the field built a robust and genuinely useful theory. Insight, the account goes, requires an impasse — a point where every move your current framing permits has been exhausted — followed by a restructuring, where the problem's representation is rebuilt so that a previously invisible move becomes obvious. Stellan Ohlsson's representational change theory is the elegant version of this, and it has held up well enough that I have leaned on it here more than once.

But look at what the instrument captures. A start state. An end state. A stopwatch. And a story. The restructuring itself — the actual moment the representation changes shape — is never observed. It is inferred, backward, from the fact that a solution appeared and the solver says it arrived suddenly. We have been studying the most interesting event in problem-solving by studying its shadow on either wall.

The move: change what the instrument records

The framework that reframed this for me is ESCAPE, from Vasanth Sarathy, Nicholas Rabb, Daniel Kasenberg, and Matthias Scheutz at Tufts, presented at the 2024 Cognitive Science Society meeting. Their diagnosis is blunt: the literature "lacks a cognitive process model of insight," and the reason is partly methodological — there has been no "unified, scalable, and tunable experimental framework" for watching creative problem-solving with any fidelity.

So they built one out of puzzle video games. Not the toy problems, but designed game environments where a solver makes move after move, and every one of those moves is logged. The result is what they call process data — not just whether you solved it and how long you took, but the entire trajectory of what you tried, in what order, and where you doubled back. The pitch is that this lets researchers "explore different theoretical frameworks of representation restructuring" directly, against a record dense enough to model computationally, rather than reasoning from a start point, an end point, and a solver's after-the-fact impression.

What I find quietly radical is the shift in altitude. The old instrument sees the answer. This one tries to see the search — the shape of the wandering before the answer, which is exactly the region where restructuring, if it is a real event, would have to leave a trace.

The wrong-instrument pattern, one more time

I keep noticing a shape across fields I read about. A discipline names a construct — creativity, intelligence, a nematode's behavior, a hacker's skill — then measures it with the most tractable instrument available, and for years cannot tell that the name and the measurement have quietly come apart. Divergent-thinking tests were supposed to measure creativity and may have been measuring something narrower. A complete wiring diagram of C. elegans has not yielded a working simulation, which suggests the connectome may be the wrong layer to have mapped. The pattern is always the same: the instrument was never neutral. It decided in advance what could count as evidence.

Insight research has its own version, and ESCAPE names it. If the only thing you can record is the before and after, then "impasse followed by restructuring" isn't just your best theory — it may be the only theory the instrument was ever capable of producing. You cannot find a gradual, many-step restructuring in data that has no middle.

But I want to be honest about the trade, because it is not free. The one thing the old, thin instrument captured that the rich one may not is the phenomenology — the report. The felt click, the sense of sudden arrival, is precisely the thing a move-by-move log cannot see. Process data can show you a solver's trajectory bend toward the answer; it cannot tell you whether, in that moment, the solver felt the floor drop away. You gain the search and you risk losing the story. And the story is not decoration. For a puzzle designer, the click is the product. A framework that finally lets us watch restructuring happen might, if we are not careful, teach us everything about the path and nothing about why arriving on it feels like grace.

So here is the question I am sitting with. If we could finally watch a representation change shape in real time — every intermediate step laid out like a transcript — would we discover that the sudden click was an illusion all along, a smooth process our awareness only sampled at the end? Or would we find the log going along steadily, steadily, and then a single discontinuous jump that no amount of resolution can dissolve — the instrument at last catching the thing it was built to deny?