
A mouse puts its nose into a port and smells an odour. The odour tells it, truthfully, whether the next few seconds will contain water. The mouse cannot act on this. The trial is already settled, and the odds sit at 25% whichever way it turns. It smells the odour anyway, and when the experimenters begin charging for the privilege, shaving water off the informative side until the preference finally breaks, the mouse keeps paying up to a price of roughly four to six microlitres.
That number is what I keep returning to. Somebody has put a price tag on wanting to know.
The paper is Representations of the intrinsic value of information in mouse orbitofrontal cortex, out in Nature Neuroscience, by Jennifer Bussell, Ryan Badman, Christian Márton, Ethan Bromberg-Martin, L.F. Abbott, Kanaka Rajan and Richard Axel, working across Columbia's Zuckerman Institute, Harvard Medical School and Washington University. The preprint is open, figures and all.
The mouse buys a spoiler
The task is a three-port nosepoke. A centre odour sends the animal left or right. One side port issues cues that reliably announce the coming water; the other issues cues that announce nothing. Reward probability is 25% at both, and the authors are careful about why that matters: the information "does not alter the structure of the task or the reward outcome." Nothing is purchased with it. Nothing is avoided. The mouse simply learns the ending a few seconds early.
Mice chose the informative port on 78% of free-choice trials, with a 95% confidence interval of 71 to 83%.
Sit with what that cue actually is. Advance, accurate, unactionable knowledge of how a thing turns out. That is a spoiler, in the precise sense the word carries in every fandom that polices it.
And puzzle culture is organised in flat opposition to it. The whole apparatus of hint systems, sealed envelopes, progressive clue tiers, the game master who watches you flounder for six minutes before saying anything at all, exists to ration exactly the commodity this mouse is buying at a premium. A first-time solver who is handed the combination has been robbed, and will say so. Enthusiast forums maintain elaborate spoiler etiquette for rooms that thousands of people will never visit. We built an entire economy around withholding the odour.
So either the mouse and the solver want different things, or they want the same thing and disagree violently about the delivery schedule. I lean toward the second, and I think the paper's neural half is where the argument gets interesting.
Two ledgers, barely touching
The team imaged 1,138 orbitofrontal neurons across seven mice with a microendoscope, tracking the same cells across weeks of training.
Roughly 19% of OFC neurons responded to the cues predicting information. Roughly 21% responded to the cues predicting water. Those two populations overlapped by only 23%, which the chi-square test could not distinguish from chance, and the coding indices for the two cue types showed no relationship across the population at all (r = 0.04). When they trained a classifier on one, it failed to read the other. The authors' phrasing is that information and water value "are encoded by largely separate ensembles of OFC cells."
Here is the tension I find genuinely hard to put down. The behaviour proves the two things are convertible — there is an exchange rate, and it is four to six microlitres. Convertibility usually implies a shared unit somewhere, a common scale on which the comparison gets made. The orbitofrontal cortex declines to provide one. It keeps two sets of books that barely reference each other, and the trade happens anyway.
The same author, seventeen years earlier
What makes this properly strange is that one of the co-authors already answered this question, and got the other result.
In 2009, Ethan Bromberg-Martin and Okihide Hikosaka published Midbrain dopamine neurons signal preference for advance information about upcoming rewards in Neuron. Monkeys chose between a cue that revealed the size of a coming reward and one that revealed nothing. As in the mouse task, the choice "had no effect on the reward size." The monkeys took the informative option on 80 to 100% of trials.
Then the recordings: the same midbrain dopamine neurons that signalled expected water also signalled expected information, bursting or falling in the same grammar, and the strength of that neural discrimination tracked how badly each individual monkey wanted to know.
Same investigator. One system merging the two currencies into a single signal, another keeping them in separate ledgers. This does not read to me as a contradiction so much as an architecture with a conversion desk in it — the midbrain producing a common-currency teaching signal, the cortex maintaining the distinction that the teaching signal collapses. Which layer you record from determines whether curiosity looks like a flavour of appetite or like its own appetite entirely.
I wrote in the 7 May post about a finding that people will freely take on Stroop and Simon tasks with no external reward, and argued it relocated puzzle satisfaction somewhere the click-and-resolution framework doesn't reach. This lands in the same territory from the other side. There, the effort paid for itself. Here, the answer does, at a rate you can measure in microlitres.
What a hint costs
If the accounts really are kept separate, a design consequence follows that I don't think the industry prices correctly.
Escape rooms hand out two distinct goods. One is the room's information — the combination, the mapping, the thing the cipher says. The other is the win: the door, the clock, the photo, the leaderboard. Standard practice treats the hint as a tax on the win, which is why hints so often come bundled with a penalty and delivered in a tone of mild disappointment.
But if the information carries value on its own books, a hint does two things at once. It debits the win, and it deposits into an account the escape metric never reads. A team that fails the room and learns how it worked has been paid in a currency the leaderboard cannot see. A team that escapes on brute-forced locks without ever understanding the mechanism has won while leaving that account empty, which might explain a review pattern I find otherwise puzzling: groups who got out, said so proudly, and rated the room poorly anyway.
The measurement problem is obvious enough. Nobody has run the mouse's experiment on a solver. What would a human pay, in something more honest than self-report, to learn the answer to a puzzle they have already failed and will never revisit? Not the answer they can still use. The one that arrives too late to matter.
I would like to know that number. I suspect it is larger than the hint-penalty schedules assume, and I notice, with some amusement, that wanting to know it is itself the thing being measured.