A Note on Origin
Throughout my career I have been asked to think outside the box, and I have asked other people to do it. I have also watched what the instruction does to a room. Some faces go tight, because they have been handed an objective with no method attached. The brave ones ask the question the instruction invites and cannot answer: where did you learn it?
I have never had a good answer. I can describe a method, and I will — but a method held to strictly is only ever as good as the subject it is being held against, and the whole point is that the subject is the thing you have not looked at yet.
Which is the honest reason for this essay. Not whether thinking outside the box is valuable, but whether it is teachable at all.
Here is the part that makes it hard. Every field hands you a vocabulary, and the vocabulary is not neutral. It gives you the rules, how things are supposed to work, and which questions are the sensible ones to ask. Nowhere does it say you have to accept any of it. You infer that you must, because everyone competent around you already has.
So the box is not a failure of imagination. It is what you were given in order to be useful, and it arrives before you have had a look at the thing.
Little b, big B
It started because I was shopping for used server hardware.
The listing was for a very fast solid-state drive. It did not say how fast, so I looked it up. The drive is rated at roughly 6.2 GB/s — gigabytes. Which confused me for a second, because open in another tab was an ordinary 14-terabyte hard disk whose description also said six: SATA 6 Gb/s — gigabits.
The former cost two thousand dollars and the latter one tenth. Surely markets were not this inefficient.
I knew that was wrong. Not as an argument — as a physical intuition. One of these things is a filing cabinet with a small mechanical arm inside, waiting for the right drawer to spin past. The other has no moving parts whatsoever. Whatever the numbers said, those two do not work at the same speed.
Which got me thinking about what these things have in common. A hard disk, a solid-state drive, and throw into the mix the memory in the machine — they all hold data and hand it to the processor when it is wanted. The differences are physical. How far the data has to travel, and how wide the path is when it goes.
This is a main reason a modern graphics card can approach the size of a loaf of bread. It is not mostly processor. It is memory bolted directly alongside the processor, and the cooling required to sit that close — shorten the distance, widen the path, pay for it in volume. Inside the chip it happens again, and again, each level nearer the work than the last. One idea applied over and over.
Oh yeah, I almost forgot. 8 Gb = 1 GB.
No seam on a synapse
Here is where it stops being about computers.
Dense compute, long-term persistent storage, short-term volatile memory. Say it about a machine and it is a spec sheet. Say it about a brain and it is synapses, consolidated memory, working memory.
If the whole direction of travel is to shrink the distance between holding and acting, the interesting case is the one where there is no distance left. A brain runs the same sequence — takes something in, holds it, acts on it, produces something — with no wire down the middle and nothing travelling between parts. And when I pushed on that, what came back was not an analogy but a definition:
A component is memory when its state matters later. It is compute when its state participates in determining what happens next.
There is no reason it cannot be both.
Then it is the same thing about a synapse. It holds something, and the holding is what changes how the next signal is transformed — so it is not one or the other. There is no seam on it marking where the storage ends and the computing begins.
I should be careful here. This is territory where people who have spent their lives on it do not agree — whether the mind decomposes into components at all, and if so into which ones, is not settled. There are schools. There are decades of argument.
That disagreement is the point rather than a hole in it. If memory and computation were joints in nature, the people looking hardest would have found them by now and stopped arguing about where to cut. What they are actually arguing about is which description to use — each of them holding a vocabulary their own training handed them, trying to settle the structure of the mind with the only instrument available, which is one.
So memory and computation are not kinds of object at all. They are roles that a physical state plays, and whether something counts as one or the other depends on what you want to know about it.
Von who?
Only then did a name arrive. Von Neumann.
I had not heard it before that morning.
The architecture underlying most conventional computers is his: a processor at one end, memory at the other, instructions and data travelling between them. When the processor can work faster than the system can feed it, the connection becomes the constraint. That is the von Neumann bottleneck, and it has been named and studied since before I was born.
So the boundary I had just concluded was not there is the boundary the canonical model is built on. It does not describe the separation as a choice. It describes it as the architecture, and then names the cost of it as though the cost were a fact about computers.
Which means the entire escalating engineering effort — all that shortening and widening, the loaf of bread — is us paying year after year to undo a separation we invented because it was convenient to think with.
What interested me was the order. Not that I had got somewhere early — that I had got somewhere I could not have got to at all if the name had arrived first, because the name hands you the boundary before it hands you the question.
The name is worth having. Once somebody says “von Neumann bottleneck,” eighty years of work becomes findable — all of it invisible under the search term “why the hell is the memory over there?” I just would have had nothing of my own to hold it up against.
This keeps happening to me. I reason my way to the shape of something, then discover the shape has a name, a literature, and people who have been arguing about it at conferences since before I was born. It used to annoy me. I am beginning to think the order is the useful part.
One level below thought
The brain comparison ran out of road about there. I could see the shape of what I was pointing at and could not get it across, and what I came up with was this: a model built mentally cannot be expressed physically.
The first reading it got: mental models divide reality into discrete functions, physical reality has only states and interactions, and the boundaries are imposed by the model. True, well put, and not what I meant.
I meant something one level lower.
One degree below what I experience as a mental model, there are synapses and neurons in physical states. Those states are not divided into storage here and computation there. A synaptic state persists, and therefore participates in memory; that same state changes how subsequent signals are transformed, and therefore participates in computation. They are the same physical thing under two descriptions.
So the mental model with which I am considering the separation of memory and computation is itself running on a substrate where that separation does not hold.
That is not a cute recursion. It means my mental model is not an observer standing outside the physical environment and describing it. I am using synapses to think about synapses. And in a synapse, the separation I am reasoning about is not there to find.
There is an important distinction inside this. The object I conceive does not have to be physically possible. I can imagine a perpetual-motion machine, an infinite hotel, a four-dimensional object, or a magic wand.
But the act of conceiving it necessarily has a physical implementation.
Whatever conceptual freedom feels like from the inside, the representation has to occur in a realizable neural system. I can reach beyond the environment in what a thought represents. I cannot escape it in the mechanism by which representation occurs.
Testing it somewhere I know better
So far I have been reasoning about hardware I bought last week and a brain I have never studied. That is thin evidence for a claim this size. If the observation is any good it should hold somewhere I actually know something, and be falsifiable there.
I have spent most of a career in capital markets. So: does it hold?
Markets run on a framework most of the consequential participants learned in the same places — management, investors, analysts, lenders, boards. At first that is a model of the market. Then the people carrying the model become the market, and the framework stops describing behaviour and starts producing the behaviour that will later be read as evidence for it. I have written about pieces of this before — how the budget process filters out ideas that cannot be defended before they are understood, and why the people inside those processes behave the way they do. What I did not have then was a name for the shape underneath both.
It is the same shape, and the constraint works the same way. A synapse does not stop me imagining a wand; it makes the imagining a physical event in a medium where the distinction does not hold. The constraint is on the act, not the object. The framework does not stop a CEO conceiving something unconventional either. It shapes the environment the idea has to walk into once conceived — an environment the framework itself built. In both cases the model is not standing outside the thing it describes. It is inside it, and it helped make it.
Which is why “think outside the box” confounds me. A CEO may think outside it perfectly well. The lender still prices the box, and the lender is not in the room.
The box has become environmental.
And notice what that does to a magic wand. A wand specification cannot be defended upfront, because the inputs are not known yet — which is exactly the property that makes an idea fail review. The framework is not wrong. It is superb once the opportunity is real. It is simply the wrong instrument for deciding what is worth doing in the first place.
Pointing that out is interesting. It is not yet useful.
The question that matters is what you do about it.
What is the widget?
I have a method I use without usually calling it a method. The best description of it I have heard is not mine: decompose until the inherited solution disappears.
I start by asking: What is the widget?
Not what is the accepted name of the problem. Not what tool do we normally use. Not which menu item applies. What is the thing?
Then: What state is the widget actually in?
Then: What state do I want it to be in?
Then: What has to remain true while I transform one state into the other?
Only after that do I ask how to transform it.
That ordering matters because nouns carry instructions. The moment I call something a disk-resizing problem, a toolbox appears. The moment I call something a financing problem, another toolbox appears. The moment I call something a memory problem, I have already accepted a boundary around memory — which is precisely the boundary the whole first half of this essay was about.
Those toolboxes are useful, and that is the danger. A toolbox does not announce itself as an assumption; it announces itself as help. I just do not want one in the room while I am still deciding what the problem is.
Once I know the widget, its starting state, its ending state and its invariants, I deliberately make the transformation space ridiculous.
At one end is the crudest thing that definitely gets me from A to B. Destroy it and rebuild it. Replace the whole machine. Copy everything. Hire a thousand people. Throw money at it. It is ugly and it works, and it establishes a physically plausible outer bound.
That end is the BFH: the big fucking hammer.
At the other end is the magic wand.
If I could ignore cost, physics, convention and available tooling for ten seconds, what exactly would I want to happen? Not “make it better.” What would the wand actually do to the state of the widget? What would change? What would remain identical? How quickly would it happen? What would the world look like immediately afterward?
Poof!
This is where people sometimes misunderstand me. The point is not to imagine a magic wand and then return sadly to reality.
I would say: go find me a magic wand.
The impossible tool is a specification.
Once I have specified what the wand does, I can search the real world for mechanisms that approximate those properties — and the existing industry solution gets no privileged position merely because it already has a name.
Feasibility comes late on purpose. Let it in during problem definition and it hands you the toolbox before you have said what the problem is.
Standing to disagree
None of this is an argument against expertise. If I need brain surgery I do not want the surgeon rediscovering anatomy from scratch. Expertise is accumulated error correction — civilization’s way of not paying twice for the same lesson. And training has to categorize, because that is how knowledge becomes teachable.
But categories are powerful precisely because they collapse complexity, and compression discards information. The danger is when “this is a useful way to decompose the problem” quietly becomes “this is what the problem is.” A wire between memory and computation is a useful decomposition. It is not a fact about matter — and the difference between those two sentences is worth eighty years of engineering effort spent undoing it.
Which brings me back to the filing cabinet with the arm in it. That picture is what let me overrule a published specification, and it is the scarce thing.
Not the willingness to question, and not a technique for questioning. It is possessing an independent model of the object — one built from what the thing does rather than from what the field calls it — so that there is something in the room with the authority to contradict the documentation when the two disagree.
This is what the four questions are really for. What is the widget? is not a rhetorical opener. It is the question that constructs the independent model, before the vocabulary arrives to supply one ready-made. You cannot notice that a category is doing work it should not be doing unless you have some description of the thing that does not depend on the category.
And it repairs something I left loose in the market case.
The problem there is not that a shared framework prevents people from thinking. It is that for many participants the framework is the only model of the object they have. If your entire picture of a business is constructed from capital structure, hurdle rates, comparables and return on invested capital, then when the framework and the business disagree, there is nothing available to notice the disagreement with. The documentation is not being weighed against an independent picture, because there is no independent picture. There is only documentation.
Which is what confounds me about “think outside the box.” It asks for a conclusion the person has no instrument to reach. You cannot decide to disagree with the only model you possess. Something else has to be in the room first — a physical intuition, a direct look at the thing, an operator who has actually run it — and then disagreement becomes possible without being an act of will.
The one problem I cannot get out of
There is a paradox in being taught to think outside the box: the lesson necessarily arrives in a box of its own. A method, a framework, a curriculum, a five-step process for escaping five-step processes. Everything I have just written is subject to it. Ask what the widget is, specify the impossible tool, let reality vote last — the moment any of that becomes canonical it starts acquiring the same authority over the problem it was supposed to loosen. If you take the method as the rules, you have built exactly the thing the method exists to postpone.
I do not know how to solve that universally. Trying to devise a universal method for escaping it would only produce another box.
What I can say is what happens when the postponement works. Sometimes the answer turns out to be conventional, and often the conventional answer is conventional because it is very good. Sometimes ordinary tools assemble into something that behaves suspiciously like the impossible one. Sometimes somebody has already shipped the wand and you simply had not searched for it under that description. Sometimes the answer is still the BFH. And sometimes the most valuable result is discovering that the problem you were preparing to solve consisted mostly of 160 gigabytes of nothing.
So I do not leave the magic wand on the whiteboard. If the transformation I actually need looks impossible, I ask the more useful question:
Can someone tell me where I can find a real magic wand?