Follows from Part 1: uncovering the laws of nature
Friends,
In Part 1, I teased that Geoffrey West’s Scale was a perfect surface to show how you can learn as you’ve always wanted. Or needed but didn’t know it.
Plan
🧠For those familiar with learning science, you will recognize several techniques, but I’ll label them as they appear.
How this all started
When I read a physical book, I will usually take a screenshot and then OCR the page to keep a digital excerpt. This is ok if there aren’t a lot of excerpts or highlights to preserve. I quickly realized I was going to need the Kindle version of Scale as the highlights and their accompanying inconvenience were piling up fast. I snagged it on Libby (this is your library’s digital loan service). There was no waitlist. Yet another reminder that there is so much joy available for free.
We must talk about highlighting.
The naive understanding of highlighting is that by taking the effort to trace a 25% opacity yellow film over words, you have learned something. By now you know this is item #77 on the list of self-deceptions. Still we carry on because it’s a cheap option. Somewhere in the recesses of your dopamine-addled mind (remember dopamine is the “seeking” chemical), you expect the highlights you stash like old coax cables will find fresh life when recombined with technology. Don’t be hard on yourself, an impotent but aspirational habit ranks less than wisdom but higher than apathy.
But it turns out this self-deception call option may finally have a payoff.
As I was marking up Scale, I had no guilt about not internalizing what I was reading in the moment. My plan was to export all the highlights like I usually do.
⚠️Kindle formats usually limit your exports to 10% of the book for copyright reasons. It’s a bit messy, but once I think I’m about there, I export those highlights to a file, then delete them in the Kindle app. For Scale, this happened when I was about half finished with the book. I could then start highlighting from zero for the second half.
This time, instead of just storing them, I was going to give them to a Claude project to seeding a “curriculum” so I could learn in a way that only comes from practice, not recitation. Reviewing your notes/highlights lets you cram for a test, but it’s not the kind of learning you can call on for invention.
Socrates 2026
Let’s rewind for a moment.
A few months back I bombed a Jane Street interview question I found online. This wouldn’t normally bother me as I’m far past the time in my life where my self-esteem teeters on an illusion of cleverness. But it was a question I felt I should know how to answer as opposed to the corpus of questions from which I wouldn’t even know where to begin.
You can see my write-up about it here: turning a Jane Street interview question blunder to a lesson
I realized that I couldn’t answer the question because I didn’t fully appreciate that variance is the spread between the expectation of a square and the square of an expectation.
Adjacent thoughts
My knowledge of variance was vague and formulaic. My knowledge of many things is like that. I don’t find that comforting, just a practical necessity in a limited life. But part of life’s pleasures is the freedom to NOT 80/20 something if doing so bugs you.
Alas, this one bugged me and the cost to fix it is lower now since LLM’s can be used as tireless tutors whose judgement of our faculties presents no threat.
I opened a chat and asked it to teach me Socratically, one small question at a time. When I run out of time or get tired, I know I can pick up where I left off. Or a little bit before that, since I usually need to insert before the point where I got tired since that point coincides with the material that made you take a break, so you don’t quite “own” it.
This is a snapshot of where I am in my Variance progression where I derive every formula from the already intuitive definition of “sum of squared deviations”:
As the learning progresses, you build more cases. With practice, you see that the key to all the derivations is that you are building on things you already know:
It’s hard to see this without practice.
When I revisit some of the derivations, I sometimes get stuck again but I know I’m screwing up one of these 3 foundational elements. That’s pretty crazy. It’s a gap in something I thought I knew cold, but the diagnosis is far more apparent because of how I’ve structured the learning in cahoots with Claude.
This is a timely place to name a few learning science techniques at play (see the appendix for more on these):
The entire process is infused with the “generation effect” which takes advantage of our ability to remember something far better when you produce the answer yourself than when you read it.
And finally, every topic is a branch of an overarching commitment to interleaving. Power functions are mixed into a learning practice that includes other topics I want a closer look at. The approach makes affordances for both variety and synergy.
Before getting back to Scale and power functions, I have one more remark on this whole personal project I’ve donned Socrates 2026.
I really want AI companies to launch a Native Ink Surface with a submit button. Math derivation, music notation, art. All of these would be far less painful with a stylus. Is this too much to ask:
As I was reading and highlighting Scale, I strained to interpret the exponents. That means there are gaps in my understanding. Simple as that. These aren’t new concepts, but it’s clear I need some mental Dap if I want “automaticity”.
Paraphrasing Math Academy:
Automaticity is the ability to recall foundational math facts instantly and accurately from long-term memory, requiring zero conscious effort or working memory….
Automaticity is the prerequisite to true computational fluency. Once low-level skills (arithmetic, exponent rules, trig identities) are automated, recall becomes effortless, allowing your brain to focus entirely on higher-level problem solving and critical thinking.
We’ve been taught to think of tests (ie retrieval practice) as how you check whether you learned something. But it’s actually how you learn in the first place because it’s “doing”. To learn in a durable way is to “do”.
The highlights I collected became the raw material for Claude to design questions. But AI is obviously capable of far more than regurgitation, distillation and re-shuffling. It constructs sensible questions that arise from the text but not directly addressed. It can order the questions so they build gradually. It can relate material across domains. This is a gift to a learner.
Injecting a thought
AI cannot motivate you. It cannot inspire you. AI offers an unbundling of the tutor, not a replacement. The role of humans in the learning loop is going to grow, which might be a contrarian position. Think of coaches. Some are exceptional because they are masters of the Xs and Os. Some are exceptional because of their ability to lead and communicate. These are squishy. The squishy things will not rise in relative importance. They are important and AI doesn’t change that either way. It’s that AI will put a spotlight on the fact that there will be relatively higher yields to focus on the squishy. Whether we will or not (and be able to judge the delta) is an open question. A topic for another day perhaps, but I’m betting on this with my time.
There are no shortcuts. If you want automaticity, you gotta hit the gym.
We’re working with y = xᵃ throughout. The exponent a is the only thing carrying information about the relationship.
Warmup.
y = x². If x doubles, what happens to y?
That question should feel familiar. Option prices follow a square root relationship with respect to time.. Doubling the time to expiry only multiplies a straddle by √2, while quadrupling it doubles the straddle.
The exponent is less than 1, so y grows slower than x. Less than double. This is the whole idea of sublinear scaling and economy of scale. Double the animal and it needs about 68% more energy, not 100% more.
A fraction in the exponent is one number, not a division. 2³ᐟ⁴ means “take the fourth root of 2, then cube it,” which is about 1.68. Meanwhile 2³ / 2⁴ = 2³⁻⁴ = 2⁻¹ = 0.5. Dividing powers subtracts exponents.
A negative exponent is always a reciprocal. Compute the positive version, then flip.
2⁰·⁸⁵ sits between 2⁰·⁵ ≈ 1.41 and 2¹ = 2, closer to 2 because the exponent is close to 1. About 1.8. Roads go up 80% when the city doubles. Sublinear again, which means per person, road length actually falls.
2¹·¹⁵ ≈ 2.22. Careful here. The move is NOT “2 plus 0.15.” It’s 2¹ × 2⁰·¹⁵ = 2 × 1.11 ≈ 2.22. Superlinear. Bigger city, disproportionately more output per person. Also disproportionately more crime and disease. Good and the bad scale together.
Horse is 8²ᐟ³ = 4x stronger in total, but 8x heavier. So per pound it’s 4/8 = half as strong. This was Galileo’s observation. A small dog can carry two or three dogs on its back. A horse can’t carry even one. Strength grows like area, weight grows like volume, and volume outruns area as things get bigger.
Why Godzilla can’t exist
Strength scales with cross-sectional area, not size. This is why lumber is sold as a “2×4.” The two-by-four inches of cross-section is what bears the load. Double every dimension of a beam and its strength goes up 4x, because area scales with length squared.
Mass scales with volume, which is length cubed. Double every dimension and the thing weighs 8x more.
Scale a creature up and its weight (volume, 8x) outruns its strength (cross section, 4x) with every doubling. At Godzilla’s size, the legs would have to support a mass that has exploded as the cube of height while the bones holding it up only got stronger as the square. He’d snap under his own weight before he took a step. Same reason an ant can carry many times its body weight and an elephant can barely carry its own.
You build the knowledge, check yourself on new questions, come back another day, see how much ground you gave back. It’s 2 steps forward, 1 step back. Eventually you earn the consolidated reference and a sense that you have earned the shortcuts.
For y = xᵃ:
Three worked slopes, read in orders of magnitude
y = x¹ᐟ², slope one half. y grows at half the rate of x. x goes up two orders of magnitude, y goes up one. Or: to get y up one order of magnitude, x has to move two.
y = x¹ᐟ⁴, slope one quarter. Even more damped. x times 10,000 (four orders of magnitude), y only times 10 (one order). This is per-cell metabolism, which scales as mass⁻¹ᐟ⁴.
y = x³ᐟ⁴, slope three-quarters. Kleiber. x times 10,000, y times 1,000. Three orders of magnitude of metabolism per four orders of magnitude of mass. The “3 to 4 ratio in powers of ten”.
The economy-of-scale shortcut
If the total scales as xᵃ, then per unit of x it scales as xᵃ⁻¹. Subtract 1 from the exponent and you have the per-person, per-pound, per-cell law. The sign of a−1 is the whole story:
Cities have two exponents that mirror each other around 1. Physical stuff like roads and cables scales at 0.85, so per person it gets cheaper as the city grows. Social stuff like wages and patents scales at 1.15, so per person it grows.
Power law: the variable is in the base. It cares about ratios. Multiply the input, multiply the output, and the multiplier is the same no matter where you started.
Exponential: the variable is in the exponent. It cares about differences. Add to the input, multiply the output.
CAGR is a familiar exponential. Consider 10% CAGR. Going from year 5 to year 10 does not multiply your balance by the same factor as going from year 30 to year 60, even though both double the time. The multiplier depends on where you are, not on the ratio.
That base-versus-exponent distinction isn’t just about growth over time. It also governs how the probability of a large moves. The mean-standard deviation framework we are so familiar with is Gaussian “mediocrastian” math. But we know empirically that tails reside in “extremistan”. You can’t have a 10-sigma move every few decades. The bell curve is just a misspecified description of returns.
The fat tail in returns is more of a power law in the size of the move.
That’s a mouthful. Let’s make it easier.
Fix a horizon, say one day. Walk out from the average toward bigger and bigger moves and ask how fast the probability drops.
Gaussian answer: probability shrinks like e−x². Not just exponential, but exponential in the square of the move. Each additional unit of move costs you more probability than the last. A few units out and it’s effectively zero.
Power-law answer: probability shrinks like x^(−α), some fixed power of the move size. Double the move and you divide the probability by a constant factor (2^α), and you keep dividing by that same factor forever. There’s no cliff like the shoulder of a Gaussian curve.
The contrast is entirely about what sits where in the decay formula:
Base means ratios and slow.
Exponent means differences and fast.
A bell curve places the move in the exponent and squares it, so its tail vanishes. A power law leaves the move in the base, extending its probability further out in the wings.
To take the message of this 2-part series seriously means it’s unlikely that simply reading it imparted knowledge that you magically internalized (assuming you’re not multiple standard deviations up the IQ curve).
Instead, I hope you can employ AI tools to learn as you need. At your pace, and until your satisfaction. All those notes and highlights were not the learning itself but the fuel for powering a custom learning engine directed to your own goals and interests.
I of course hope that scaling laws, exponents, and variance were a desirable canvas to demonstrate the learning process, but they were not the point themselves. You can use any book as a starting point to go deeper. The combination of disaggregated training knowledge embedded in LLMs with narrower, concentrated material from an author or group of authors promises the best of 2 different advances against our humble ignorance.
Feel free to feed this post into your favorite agent to seed your own learning quest.
I usually have a concrete plan in advance of writing, but today’s letter is totally…
Return Stacked’s RSSB gives you a dollar of global equities and a dollar of Treasuries…
In this issue: a showdown with yourself unique job opening real-returns trading stock-bond correlation confidence…
At the end of July, Dean Curnutt tweeted: The thread should sound familiar. Weeks earlier, Dean…
Friends, I’m visiting my family in NJ so Moontower will return August 12th. That said,…
The market is 12-15. 12 bid. 15 offer. The broker sizes up the offer. “How…