Newsletter Thoughts

Socrates 2026: how to use highlights

Follows from Part 1: uncovering the laws of nature


Friends,

In Part 1, I teased that Geoffrey West’s Scale was a perfect surface to show how you can learn as you’ve always wanted. Or needed but didn’t know it.

Plan

  1. Cover how I used LLM to self-teach, which you can use for learning or re-learning anything.
  2. Cement and practice our understanding of power functions

🧠For those familiar with learning science, you will recognize several techniques, but I’ll label them as they appear.

How this all started

When I read a physical book, I will usually take a screenshot and then OCR the page to keep a digital excerpt. This is ok if there aren’t a lot of excerpts or highlights to preserve. I quickly realized I was going to need the Kindle version of Scale as the highlights and their accompanying inconvenience were piling up fast. I snagged it on Libby (this is your library’s digital loan service). There was no waitlist. Yet another reminder that there is so much joy available for free.

We must talk about highlighting.

The naive understanding of highlighting is that by taking the effort to trace a 25% opacity yellow film over words, you have learned something. By now you know this is item #77 on the list of self-deceptions. Still we carry on because it’s a cheap option. Somewhere in the recesses of your dopamine-addled mind (remember dopamine is the “seeking” chemical), you expect the highlights you stash like old coax cables will find fresh life when recombined with technology. Don’t be hard on yourself, an impotent but aspirational habit ranks less than wisdom but higher than apathy.

But it turns out this self-deception call option may finally have a payoff.

As I was marking up Scale, I had no guilt about not internalizing what I was reading in the moment. My plan was to export all the highlights like I usually do.

⚠️Kindle formats usually limit your exports to 10% of the book for copyright reasons. It’s a bit messy, but once I think I’m about there, I export those highlights to a file, then delete them in the Kindle app. For Scale, this happened when I was about half finished with the book. I could then start highlighting from zero for the second half.

This time, instead of just storing them, I was going to give them to a Claude project to seeding a “curriculum” so I could learn in a way that only comes from practice, not recitation. Reviewing your notes/highlights lets you cram for a test, but it’s not the kind of learning you can call on for invention.

Socrates 2026

Let’s rewind for a moment.

A few months back I bombed a Jane Street interview question I found online. This wouldn’t normally bother me as I’m far past the time in my life where my self-esteem teeters on an illusion of cleverness. But it was a question I felt I should know how to answer as opposed to the corpus of questions from which I wouldn’t even know where to begin.

You can see my write-up about it here: turning a Jane Street interview question blunder to a lesson

I realized that I couldn’t answer the question because I didn’t fully appreciate that variance is the spread between the expectation of a square and the square of an expectation.

Adjacent thoughts

  1. That variance is always non-negative is a demonstration of Jensen’s inequality operating on a function that takes the sum of squared deviations.
  2. Variance can actually be a little easier to appreciate as an instance of covariance between a random variable and itself!

My knowledge of variance was vague and formulaic. My knowledge of many things is like that. I don’t find that comforting, just a practical necessity in a limited life. But part of life’s pleasures is the freedom to NOT 80/20 something if doing so bugs you.

Alas, this one bugged me and the cost to fix it is lower now since LLM’s can be used as tireless tutors whose judgement of our faculties presents no threat.

I opened a chat and asked it to teach me Socratically, one small question at a time. When I run out of time or get tired, I know I can pick up where I left off. Or a little bit before that, since I usually need to insert before the point where I got tired since that point coincides with the material that made you take a break, so you don’t quite “own” it.

This is a snapshot of where I am in my Variance progression where I derive every formula from the already intuitive definition of “sum of squared deviations”:

 

As the learning progresses, you build more cases. With practice, you see that the key to all the derivations is that you are building on things you already know:

  1. The FOIL method from algebra
  2. PEMDAS from arithmetic
  3. The substitution that comes from seeing expectation or E[X] as nothing but a weighted average which means it’s equivalent to x_bar when each sample has equal weight

It’s hard to see this without practice.

When I revisit some of the derivations, I sometimes get stuck again but I know I’m screwing up one of these 3 foundational elements. That’s pretty crazy. It’s a gap in something I thought I knew cold, but the diagnosis is far more apparent because of how I’ve structured the learning in cahoots with Claude.

This is a timely place to name a few learning science techniques at play (see the appendix for more on these):

  • deliberate practice — deriving every variance formula from the definition of “sum of squared deviations,” over and over, refining each pass
  • desirable difficulty — doing the algebra by hand and taking pictures of the scratch work instead of watching it get done
  • spaced repetition — returning to the same derivation threads over days and weeks, not one sitting
  • expert guidance — Claude posing the next question and catching my errors with numerical counterexamples
  • layering skills — building each new case on FOIL, PEMDAS, and E[X]-as-weighted-average, things I already own
  • expertise reversal effect — starting with scaffolded one-question-at-a-time prompts rather than open-ended problems
  • consolidation — having Claude summarize what stuck, weighted to my actual gaps and the spots I tripped

The entire process is infused with the “generation effect” which takes advantage of our ability to remember something far better when you produce the answer yourself than when you read it.

And finally, every topic is a branch of an overarching commitment to interleaving. Power functions are mixed into a learning practice that includes other topics I want a closer look at. The approach makes affordances for both variety and synergy.

Before getting back to Scale and power functions, I have one more remark on this whole personal project I’ve donned Socrates 2026.

I really want AI companies to launch a Native Ink Surface with a submit button. Math derivation, music notation, art. All of these would be far less painful with a stylus. Is this too much to ask:


Automaticity

As I was reading and highlighting Scale, I strained to interpret the exponents. That means there are gaps in my understanding. Simple as that. These aren’t new concepts, but it’s clear I need some mental Dap if I want “automaticity”.

Paraphrasing Math Academy:

Automaticity is the ability to recall foundational math facts instantly and accurately from long-term memory, requiring zero conscious effort or working memory….

Automaticity is the prerequisite to true computational fluency. Once low-level skills (arithmetic, exponent rules, trig identities) are automated, recall becomes effortless, allowing your brain to focus entirely on higher-level problem solving and critical thinking.

We’ve been taught to think of tests (ie retrieval practice) as how you check whether you learned something. But it’s actually how you learn in the first place because it’s “doing”. To learn in a durable way is to “do”.

The highlights I collected became the raw material for Claude to design questions. But AI is obviously capable of far more than regurgitation, distillation and re-shuffling. It constructs sensible questions that arise from the text but not directly addressed. It can order the questions so they build gradually. It can relate material across domains. This is a gift to a learner.

Injecting a thought

AI cannot motivate you. It cannot inspire you. AI offers an unbundling of the tutor, not a replacement. The role of humans in the learning loop is going to grow, which might be a contrarian position. Think of coaches. Some are exceptional because they are masters of the Xs and Os. Some are exceptional because of their ability to lead and communicate. These are squishy. The squishy things will not rise in relative importance. They are important and AI doesn’t change that either way. It’s that AI will put a spotlight on the fact that there will be relatively higher yields to focus on the squishy. Whether we will or not (and be able to judge the delta) is an open question. A topic for another day perhaps, but I’m betting on this with my time.

There are no shortcuts. If you want automaticity, you gotta hit the gym.

The reps

We’re working with y = xᵃ throughout. The exponent a is the only thing carrying information about the relationship.

Warmup.

y = x². If x doubles, what happens to y?

POLL

y = x². If x doubles, what happens to y?

Goes up by 2
Goes up by a factor of 4
Goes up by a factor of 8
Stays the same
57 VOTES · · SHOW RESULTS

 

POLL

The rule: multiply x by some factor F, multiply y by F to the exponent. Here F is 2 and the exponent is 2, so y goes up by 2² = 4. Same law, y = x². If x triples?

Goes up b 3
Goes up by 6
Goes up by a factor of 9
Goes up by a factor of 27
38 VOTES · · SHOW RESULTS
POLL

In the last question, the exponent is fixed. You just swap the multiplier. Now a square root. y = √x, which is y = x¹ᐟ². If x quadruples, what happens to y?

Quadruples
Doubles
Goes up by a factor of 8
Halves
35 VOTES · · SHOW RESULTS

That question should feel familiar. Option prices follow a square root relationship with respect to time.. Doubling the time to expiry only multiplies a straddle by √2, while quadrupling it doubles the straddle.

POLL

Kleiber’s law. Metabolic rate scales as mass³ᐟ⁴. A mammal’s mass doubles. Its metabolic rate goes up by:

Exactly 2x (it doubles)
More than 2x
Less than 2x
It halves
30 VOTES · · SHOW RESULTS

The exponent is less than 1, so y grows slower than x. Less than double. This is the whole idea of sublinear scaling and economy of scale. Double the animal and it needs about 68% more energy, not 100% more.

POLL

Is 2³ᐟ⁴ the same as 2³ / 2⁴?

Yes, both equal 0.5
No, they’re different operations
Yes, both equal about 1.68
They’re both undefined
27 VOTES · · SHOW RESULTS

 

A fraction in the exponent is one number, not a division. 2³ᐟ⁴ means “take the fourth root of 2, then cube it,” which is about 1.68. Meanwhile 2³ / 2⁴ = 2³⁻⁴ = 2⁻¹ = 0.5. Dividing powers subtracts exponents.

POLL

What is 2⁻¹ᐟ⁴?

1/16
About .84
-1.19
-16
25 VOTES · · SHOW RESULTS

 

A negative exponent is always a reciprocal. Compute the positive version, then flip.

POLL

City infrastructure scales as population⁰·⁸⁵. A city’s population doubles. Total road length multiplies by roughly:

2.0
1.8
1.4
.85
22 VOTES · · SHOW RESULTS

2⁰·⁸⁵ sits between 2⁰·⁵ ≈ 1.41 and 2¹ = 2, closer to 2 because the exponent is close to 1. About 1.8. Roads go up 80% when the city doubles. Sublinear again, which means per person, road length actually falls.

POLL

City wages and output scale as population¹·¹⁵. Population doubles. Total wages multiply by roughly:

2.15
2.0
2.2
4.0
23 VOTES · · SHOW RESULTS

 

2¹·¹⁵ ≈ 2.22. Careful here. The move is NOT “2 plus 0.15.” It’s 2¹ × 2⁰·¹⁵ = 2 × 1.11 ≈ 2.22. Superlinear. Bigger city, disproportionately more output per person. Also disproportionately more crime and disease. Good and the bad scale together.

POLL

Strength scales as weight²ᐟ³. A horse weighs 8x what a small dog weighs. Per pound of body weight, the horse is:

Stronger than the dog
Exactly as strong per pound
Half as strong per pound
A quarter as strong per pound
22 VOTES · · SHOW RESULTS

 

Horse is 8²ᐟ³ = 4x stronger in total, but 8x heavier. So per pound it’s 4/8 = half as strong. This was Galileo’s observation. A small dog can carry two or three dogs on its back. A horse can’t carry even one. Strength grows like area, weight grows like volume, and volume outruns area as things get bigger.

Why Godzilla can’t exist

Strength scales with cross-sectional area, not size. This is why lumber is sold as a “2×4.” The two-by-four inches of cross-section is what bears the load. Double every dimension of a beam and its strength goes up 4x, because area scales with length squared.

Mass scales with volume, which is length cubed. Double every dimension and the thing weighs 8x more.

Scale a creature up and its weight (volume, 8x) outruns its strength (cross section, 4x) with every doubling. At Godzilla’s size, the legs would have to support a mass that has exploded as the cube of height while the bones holding it up only got stronger as the square. He’d snap under his own weight before he took a step. Same reason an ant can carry many times its body weight and an elephant can barely carry its own.

The toolkit

You build the knowledge, check yourself on new questions, come back another day, see how much ground you gave back. It’s 2 steps forward, 1 step back. Eventually you earn the consolidated reference and a sense that you have earned the shortcuts.

For y = xᵃ:

  1. Multiply x by F, and y multiplies by Fᵃ. This is scale invariance.
  2. For every order of magnitude in x, y changes by a orders of magnitude. This is the log-log slope reading.
  3. Double x, and y changes by a factor of 2ᵃ. This is the doubling sentence, the one West uses constantly. It’s natural for us to think of scaling with respect to doubling.
  4. Per unit of x, the quantity scales as xᵃ⁻¹. This is economy of scale versus increasing returns.

 

Interpreting the power

  • The sign tells you direction.
  • The magnitude tells you speed.
  • The distance from 1 tells you how it compares to a plain linear relationship.

Three worked slopes, read in orders of magnitude

y = x¹ᐟ², slope one half. y grows at half the rate of x. x goes up two orders of magnitude, y goes up one. Or: to get y up one order of magnitude, x has to move two.

y = x¹ᐟ⁴, slope one quarter. Even more damped. x times 10,000 (four orders of magnitude), y only times 10 (one order). This is per-cell metabolism, which scales as mass⁻¹ᐟ⁴.

y = x³ᐟ⁴, slope three-quarters. Kleiber. x times 10,000, y times 1,000. Three orders of magnitude of metabolism per four orders of magnitude of mass. The “3 to 4 ratio in powers of ten”.

The economy-of-scale shortcut

If the total scales as xᵃ, then per unit of x it scales as xᵃ⁻¹. Subtract 1 from the exponent and you have the per-person, per-pound, per-cell law. The sign of a−1 is the whole story:

  • a > 1: per-unit grows. Increasing returns.
  • a = 1: per-unit flat. Constant returns.
  • a < 1: per-unit shrinks. Economy of scale.

Cities have two exponents that mirror each other around 1. Physical stuff like roads and cables scales at 0.85, so per person it gets cheaper as the city grows. Social stuff like wages and patents scales at 1.15, so per person it grows.

Power law versus exponential: application to tail probability

Power law: the variable is in the base. It cares about ratios. Multiply the input, multiply the output, and the multiplier is the same no matter where you started.

Exponential: the variable is in the exponent. It cares about differences. Add to the input, multiply the output.

CAGR is a familiar exponential. Consider 10% CAGR. Going from year 5 to year 10 does not multiply your balance by the same factor as going from year 30 to year 60, even though both double the time. The multiplier depends on where you are, not on the ratio.

That base-versus-exponent distinction isn’t just about growth over time. It also governs how the probability of a large moves. The mean-standard deviation framework we are so familiar with is Gaussian “mediocrastian” math. But we know empirically that tails reside in “extremistan”. You can’t have a 10-sigma move every few decades. The bell curve is just a misspecified description of returns.

The fat tail in returns is more of a power law in the size of the move.

That’s a mouthful. Let’s make it easier.

Fix a horizon, say one day. Walk out from the average toward bigger and bigger moves and ask how fast the probability drops.

Gaussian answer: probability shrinks like e−x². Not just exponential, but exponential in the square of the move. Each additional unit of move costs you more probability than the last. A few units out and it’s effectively zero.

Power-law answer: probability shrinks like x^(−α), some fixed power of the move size. Double the move and you divide the probability by a constant factor (2^α), and you keep dividing by that same factor forever. There’s no cliff like the shoulder of a Gaussian curve.

The contrast is entirely about what sits where in the decay formula:

  • Gaussian: the move x is up in the exponent (e−x²). Move in the exponent means probability responds to differences in move size.
  • Power law: the move x is in the base (x−α). Move in the base means probability responds to ratios of move size.

Base means ratios and slow.

Exponent means differences and fast.

A bell curve places the move in the exponent and squares it, so its tail vanishes. A power law leaves the move in the base, extending its probability further out in the wings.

Handy intuition (and trivia!)

 

Wrapping up

To take the message of this 2-part series seriously means it’s unlikely that simply reading it imparted knowledge that you magically internalized (assuming you’re not multiple standard deviations up the IQ curve).

Instead, I hope you can employ AI tools to learn as you need. At your pace, and until your satisfaction. All those notes and highlights were not the learning itself but the fuel for powering a custom learning engine directed to your own goals and interests.

I of course hope that scaling laws, exponents, and variance were a desirable canvas to demonstrate the learning process, but they were not the point themselves. You can use any book as a starting point to go deeper. The combination of disaggregated training knowledge embedded in LLMs with narrower, concentrated material from an author or group of authors promises the best of 2 different advances against our humble ignorance.

Feel free to feed this post into your favorite agent to seed your own learning quest.

editor

Share
Published by
editor

Recent Posts

“why should I learn this?”

I usually have a concrete plan in advance of writing, but today’s letter is totally…

2 hours ago

stock-bond correlation

Return Stacked’s RSSB gives you a dollar of global equities and a dollar of Treasuries…

3 days ago

Moontower #324

In this issue: a showdown with yourself unique job opening real-returns trading stock-bond correlation confidence…

3 days ago

collar shopping

At the end of July, Dean Curnutt tweeted: The thread should sound familiar. Weeks earlier, Dean…

1 week ago

Moontower #323

Friends, I’m visiting my family in NJ so Moontower will return August 12th. That said,…

2 weeks ago

the sound of inevitability

The market is 12-15. 12 bid. 15 offer. The broker sizes up the offer. “How…

2 weeks ago