after this post you will be sizing bets in your head

One of the most important concepts in risk-taking is bet sizing. Which is unfortunate because people are quite bad at it, while the effort to be way above average is quite low.

A jarring and famous demonstration of this is the Haghani-Dewey Coin Flipping study, which showed how even college grads with business, economic, and technical backgrounds incinerated their capital or massively underperformed the expected profits presented to them by a game they knew was rigged in their favor.

You can read my synopsis in Bet Sizing Is Not Intuitive.

For a binary wager (ie win or lose), if you know the payoffs and the probability of winning, both of which were known to the participants, the solution is to use the Kelly Criterion.

The tragedy is that it is incredibly simple to compute and applies to many conventional gambles and decisions (the examples in the quiz will span various life situations!).

If something is both easy and widely relevant, it should be common knowledge. So let’s fix that today. I’ll show you how easy it is to use, and you’ll forever be able to do it in your head.

First, a succinct definition:

Kelly is the bet size, as a fraction of bankroll, that maximizes the long-run compounded growth rate of your wealth. It’s a mathematical solution to bet size that doesn’t seek to maximize expected profit per trial, but the size that optimally balances compounding rate and survival.

If you want to go deep on this, see Moontowerquant’s Kelly Criterion Resources, but today’s focus is on getting straight to usability.

We will use this formulation of Kelly because it’s general:

f* = p − q/b

where:

p = probability of winning

q = 1−p or probability of losing

b = the odds you’re getting → what you win divided by what you risk

The easiest way to learn it is just jump right in with a few worked examples:

Fair coin wager (even odds style bet)

p =50%

q= 50%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 50% – 50% / 1 = 0 → bet nothing, zero edge

Coin biased in your favor (even odds style bet)

p =60%

q= 40%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 60% – 40% / 1 = .20 → bet 20% of your bankroll

Roll a 6 on a die (underdog bet where you get odds)

p =1/6

q= 5/6

b =8 (for a $1 bet, you either lose a $1 or make an $8 profit)

f* = 1/6 – (5/6) / 8

f* =8/48 – 5/48 = 3/48 → bet 6.25% of your bankroll

If f* is 0 or negative, you have no edge, so not betting is prescribed

Sports moneyline (betting as a favorite where you lay odds)

A −200 favorite. You risk $2 to win $1, and the line implies 2/3, but you think it’s closer to 3 in 4.

p = 75%
q = 25%
b = 0.5 (getting 50% return on the amount you risk)

f* = 75% − 25% / 0.5 = 75% − 50% = 25% → bet 25% of your bankroll

Wait a minute, these are large bets?!!

If these bet sizes seem surprisingly large for the given advantages, then your senses are well-tuned. For most people, “full” Kelly is too big!

Kelly maximizes long-run growth on the assumption your probability is correct. Well, it probably isn’t because the world is messy. We can inject some humility by using a fraction of Kelly:

  • “Half Kelly” gives up about a quarter of the growth rate and roughly halves the drawdowns. If you invert that, you see that the Kelly scaling law means as you bet bigger, you get diminishing returns per unit of risk. Extrapolating that logic, betting more than “full Kelly” is incinerating compounded wealth even if the individual bet has positive EV.
  • “Quarter Kelly” or less is far more common in practice.

Please don’t let the equation scare you, it’s intuitive and easy to remember

Look at the equation again:

f* = p − q/b

It’s just “how often you win” minus “how often you lose.” It’s just that the second term incorporates the payoff. The loss term gets divided by b, which represents the return you collect when you’re right.

  • When b = 1, you’re getting even money. A 100% return. Dividing by 1 leaves q alone, and the equation collapses to pure hit rate: p − q. That’s the coin case where you bet $1 to make $1.
  • When b > 1, you’re getting long odds. The division shrinks the loss term. This is why the die works. You lose 5 out of 6 rolls. Straight subtraction says you’re down 66 cents on the dollar and should never play, but you’re paid 8-to-1, so that 5/6 becomes 5/48, and suddenly the 1/6 win percentage is the bigger number. Long odds forgive a bad hit rate.
  • When b < 1, you’re laying odds. Now division stretches the loss term. The moneyline: you only lose a quarter of the time, but at −200 each loss costs you two units to earn back one, so that 25% loss percentage behaves like 50%. Being right three times out of four barely clears the bar. Lay enough odds and even a very good record is a losing proposition.

The graphic shows how you can think of the odds (the denominator) as shrinking or inflating q, as you collapse your thinking to a comparison of p vs q.

A word on b

b trips people up because “odds” is loaded gambler jargon. A wider interpretation of b is that it’s a percent return.

It’s what you make divided by what you risk. Even money is b = 1: risk a dollar, make a dollar. That’s a 100% return on the amount at stake. 3-to-1 is b = 3, a 300% return. Laying −200 is b = 0.5 because if you risk two to make one, it’s a 50% return.

[Return is a profit, while multiples don’t subtract your initial risk. It’s the difference between “I 2x’d my money” vs “I made 100%” or “I 10x’d my money” vs “I made 900%”. The percent return is the multiple minus one because we subtract our initial risk.]

The reason it’s a return and not just “the odds” is that Kelly assumes a loss wipes out the whole stake. The denominator is always the same number: everything you put up. b is comparable across a coin, a die, and a moneyline because it’s the return on risk, always measured against a total loss.

b = 1 is a natural reference point. At 100% return, a win exactly cancels a loss, so you need to win more than half the time. The breakeven hit rate changes with b.

Set f* = 0 and you get p = 1/(1+b).

Read the table as a menu of the hit rates you’re allowed to have. At b = 24 you can be wrong 24 times out of 25 and still be flat. At b = 0.25 you can be right four out of five and still be flat

Applying to real life: when is Kelly the right tool?

Kelly needs a few inputs: a bankroll, a payoff you know, and a probability estimate.

Which of these is a Kelly problem?

  1. A prediction market contract trading at 30¢. You think it’s worth 45¢.
  2. You’re all-in-or-fold on the river with a read that you’re good 40% of the time, getting 3-to-1 from the pot.
  3. How much of your 401(k) to put in equities.
  4. Writing checks as an angel investor across 30 startups.
  5. Buying a weekly call on a biotech ahead of an FDA decision date.
  6. Whether to take the new job.
  7. Your buddy offers you 5-to-1 that it rains in Oakland tomorrow. The forecast says 30%.
  8. Buying homeowners insurance. The premium is clearly more than the expected loss — that’s how the insurer stays in business.
  9. Your neighbor doesn’t carry homeowners coverage. She banks the premium instead.
  10. Your auto policy offers a $500 deductible or a $2,500 deductible, for a $340/yr discount on the premium.
  11. You’ve got vested startup options. Exercising costs $40k out of pocket in strike, and you think there’s maybe a 15% chance the company gets somewhere that makes them worth $1M.
  12. A merger arb spread. Target’s at $46, deal price is $50, and it trades back to $38 if the deal breaks. You think it closes 90% of the time.
  13. Your agency spends 20 hours of unbilled time on a speculative pitch. You win about a quarter of them, and a win is worth 80 billable hours.

Solutions to Kelly Problems

  1. Yes. Cleanest case there is. Binary, known payoff, and the price provides b directly. Risk 30¢ to make 70¢, so b = 2.33. f* = 45% − 55%/2.33 = 21%.
  2. Yes. This is the canonical one. p = 40%, b = 3, f* = 40% − 60%/3 = 20% of your stack. The wrinkle is that in poker your stack isn’t really your bankroll. There’s a whole literature on pros using Kelly for bankroll management across sessions rather than for a single river decision.
  3. No. Not this version of it. Stock returns aren’t generally binary so there’s no p, q, or b. There’s a continuous analog called Merton’s Share, which is similarly rooted in reward vs variance. What Gamblers Can Teach the Buy-and-Hold Crowd can get you started.
  4. Sort of. The structure is right: repeated, roughly binary, long odds. The problem is that p is a guess and b is a bigger guess, and Kelly is violently sensitive to overestimating your edge. Garbage in, garbage out.
  5. Approximately. If you treat it as approve/reject it’s binary enough to size with. If the expiry aligns with the date such that you are betting strictly on the terminal intrinsic value, the option piece will inherit the binary modeling you imposed on the stock.
  6. No. The variables are too opaque.
  7. Yes. p = 30%, b = 5, f* = 30% − 70%/5 = 16%. Note, you’ll lose this bet more than twice as often as you win it, so your most likely scenario is losing 16%. You can shrink the Kelly fraction if this makes you uncomfortable.
  8. Wrong side of the equation. Run f* on this and you get a negative number, because you’re buying a negative-EV bet.
  9. Yes. It’s the same policy, so notice that the bet only exists on the insurer’s side of it! Every year your neighbor doesn’t buy, she collects a premium and writes a tail. Rebuild cost $500k, premium $3,000, call it a 1-in-500 chance of a total loss.

    p = 99.8%
    q = 0.2%
    b = 3,000 / 500,000 = 0.006 (risk $500k to win $3,000)

    f* = 99.8% − 0.2%/0.006 = 99.8% − 33.3% = 66.5%

    Positive, as expected since insurers price premiums well above fair value. In this case, her bet size is the $500k house. If the house is most of her net worth, she’s at 100% on a bet capped at 66%. Rather than overbet, she should buy the policy. If she’s worth $5M, she’s betting 10% when she’s allowed 66%, which puts her near quarter Kelly (~14%), and she could skip the insurance. If she’s worth $1mm, it’s a 50% bet, which is more than half Kelly. I’d say take the insurance but I’m a wimp. There are other considerations (would she have the liquidity to rebuild the home or maybe taking the insurance with a high deductible is a better fit), but just doing this exercise gives you a sense of how risky or conservative your choices are relative to the bet share that maximizes long-term wealth.

  10. Yes. Raising the deductible is like you writing a $2,000 policy and collecting $340 a year for it. You’re the insurer again, so work out what you need to believe. Take the high deductible and save $340. Have a claim, and you’re out $2,000 more, but you already banked the $340, so the loss is $1,660.

    b = 340 / 1,660 = 0.205

    f* = p − q/0.205 = p − 4.88q

    Set that to zero, and you get p = 4.88q, which, with p + q = 1, means q = 17%. Your breakeven is a claim every 5.9 years. Anything less frequent and you’re the one with the edge.

    Let’s say real-world collision frequency is more like 6%. So p = 94%:

    f* = 94% − 6%/0.205 = 94% − 29.3% = 64.7%

    Which says risk at most ~65% of your bankroll. The risk here is $1,660. That clears as long as you have about $2,600 in liquid savings, which is to say the sizing check is trivially satisfied for almost everyone so you should generally opt for the higher deductible. For quarter Kelly, we’d need savings of $1,660/(.647 * .25) = $10,262.

  11. Approximately. It’s not truly binary, but if you frame it in a way where you are comfortable with the no consolation prize of a medium outcome, you can see it as paying $40k for a 15% shot at $1M. b = 24, so f* = 15% − 85%/24 = 11.5% of your liquid net worth. This is quite sensitive to your estimate of p of course.
  12. Yes, a classic example of binary-type risk in markets. You risk $8 to make $4, so b = 0.5 — you’re laying odds, same as the moneyline. f* = 90% − 10%/0.5 = 70%. That number is only as good as the 90%. Revise p to 75% and f* is 25%.
  13. Yes, in a subtle way! Your bankroll is capacity, not cash. b = 80/20 = 4, so f* = 25% − 75%/4 = 6.25%. Twenty hours has to be 6% of what you’re working with, which means you can’t run this pitch out of a 100-hour month. If you’re the manager, you can put it in dollar terms by converting to wages.

Finally, I strongly recommend William Poundstone’s book Fortune’s Formula: The Untold Story of the Scientific Betting System That Beat the Casinos and Wall Street

Description:

In 1956, two Bell Labs scientists discovered the scientific formula for getting rich. One was mathematician Claude Shannon, neurotic father of our digital age, whose genius is ranked with Einstein’s. The other was John L. Kelly Jr., a Texas-born, gun-toting physicist. Together they applied the science of information theory—the basis of computers and the Internet—to the problem of making as much money as possible, as fast as possible.

Shannon and MIT mathematician Edward O. Thorp took the “Kelly formula” to Las Vegas. It worked. They realized that there was even more money to be made in the stock market. Thorp used the Kelly system with his phenomenally successful hedge fund, Princeton-Newport Partners. Shannon became a successful investor, too, topping even Warren Buffett’s rate of return. Fortune’s Formula traces how the Kelly formula sparked controversy even as it made fortunes at racetracks, casinos, and trading desks. It reveals the dark side of this alluring scheme, which is founded on exploiting an insider’s edge.

Shannon believed it was possible for a smart investor to beat the market—and William Poundstone’s Fortune’s Formula will convince you that he was right.

And this is from my notes, Insights From Fortune’s Formula:

a gripping narrative full of 20th century trivia that ties together the birth of information theory, some of the greatest scientific minds of the 1900s, the rise of quantitative finance, and the role of organized crime. These topics come alive in a fresh, memorable way when discovered through the lens of its colorful characters.

It chronicles the history of the efficient market hypothesis (MIT, U Chicago, Paul Samuelson). You can organize its conclusion around this excerpt:

There is much truth in the efficient market hypothesis. The controversy has always been over just how far the claim can be pressed. Asking whether markets are efficient is like asking whether the world is round. The best way to answer depends on the expectations and sophistication of the questioner. If someone is asking whether the world is round or flat, as fifteenth-century Europeans might have asked, then “round” is a better answer. If someone knows that and is asking whether the earth is a geometrically perfect sphere, the answer is no.

Moontower #327

In this issue:

  • slow burn
  • mental bet sizing for everyone
  • sizing for real-life decisions

Friends,

Riddle me this:

Reading moontower makes me feel…

⏹️Smart

⏹️Stupid

(read the article on substack to actually submit an answer)

Over the years, I’ve administered 2 reader polls to take the pulse of why you bother reading this newsletter. Both times I’ve asked the above question but with an open-ended fill-in-the-blank template. The answers come back in all shapes and sizes, but most of them can be boiled down to “smart” or “stupid”.

My reactions to the “stupid” responses:

My writing has a style that can be, at times, counterproductive. It’s often unfiltered and parenthetical. My wife constantly tells me that if I write more clearly, I’d have a bigger audience, and she’s, of course, right. The writing is not optimized for clarity. Scott H Young is my canonical example of such writing. I’m a big fan, but my lack of self-control forbids this common-sense approach.

My writing is aspirationally in the spirit of a Pixar movie where it serves 2 masters: the child and the adult who brought them to the theater. I’m trying to teach relatively basic or intermediate concepts while leaving Easter eggs for readers who arrive with more finance/trading context. This will leave the more basic readers feeling like they are missing a joke.

Fellow trader and writer RobotJames, once told me that Moontower is a “slow burn”. There’s no single banger that you could send to someone and say “this is why you should read this”, but reading it week in and week out leads to something more than the sum of its parts. My lament over never having a truly massive hit notwithstanding, I will not complain about the “slow burn” status. Many of my favorite things often took time to warm up to. Acquired tastes are rewards for crossing rugged terrain. (Except when they’re status cosplays. Our wants are even opaque to ourselves. There’s a thin line between appreciation and snobbery and its placement is up to the onlookers, not the admirer.)

To feel stupid when reading moontower would be natural. There’s technical or domain-specific minutiae mixed in with the meat and potatoes. This is further exacerbated by the cohort of readers who only show up for the personal sections and whose interests lie especially far from any of the material I cover. But if I read the blog of a doctor friend because it’s a way to follow along their life, the object-level material is gonna make me feel stupid.

And of course sometimes you feel stupid because I fail. The sequence of words fails to unlock the concept. Today I’m actively trying not to f this up. I want to teach you a neat topic. My goal is to really cement it for you. To make it accessible in real life for you in the same way multiplication tables are at your disposal.

If you don’t regularly read Money Angle, maybe see if today works for you. It’s long, but that’s because we are going to drill the concept. If you succeed in walking away with a new mental tool to practice, then the ROI on the reading time will look like a steal. If you’re not interested, then I’ll channel one of my college roommates whenever I tried to bail on hitting the gym, ”That’s cool man, I’ll just take you off my list of successful people for today” and leave you one of the OG memes that deserves annual visitation.

 

Commencing demystifications now…

Money Angle

One of the most important concepts in risk-taking is bet sizing. Which is unfortunate because people are quite bad at it, while the effort to be way above average is quite low.

A jarring and famous demonstration of this is the Haghani-Dewey Coin Flipping study, which showed how even college grads with business, economic, and technical backgrounds incinerated their capital or massively underperformed the expected profits presented to them by a game they knew was rigged in their favor.

You can read my synopsis in Bet Sizing Is Not Intuitive.

For a binary wager (ie win or lose), if you know the payoffs and the probability of winning, both of which were known to the participants, the solution is to use the Kelly Criterion.

The tragedy is that it is incredibly simple to compute and applies to many conventional gambles and decisions (the examples in the quiz will span various life situations!).

If something is both easy and widely relevant, it should be common knowledge. So let’s fix that today. I’ll show you how easy it is to use, and you’ll forever be able to do it in your head.

First, a succinct definition:

Kelly is the bet size, as a fraction of bankroll, that maximizes the long-run compounded growth rate of your wealth. It’s a mathematical solution to bet size that doesn’t seek to maximize expected profit per trial, but the size that optimally balances compounding rate and survival.

If you want to go deep on this, see Moontowerquant’s Kelly Criterion Resources, but today’s focus is on getting straight to usability.

We will use this formulation of Kelly because it’s general:

f* = p − q/b

where:

p = probability of winning

q = 1−p or probability of losing

b = the odds you’re getting → what you win divided by what you risk

The easiest way to learn it is just jump right in with a few worked examples:

Fair coin wager (even odds style bet)

p =50%

q= 50%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 50% – 50% / 1 = 0 → bet nothing, zero edge

Coin biased in your favor (even odds style bet)

p =60%

q= 40%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 60% – 40% / 1 = .20 → bet 20% of your bankroll

Roll a 6 on a die (underdog bet where you get odds)

p =1/6

q= 5/6

b =8 (for a $1 bet, you either lose a $1 or make an $8 profit)

f* = 1/6 – (5/6) / 8

f* =8/48 – 5/48 = 3/48 → bet 6.25% of your bankroll

If f* is 0 or negative, you have no edge, so not betting is prescribed

Sports moneyline (betting as a favorite where you lay odds)

A −200 favorite. You risk $2 to win $1, and the line implies 2/3, but you think it’s closer to 3 in 4.

p = 75%
q = 25%
b = 0.5 (getting 50% return on the amount you risk)

f* = 75% − 25% / 0.5 = 75% − 50% = 25% → bet 25% of your bankroll

Wait a minute, these are large bets?!!

If these bet sizes seem surprisingly large for the given advantages, then your senses are well-tuned. For most people, “full” Kelly is too big!

Kelly maximizes long-run growth on the assumption your probability is correct. Well, it probably isn’t because the world is messy. We can inject some humility by using a fraction of Kelly:

  • “Half Kelly” gives up about a quarter of the growth rate and roughly halves the drawdowns. If you invert that, you see that the Kelly scaling law means as you bet bigger, you get diminishing returns per unit of risk. Extrapolating that logic, betting more than “full Kelly” is incinerating compounded wealth even if the individual bet has positive EV.
  • “Quarter Kelly” or less is far more common in practice.

Please don’t let the equation scare you, it’s intuitive and easy to remember

Look at the equation again:

f* = p − q/b

It’s just “how often you win” minus “how often you lose.” It’s just that the second term incorporates the payoff. The loss term gets divided by b, which represents the return you collect when you’re right.

  • When b = 1, you’re getting even money. A 100% return. Dividing by 1 leaves q alone, and the equation collapses to pure hit rate: p − q. That’s the coin case where you bet $1 to make $1.
  • When b > 1, you’re getting long odds. The division shrinks the loss term. This is why the die works. You lose 5 out of 6 rolls. Straight subtraction says you’re down 66 cents on the dollar and should never play, but you’re paid 8-to-1, so that 5/6 becomes 5/48, and suddenly the 1/6 win percentage is the bigger number. Long odds forgive a bad hit rate.
  • When b < 1, you’re laying odds. Now division stretches the loss term. The moneyline: you only lose a quarter of the time, but at −200 each loss costs you two units to earn back one, so that 25% loss percentage behaves like 50%. Being right three times out of four barely clears the bar. Lay enough odds and even a very good record is a losing proposition.

The graphic shows how you can think of the odds (the denominator) as shrinking or inflating q, as you collapse your thinking to a comparison of p vs q.

A word on b

b trips people up because “odds” is loaded gambler jargon. A wider interpretation of b is that it’s a percent return.

It’s what you make divided by what you risk. Even money is b = 1: risk a dollar, make a dollar. That’s a 100% return on the amount at stake. 3-to-1 is b = 3, a 300% return. Laying −200 is b = 0.5 because if you risk two to make one, it’s a 50% return.

[Return is a profit, while multiples don’t subtract your initial risk. It’s the difference between “I 2x’d my money” vs “I made 100%” or “I 10x’d my money” vs “I made 900%”. The percent return is the multiple minus one because we subtract our initial risk.]

The reason it’s a return and not just “the odds” is that Kelly assumes a loss wipes out the whole stake. The denominator is always the same number: everything you put up. b is comparable across a coin, a die, and a moneyline because it’s the return on risk, always measured against a total loss.

b = 1 is a natural reference point. At 100% return, a win exactly cancels a loss, so you need to win more than half the time. The breakeven hit rate changes with b.

Set f* = 0 and you get p = 1/(1+b).

Read the table as a menu of the hit rates you’re allowed to have. At b = 24 you can be wrong 24 times out of 25 and still be flat. At b = 0.25 you can be right four out of five and still be flat

Applying to real life: when is Kelly the right tool?

Kelly needs a few inputs: a bankroll, a payoff you know, and a probability estimate.

Which of these is a Kelly problem?

  1. A prediction market contract trading at 30¢. You think it’s worth 45¢.
  2. You’re all-in-or-fold on the river with a read that you’re good 40% of the time, getting 3-to-1 from the pot.
  3. How much of your 401(k) to put in equities.
  4. Writing checks as an angel investor across 30 startups.
  5. Buying a weekly call on a biotech ahead of an FDA decision date.
  6. Whether to take the new job.
  7. Your buddy offers you 5-to-1 that it rains in Oakland tomorrow. The forecast says 30%.
  8. Buying homeowners insurance. The premium is clearly more than the expected loss — that’s how the insurer stays in business.
  9. Your neighbor doesn’t carry homeowners coverage. She banks the premium instead.
  10. Your auto policy offers a $500 deductible or a $2,500 deductible, for a $340/yr discount on the premium.
  11. You’ve got vested startup options. Exercising costs $40k out of pocket in strike, and you think there’s maybe a 15% chance the company gets somewhere that makes them worth $1M.
  12. A merger arb spread. Target’s at $46, deal price is $50, and it trades back to $38 if the deal breaks. You think it closes 90% of the time.
  13. Your agency spends 20 hours of unbilled time on a speculative pitch. You win about a quarter of them, and a win is worth 80 billable hours.

Money Angle For Masochists

Solutions to Kelly Problems

  1. Yes. Cleanest case there is. Binary, known payoff, and the price provides b directly. Risk 30¢ to make 70¢, so b = 2.33. f* = 45% − 55%/2.33 = 21%.
  2. Yes. This is the canonical one. p = 40%, b = 3, f* = 40% − 60%/3 = 20% of your stack. The wrinkle is that in poker your stack isn’t really your bankroll. There’s a whole literature on pros using Kelly for bankroll management across sessions rather than for a single river decision.
  3. No. Not this version of it. Stock returns aren’t generally binary so there’s no p, q, or b. There’s a continuous analog called Merton’s Share, which is similarly rooted in reward vs variance. What Gamblers Can Teach the Buy-and-Hold Crowd can get you started.
  4. Sort of. The structure is right: repeated, roughly binary, long odds. The problem is that p is a guess and b is a bigger guess, and Kelly is violently sensitive to overestimating your edge. Garbage in, garbage out.
  5. Approximately. If you treat it as approve/reject it’s binary enough to size with. If the expiry aligns with the date such that you are betting strictly on the terminal intrinsic value, the option piece will inherit the binary modeling you imposed on the stock.
  6. No. The variables are too opaque.
  7. Yes. p = 30%, b = 5, f* = 30% − 70%/5 = 16%. Note, you’ll lose this bet more than twice as often as you win it, so your most likely scenario is losing 16%. You can shrink the Kelly fraction if this makes you uncomfortable.
  8. Wrong side of the equation. Run f* on this and you get a negative number, because you’re buying a negative-EV bet.
  9. Yes. It’s the same policy, so notice that the bet only exists on the insurer’s side of it! Every year your neighbor doesn’t buy, she collects a premium and writes a tail. Rebuild cost $500k, premium $3,000, call it a 1-in-500 chance of a total loss.

    p = 99.8%
    q = 0.2%
    b = 3,000 / 500,000 = 0.006 (risk $500k to win $3,000)

    f* = 99.8% − 0.2%/0.006 = 99.8% − 33.3% = 66.5%

    Positive, as expected since insurers price premiums well above fair value. In this case, her bet size is the $500k house. If the house is most of her net worth, she’s at 100% on a bet capped at 66%. Rather than overbet, she should buy the policy. If she’s worth $5M, she’s betting 10% when she’s allowed 66%, which puts her near quarter Kelly (~14%), and she could skip the insurance. If she’s worth $1mm, it’s a 50% bet, which is more than half Kelly. I’d say take the insurance but I’m a wimp. There are other considerations (would she have the liquidity to rebuild the home or maybe taking the insurance with a high deductible is a better fit), but just doing this exercise gives you a sense of how risky or conservative your choices are relative to the bet share that maximizes long-term wealth.

  10. Yes. Raising the deductible is like you writing a $2,000 policy and collecting $340 a year for it. You’re the insurer again, so work out what you need to believe. Take the high deductible and save $340. Have a claim, and you’re out $2,000 more, but you already banked the $340, so the loss is $1,660.

    b = 340 / 1,660 = 0.205

    f* = p − q/0.205 = p − 4.88q

    Set that to zero, and you get p = 4.88q, which, with p + q = 1, means q = 17%. Your breakeven is a claim every 5.9 years. Anything less frequent and you’re the one with the edge.

    Let’s say real-world collision frequency is more like 6%. So p = 94%:

    f* = 94% − 6%/0.205 = 94% − 29.3% = 64.7%

    Which says risk at most ~65% of your bankroll. The risk here is $1,660. That clears as long as you have about $2,600 in liquid savings, which is to say the sizing check is trivially satisfied for almost everyone so you should generally opt for the higher deductible. For quarter Kelly, we’d need savings of $1,660/(.647 * .25) = $10,262.

  11. Approximately. It’s not truly binary, but if you frame it in a way where you are comfortable with the no consolation prize of a medium outcome, you can see it as paying $40k for a 15% shot at $1M. b = 24, so f* = 15% − 85%/24 = 11.5% of your liquid net worth. This is quite sensitive to your estimate of p of course.
  12. Yes, a classic example of binary-type risk in markets. You risk $8 to make $4, so b = 0.5 — you’re laying odds, same as the moneyline. f* = 90% − 10%/0.5 = 70%. That number is only as good as the 90%. Revise p to 75% and f* is 25%.
  13. Yes, in a subtle way! Your bankroll is capacity, not cash. b = 80/20 = 4, so f* = 25% − 75%/4 = 6.25%. Twenty hours has to be 6% of what you’re working with, which means you can’t run this pitch out of a 100-hour month. If you’re the manager, you can put it in dollar terms by converting to wages.

Finally, I strongly recommend William Poundstone’s book Fortune’s Formula: The Untold Story of the Scientific Betting System That Beat the Casinos and Wall Street

Description:

In 1956, two Bell Labs scientists discovered the scientific formula for getting rich. One was mathematician Claude Shannon, neurotic father of our digital age, whose genius is ranked with Einstein’s. The other was John L. Kelly Jr., a Texas-born, gun-toting physicist. Together they applied the science of information theory—the basis of computers and the Internet—to the problem of making as much money as possible, as fast as possible.

Shannon and MIT mathematician Edward O. Thorp took the “Kelly formula” to Las Vegas. It worked. They realized that there was even more money to be made in the stock market. Thorp used the Kelly system with his phenomenally successful hedge fund, Princeton-Newport Partners. Shannon became a successful investor, too, topping even Warren Buffett’s rate of return. Fortune’s Formula traces how the Kelly formula sparked controversy even as it made fortunes at racetracks, casinos, and trading desks. It reveals the dark side of this alluring scheme, which is founded on exploiting an insider’s edge.

Shannon believed it was possible for a smart investor to beat the market—and William Poundstone’s Fortune’s Formula will convince you that he was right.

And this is from my notes, Insights From Fortune’s Formula:

a gripping narrative full of 20th century trivia that ties together the birth of information theory, some of the greatest scientific minds of the 1900s, the rise of quantitative finance, and the role of organized crime. These topics come alive in a fresh, memorable way when discovered through the lens of its colorful characters.

It chronicles the history of the efficient market hypothesis (MIT, U Chicago, Paul Samuelson). You can organize its conclusion around this excerpt:

There is much truth in the efficient market hypothesis. The controversy has always been over just how far the claim can be pressed. Asking whether markets are efficient is like asking whether the world is round. The best way to answer depends on the expectations and sophistication of the questioner. If someone is asking whether the world is round or flat, as fifteenth-century Europeans might have asked, then “round” is a better answer. If someone knows that and is asking whether the earth is a geometrically perfect sphere, the answer is no.

Stay groovy

☮️


Moontower Weekly Recap

Moontower #326

In this issue:

  • energy as the ultimate currency
  • “middle manager” fiction

Friends,

A year ago I wrote a piece called Capitalism Is a Temporary Condition. Buried in it was a thought:

An investor in an age of acceleration watches as commodity prices, real things, hover in familiar territory while “future cash flows discounted” reward attention — in some cases because of optimistic stories, in some cases because the company exchanges cash for BTC (an act which adds no economic value), and in some cases for nostalgic lolz.

Today, it’s embarrassing to think of actual value. The most basic form is stored energy**. An obvious example is oil. Slightly more abstract —a building is a collection of atoms that took energy to arrange and provides utility. It’s proof of work. Same with a good reputation.

The ** footnote:

I increasingly think that a durable concept of a risk-free or least-risky rate is more likely to come from Vaclav Smil than U Chicago.

I knew about Smil from a profile piece that celebrated his suffer-no-fools rigor in making sense of the world.

I’m finally getting around to How the World Really Works, and I dig how quickly it’s zeroing in on what I anticipated to be the biggest muscle movement of all. I’ll let it unfold as he does the book’s intro.

Smil asks us to imagine an alien probe watching Earth, programmed to wake up whenever something important changes in the way energy is captured or used.

For a very long time, it sleeps.

Then life figures out photosynthesis: sunlight can be captured and stored as chemical energy. Much later, humans learn to control fire. Then agriculture. Draft animals. Wind and water.

The gaps are enormous. Billions of years. Millions. Thousands.

Then they start collapsing.

Coal gives humans access to immense stores of ancient solar energy. You can see a parallel to wealth in that phrase alone. Steam will turn that energy into mechanical work, replacing what he calls prime mover energy (human and animal labor). Oil makes enormous amounts of energy portable while electricity makes it transmissible and almost infinitely adaptable.

Smil puts numbers on this.

In 1800, the world had access to roughly 0.05 gigajoules of useful energy per person each year. By 1900 it was 2.7. By 1950, about 10. By 2000, 28. By 2020, roughly 34.

In a little over two centuries, useful energy per person increased nearly 700-fold.

The average person today commands roughly the energy equivalent of 60 adults working continuously, day and night. In affluent countries, the equivalent is closer to 200–240 permanent laborers.

And the averages hide enormous differences.

Measured in primary energy, someone in one of the world’s poorest countries may use less than 10 GJ a year. India is around 20–30. The global average is roughly 75–80. China is above 100. Western Europe and Japan generally sit around 120–160. The United States is closer to 250–280.

It sounds a lot like how GDP per capita is distributed. We measure GDP in dollars, but money can be printed, thus distorting its exchange rate versus the stored value of work. In that sense, energy may be civilization’s only uncorruptible currency.

If we continue pulling on the thread, wealth might be thought of as an account from which we can spend to hold entropy at bay.

  • A 72-degree house in Scottsdale.
  • Mango salad in during a Minnesota January.
  • Clean water pumped up the Hollywood Hills.
  • A sterile operating room.
  • Skyscrapers defying gravity.
  • Data centers (I couldn’t resist).

Much of the AI commotion revolves around what it means for humans to directly turn electricity into intelligence. That the product is intelligence which can recursively accelerate knowledge instinctively unsettles us because it violates our sensibilities around balance. Like energy is somehow not being conserved, leading to the type of divergence we associate with chain reactions.

I’m actually struggling to put my finger on the right analogy, which validates another point Smil makes. He argues that at precisely the moment when we have become capable of commanding extraordinary quantities of energy, most of us have become almost completely detached from how any of it happens.

Smil calls this our “comprehension deficit.”

Modernity is a black box. Light comes from a switch, and meat comes from a carton lined with a maxi pad. It’s all so effortless. Students of economics read I, Pencil to appreciate how capitalism’s profit-motive and competition actually lead to complex chains of cooperation. The side effect of the story is a (demoralizing?) inference that the world is so invisibly complicated now that you should feel good in your choice of major because it’s strategic to be a symbol-pusher since details are hopelessly infinite. (Sorry, did I go to far here? I’m an econ major, so it’s kind of like Chappelle using the n-word. Kinda? Nevermind.)

Smil’s punchy take on specialization:

You could meet real Renaissance men on Florence’s Piazza Signoria in 1500, but not for too long after that…By the middle of the 18th century, Diderot and d’Alembert could still assemble a group capable of summing up much of their era’s knowledge in the Encyclopédie…In 1872, a century after the appearance of the last volume of the Encyclopédie, any collection of knowledge had to resort to the superficial treatment of a rapidly expanding range of topics…today, it is impossible to sum up our understanding even within narrowly circumscribed specialties…experts in particle physics would find it very hard to understand even the first page of a new research paper in viral immunology …Highly specialized branches of modern science have become so arcane” that many practitioners must train into their thirties “in order to join the new priesthood.

All I’ve done is read the introduction to the book, but the “comprehension deficit”, which he’s fixin’ to remedy with respect to how the world works (at least from a physics/chemistry macro point of view), is just a fascinating idea of its own. Before the printing process and hyper-connectedness, it was hard to learn much of what was a relatively small body of knowledge, but since then, even though you could learn more in a lifetime, it was a smaller percent of what could be known.

But that I can sit here in an air-conditioned house on a 90-degree day typing contemplations about abstractions like how wealth is really the temporary abatement of nature’s randomness and how the accounting of such wealth should be denominated in units of work, which is traditionally measured by money and its continuously leaking exchange rate to said work is itself a wild collective achievement fueled by eons of carbon remnants from lives I never knew.

I guess I feel like I owe it to the universe to be curious about its ways even if it’s a Sisyphean fantasy to close the comprehension deficit. Looking forward to continuing:

Article content

A thought worth sharing more widely from hemispheres this week’s post about discretionary trading:

In this AI era, brains and consciousness are major topics du jour. I increasingly see us as technology centaurs with both a meat and silicon brain mediating structured and unstructured data, which in turn can be private or public.

On the private side, I have Claude connected to:

  • my personal knowledge management or PKM (Notion)
  • MCP to moontower.ai data
  • git repos across work and personal
  • email (which means it can access all my writing)
  • project management software (Linear)
  • Google drive
  • Google calendar

And of course, the entire world of public info lives in the model’s training data and ability to go online.

You can be taking a walk with voice mode on spec’ing a prototype that will connect to data with nearly any context you want to provide from your private knowledge, the internet, or what’s in your head that moment.

The future is here for anyone who feels haunted by dozens of ideas a day that they previously couldn’t or wouldn’t act on.

And a discretionary trader is nobody if not someone who starts 50 sentences a day with “I wonder what happens if…”

Related:

Article content

Money Angle

I ask for your indulgence, for what follows is an informal word-wall connecting conversations I’ve had with some W2 friends.

Pick your favorite gospel on disruption. Schumpeter’s creative destruction, Christensen’s innovator’s dilemma. Most companies, sometimes even industries, will eventually recognize themselves as a melting ice cube.

In a great talk back in 2009, media investor Peter Chernin warned:

One of the things I always used to say to the people who work for me is that you can’t protect your business, and your job is not to protect your business. Your job isn’t to protect. Your job is to maximize your business at any given time, but your real job is to grow new businesses faster than the old ones decay.

He uses record labels as the cautionary example. They set out to protect their business, but “to the degree you’re going to try and do it you’re going to get killed because technology is going to liberate audiences in such a way that they’re going to get what they want regardless.”

It feels like there are only bad choices for companies that are now run-off businesses. Re-investment feels too speculative, but panic will preclude any chance of a soft or at least dignified landing. Navigating these moments well is a sign of grace, but if grace were easy we’d call it something else.

Instead we get to witness corporate pathology.

The ice cube is melting. But it started as a very big ice cube. Big enough to bridge a middle-aged middle-manager’s soft landing into an early retirement if they can claim a larger share of a shrinking cube. Welcome to corporate hunger games on steroids.

The middle manager is a risk-averse incrementalist. That’s WHY they are middle managers. They will do no such thing as grow new businesses faster than the old ones decay. They need an accomplice to secure their spot on the life raft. This accomplice comes from an unexpected place. The partners.

You see, entrepreneurs have deranged brain chemistry. My friend Jason Buck and I were drinking coffee in my backyard on Wednesday when he described an entrepreneur as someone who works 100 hours a week for themselves to not have to work an hour for someone else.

The entrepreneur, the founder, the partner is a delusional optimist. A melting ice cube can be refrozen if you just find the right segment to sell to, the one that somehow eluded you when things were going well. The middle manager latches on to this hope.

He crafts a business plan to do things “radically” differently. Of course, radical in this context is a mere gesture in comparison to the extinction-level shift in the business environment. The partners, never ones to back down from a fight, support this can-do attitude from the manager who stepped up.

The manager, energized and enabled, moves to consolidate influence. By promoting their plan, any plan, doomed as it is, they portray unsupportive colleagues as quitters by virtue of their dissent.

“If you’re not on board with my [dumb] plan, you clearly hate the company and are not a team player.”

There’s no serious appraisal of the dissenting argument’s merit. And that’s probably because the dissenting argument is “Umm, we’re f’d, so we should focus on retention, not growth, because [insert analysis that actually makes sense]”. Any analysis that maximizes EV in a losing game will be hard to sell against a bad strategy employed with hope. We gamble to get back to even and we’re risk-averse when we’re ahead.

Our middle-manager doesn’t just knife out the dissenters. They fully larp optimism with expansion. They hire loyalists, veterans of this obsolete-but-not-yet-obviously-so strategy. The loyalists are relieved. They, themselves, were cooked seeing the same writing-on-the-wall, but they just got thrown a lifeline by the last firm running the old plays. The only plays they’re familiar with. The reunion with their middle-manager buddy is as predictable as you expect, complete with brown liquor and nostalgia for those nights during training when they were chasing tail in Murray Hill and scarfing khati roll together if they struck out that night. Ah, Indian food before bed. To be young again.

The loyalist cluster is pragmatic. Every year of health insurance and private school tuition extends the polyester harmony that is their home life. But for the middle manager, the loyalist cluster is strategic. He’s like an old piece of tile covering himself in linoleum. It’s all wrong but a bigger nuisance to remove. Become hard to kill by entangling yourself deeper in a hierarchy that you created and from which you are a convenient buffer between partners and minions whose names they never want to know.

The misalignment, politics, and the waste of human life force in the name of self-preservation of an artificial environment (that’s a big one to unpack, but y’all might have to come have coffee or maybe something a bit stiffer for that convo) are regrettable.

But you know what I find worse?

That our middle-manager protagonist was not as cunning as he seems. That his primary offense was stupidity. That he sincerely mistook the situation for something he could fix instead of a noble impulse towards calculated self-preservation.

Stupidity is uncivilized. It’s unpredictable. “Say what you want about the tenets of national socialism, at least it’s an ethos. These guys are nihilists.” That’s how I feel about our manager if he is stupid. That’s a barbarian. It might be unpleasant, but you can reason with the merely deceptive.

I’m not sure if our middle-manager character (whose likeness to any real-life individual is purely coincidental) is sincere and stupid or just trying to survive, but those accomplices on high, blinded by desperate optimism, channeled the old guy at the club instead of having the courage to reinvent or the grace to just go home.

Then again, being who they are is what got them to giant ice cube status in the first place.

Money Angle For Masochists

A recommendation

Besides writing, part of my role here is to find good stuff to share with you.

One of my favorite online discoveries this year has been the X account of @SowingAlphaSeed. It’s an anonymous account that goes by “Farmer”.

The bio says:

Retail investor w/ a stretch goal of 2 Sharpe + 30% CAGR. Please DM me if you have fund recommendations. Live portfolio and track record on my website.

The main draw is that this account has steadily been learning and doing “in public”. It’s a constant source of interesting and smart ideas. It’s also focused on what it says in the bio, so the percentage of useful posts to total posts is extremely high.

Tip of the hat to Farmer (who I never met and don’t know).


An observation:

1-year collars in MU are close to the cheapest they’ve in the past year. If you’re bullish you can budget a trade that can offer multiples of return per dollar bet. If MU already made you rich, the cost to hedge is actuarially small.

Then again, you didn’t get rich by hedging amirite

This week:

Article content

Zooming in:

Article content
moontower.ai
Article content
moontower.ai

Moontower Weekly Recap

the “middle manager”

I ask for your indulgence, for what follows is an informal word-wall connecting conversations I’ve had with some W2 friends.

Pick your favorite gospel on disruption. Schumpeter’s creative destruction, Christensen’s innovator’s dilemma. Most companies, sometimes even industries, will eventually recognize themselves as a melting ice cube.

In a great talk back in 2009, media investor Peter Chernin warned:

One of the things I always used to say to the people who work for me is that you can’t protect your business, and your job is not to protect your business. Your job isn’t to protect. Your job is to maximize your business at any given time, but your real job is to grow new businesses faster than the old ones decay.

He uses record labels as the cautionary example. They set out to protect their business, but “to the degree you’re going to try and do it you’re going to get killed because technology is going to liberate audiences in such a way that they’re going to get what they want regardless.”

It feels like there are only bad choices for companies that are now run-off businesses. Re-investment feels too speculative, but panic will preclude any chance of a soft or at least dignified landing. Navigating these moments well is a sign of grace, but if grace were easy we’d call it something else.

Instead we get to witness corporate pathology.

The ice cube is melting. But it started as a very big ice cube. Big enough to bridge a middle-aged middle-manager’s soft landing into an early retirement if they can claim a larger share of a shrinking cube. Welcome to corporate hunger games on steroids.

The middle manager is a risk-averse incrementalist. That’s WHY they are middle managers. They will do no such thing as grow new businesses faster than the old ones decay. They need an accomplice to secure their spot on the life raft. This accomplice comes from an unexpected place. The partners.

You see, entrepreneurs have deranged brain chemistry. My friend Jason Buck and I were drinking coffee in my backyard on Wednesday when he described an entrepreneur as someone who works 100 hours a week for themselves to not have to work an hour for someone else.

The entrepreneur, the founder, the partner is a delusional optimist. A melting ice cube can be refrozen if you just find the right segment to sell to, the one that somehow eluded you when things were going well. The middle manager latches on to this hope.

He crafts a business plan to do things “radically” differently. Of course, radical in this context is a mere gesture in comparison to the extinction-level shift in the business environment. The partners, never ones to back down from a fight, support this can-do attitude from the manager who stepped up.

The manager, energized and enabled, moves to consolidate influence. By promoting their plan, any plan, doomed as it is, they portray unsupportive colleagues as quitters by virtue of their dissent.

“If you’re not on board with my [dumb] plan, you clearly hate the company and are not a team player.”

There’s no serious appraisal of the dissenting argument’s merit. And that’s probably because the dissenting argument is “Umm, we’re f’d, so we should focus on retention, not growth, because [insert analysis that actually makes sense]”. Any analysis that maximizes EV in a losing game will be hard to sell against a bad strategy employed with hope. We gamble to get back to even and we’re risk-averse when we’re ahead.

Our middle-manager doesn’t just knife out the dissenters. They fully larp optimism with expansion. They hire loyalists, veterans of this obsolete-but-not-yet-obviously-so strategy. The loyalists are relieved. They, themselves, were cooked seeing the same writing-on-the-wall, but they just got thrown a lifeline by the last firm running the old plays. The only plays they’re familiar with. The reunion with their middle-manager buddy is as predictable as you expect, complete with brown liquor and nostalgia for those nights during training when they were chasing tail in Murray Hill and scarfing khati roll together if they struck out that night. Ah, Indian food before bed. To be young again.

The loyalist cluster is pragmatic. Every year of health insurance and private school tuition extends the polyester harmony that is their home life. But for the middle manager, the loyalist cluster is strategic. He’s like an old piece of tile covering himself in linoleum. It’s all wrong but a bigger nuisance to remove. Become hard to kill by entangling yourself deeper in a hierarchy that you created and from which you are a convenient buffer between partners and minions whose names they never want to know.

The misalignment, politics, and the waste of human life force in the name of self-preservation of an artificial environment (that’s a big one to unpack, but y’all might have to come have coffee or maybe something a bit stiffer for that convo) are regrettable.

But you know what I find worse?

That our middle-manager protagonist was not as cunning as he seems. That his primary offense was stupidity. That he sincerely mistook the situation for something he could fix instead of a noble impulse towards calculated self-preservation.

Stupidity is uncivilized. It’s unpredictable. “Say what you want about the tenets of national socialism, at least it’s an ethos. These guys are nihilists.” That’s how I feel about our manager if he is stupid. That’s a barbarian. It might be unpleasant, but you can reason with the merely deceptive.

I’m not sure if our middle-manager character (whose likeness to any real-life individual is purely coincidental) is sincere and stupid or just trying to survive, but those accomplices on high, blinded by desperate optimism, channeled the old guy at the club instead of having the courage to reinvent or the grace to just go home.

Then again, being who they are is what got them to giant ice cube status in the first place.

energy is the ultimate currency

A year ago I wrote a piece called Capitalism Is a Temporary Condition. Buried in it was a thought:

An investor in an age of acceleration watches as commodity prices, real things, hover in familiar territory while “future cash flows discounted” reward attention — in some cases because of optimistic stories, in some cases because the company exchanges cash for BTC (an act which adds no economic value), and in some cases for nostalgic lolz.

Today, it’s embarrassing to think of actual value. The most basic form is stored energy**. An obvious example is oil. Slightly more abstract —a building is a collection of atoms that took energy to arrange and provides utility. It’s proof of work. Same with a good reputation.

The ** footnote:

I increasingly think that a durable concept of a risk-free or least-risky rate is more likely to come from Vaclav Smil than U Chicago.

I knew about Smil from a profile piece that celebrated his suffer-no-fools rigor in making sense of the world.

I’m finally getting around to How the World Really Works, and I dig how quickly it’s zeroing in on what I anticipated to be the biggest muscle movement of all. I’ll let it unfold as he does the book’s intro.

Smil asks us to imagine an alien probe watching Earth, programmed to wake up whenever something important changes in the way energy is captured or used.

For a very long time, it sleeps.

Then life figures out photosynthesis: sunlight can be captured and stored as chemical energy. Much later, humans learn to control fire. Then agriculture. Draft animals. Wind and water.

The gaps are enormous. Billions of years. Millions. Thousands.

Then they start collapsing.

Coal gives humans access to immense stores of ancient solar energy. You can see a parallel to wealth in that phrase alone. Steam will turn that energy into mechanical work, replacing what he calls prime mover energy (human and animal labor). Oil makes enormous amounts of energy portable while electricity makes it transmissible and almost infinitely adaptable.

Smil puts numbers on this.

In 1800, the world had access to roughly 0.05 gigajoules of useful energy per person each year. By 1900 it was 2.7. By 1950, about 10. By 2000, 28. By 2020, roughly 34.

In a little over two centuries, useful energy per person increased nearly 700-fold.

The average person today commands roughly the energy equivalent of 60 adults working continuously, day and night. In affluent countries, the equivalent is closer to 200–240 permanent laborers.

And the averages hide enormous differences.

Measured in primary energy, someone in one of the world’s poorest countries may use less than 10 GJ a year. India is around 20–30. The global average is roughly 75–80. China is above 100. Western Europe and Japan generally sit around 120–160. The United States is closer to 250–280.

It sounds a lot like how GDP per capita is distributed. We measure GDP in dollars, but money can be printed, thus distorting its exchange rate versus the stored value of work. In that sense, energy may be civilization’s only uncorruptible currency.

If we continue pulling on the thread, wealth might be thought of as an account from which we can spend to hold entropy at bay.

  • A 72-degree house in Scottsdale.
  • Mango salad in during a Minnesota January.
  • Clean water pumped up the Hollywood Hills.
  • A sterile operating room.
  • Skyscrapers defying gravity.
  • Data centers (I couldn’t resist).

Much of the AI commotion revolves around what it means for humans to directly turn electricity into intelligence. That the product is intelligence which can recursively accelerate knowledge instinctively unsettles us because it violates our sensibilities around balance. Like energy is somehow not being conserved, leading to the type of divergence we associate with chain reactions.

I’m actually struggling to put my finger on the right analogy, which validates another point Smil makes. He argues that at precisely the moment when we have become capable of commanding extraordinary quantities of energy, most of us have become almost completely detached from how any of it happens.

Smil calls this our “comprehension deficit.”

Modernity is a black box. Light comes from a switch, and meat comes from a carton lined with a maxi pad. It’s all so effortless. Students of economics read I, Pencil to appreciate how capitalism’s profit-motive and competition actually lead to complex chains of cooperation. The side effect of the story is a (demoralizing?) inference that the world is so invisibly complicated now that you should feel good in your choice of major because it’s strategic to be a symbol-pusher since details are hopelessly infinite. (Sorry, did I go to far here? I’m an econ major, so it’s kind of like Chappelle using the n-word. Kinda? Nevermind.)

Smil’s punchy take on specialization:

You could meet real Renaissance men on Florence’s Piazza Signoria in 1500, but not for too long after that…By the middle of the 18th century, Diderot and d’Alembert could still assemble a group capable of summing up much of their era’s knowledge in the Encyclopédie…In 1872, a century after the appearance of the last volume of the Encyclopédie, any collection of knowledge had to resort to the superficial treatment of a rapidly expanding range of topics…today, it is impossible to sum up our understanding even within narrowly circumscribed specialties…experts in particle physics would find it very hard to understand even the first page of a new research paper in viral immunology …Highly specialized branches of modern science have become so arcane” that many practitioners must train into their thirties “in order to join the new priesthood.

All I’ve done is read the introduction to the book, but the “comprehension deficit”, which he’s fixin’ to remedy with respect to how the world works (at least from a physics/chemistry macro point of view), is just a fascinating idea of its own. Before the printing process and hyper-connectedness, it was hard to learn much of what was a relatively small body of knowledge, but since then, even though you could learn more in a lifetime, it was a smaller percent of what could be known.

But that I can sit here in an air-conditioned house on a 90-degree day typing contemplations about abstractions like how wealth is really the temporary abatement of nature’s randomness and how the accounting of such wealth should be denominated in units of work, which is traditionally measured by money and its continuously leaking exchange rate to said work is itself a wild collective achievement fueled by eons of carbon remnants from lives I never knew.

I guess I feel like I owe it to the universe to be curious about its ways even if it’s a Sisyphean fantasy to close the comprehension deficit. Looking forward to continuing:

a common marginal thinking blindspot

In Chapter III of The Theory of Political Economy, William Jevons describes the economic notion of marginal value in 1871.

Nor, when we consider the matter closely, can we say that all portions of the same commodity possess equal utility. Water, for instance, may be roughly described as the most useful of all substances. A quart of water per day has the high utility of saving a person from dying in a most distressing manner. Several gallons a day may possess much utility for such purposes as cooking and washing; but after an adequate supply is secured for these uses, any additional quantity is a matter of comparative indifference. All that we can say, then, is, that water, up to a certain quantity, is indispensable; that further quantities will have various degrees of utility; but that beyond a certain quantity the utility sinks gradually to zero; it may even become negative, that is to say, further supplies of the same substance may become inconvenient and hurtful…

The final degree of utility is that function upon which the Theory of Economics will be found to turn. Economists, generally speaking, have failed to discriminate between this function and the total utility, and from this confusion has arisen much perplexity. Many commodities which are most useful to us are esteemed and desired but little. We cannot live without water, and yet in ordinary circumstances we set no value on it. Why is this? Simply because we usually have so much of it that its final degree of utility is reduced nearly to zero.

Examples of the margin setting the price are all around us. The price of power is set by the last unit of demand, which necessitates the use of “peakers” or generation with the most expensive fuel inputs. The most optimistic bidder sets the price in a classic auction.

Heck, just imagine a neighborhood where everyone had to recertify the price of their home by being willing to pay the current market value. I live in CA. Most people living here would not and could not pay the current marginal price to live here. That many of those same won’t sell is not irrational; it’s called making a wide market. You can drive a truck through their bid/ask, but so what? The point is that marginal utility is setting the price, not the average. It makes sense. Real estate is an auction.

The idea that the margin sets the price is obvious when it’s pointed out. But it’s also obvious that people haven’t internalized it.

[This is an example of a general phenomenon in which deriving an answer is much harder than verifying whether it is correct. It’s harder to compute the cube root of 729 than to verify that 5 is wrong or 9 is right. This is the same principle behind crypto mining. It’s computationally expensive to find a specific nonce, but easy to confirm whether a solution satisfies the requirements.]

Your uncle, your cousin, your brother-in-law will, with a straight face, tell you they know more about a stock than the average person as if the average person’s opinion matters when it comes to making money. The consensus price is set by well-resourced, well-connected wiseguys. If you can beat them, you’re rich. If you can do it reliably and on large sums, you’re on your way to becoming one of them, but the standard for feeling like a winner at trading is nothing short of being able to consistently beat the marginal price. It’s your record against the spread.

Marginal thinking blind spot

People often take up my offer for paid calls because they want to bounce a pitch off me. They want a sparring partner, criticism, feedback, and tips. Nobody would say they were trying to persuade me, but it’s only human to assume that they would feel greatly encouraged if they did. But the best-case scenario on that front is a stalemate. The shape of everything that works, at least in public markets trading, is nuanced and requires handling many details at a high level. There’s no magic bean machine where someone is gonna make me say “take all my money”.

[If the idea is a massive layup, it’s resting on a special, likely fleeting access. I’d even shove high-barrier-to-entry infrastructure in this category even if the details vary. In all these cases, the provider of the strategy has a bargaining position that makes it a layup for them but not the capital provider whose results will be haircutted by the cost of the bargain.]

Many arriving with a pitch aren’t even getting to the stalemate. They are confusing their extensive knowledge, knowledge far better than average, with sufficient-to-be-competitive knowledge. To be able to trade, they must have a strong concept of marginal price, yet there’s a blind spot when the concept is applied to themselves.

You can’t go wrong by drilling the concept of value over replacement into your brain. It should be part of your mental OS to apply that lens to all of your major decisions. If you think of replacement value as a strike price, VOR prompts you to think of the moneyness of the strike and the possible upside above it. It goes hand-in-hand with the economic concept of opportunity cost as the strike of a put you’re giving up.

Ask yourself

Back to markets for a moment.

2 key questions to ask yourself:

Can the sharpest players here warehouse this entire risk?

If so, the price already reflects their view, and you have to beat them to make money.

If not, is it because they don’t know about it (maybe it’s very small capacity, but even then, you could reframe as “the sharpest players who would know about this”). Or maybe, it’s just because there’s a general “wrong way” risk premium. SP500 puts are overpriced since there’s no hedge to the economy actually imploding. All such hedges, at the limit, are constrained by general credit and solvency, making a short SP500 put trade an ever-present and therefore uninteresting trade, just as being long passive indices is underwriting an ever-present risk premia (notwithstanding any comparisons with implied forward returns with real “risk-free” rates).

Are you on par with the sharpest players who are setting the price?

This second question is a prompt to think strategically about your role in the system, not an invitation to consider your IQ (which in the range-restricted sample of “competitive investors” is uncorrelated with performance and perhaps even anti-correlated in a Berkson Paradox-style inversion). The prompt should make you consider who competes at your chosen horizon, what constraints you have versus them, and how incentives shape both the chosen horizon and constraints.

Finally, a meta-consideration that sits above these questions.

Notice how comparative trading is. It is inextricable from VOR, competition, and the zero-sumness of alpha.

The entire decision to care about this is a container that all your thoughts will need to live in. This isn’t to say that trading is uniquely competitive, but like sport it’s defined by it. Life is always competition in some abstract sense, but in trading you’re a crab in a bucket. An artist competes by being N of 1 and achieves that by competing with themselves. I’m not saying this is easier or harder, I have no such authority on which to even speculate, but it is different.

I’ll stop well short of lionizing infinite games over finite ones. We have both and I’d rather people just fit where they flourish rather than serve some high-minded ideal (the statement of “I’d rather” is itself some high-minded ideal, but I have no idea how to hit ctrl-break on the recursion so take it easy on me). I only make the distinction between game types in case it strikes a chord for someone debating what to focus on.

futures premium cell

Back in my SIG days, every trader had a cell on their spreadsheet showing the SPUs (the name for SPX futures back in the day; it’s derived from the September symbol) premium or discount to the “cash”.

A little background

The cash is the SPX index price based on its components’ spot prices. You can compute the index value from the stock bids, and that would be the “SPX cash bid”. You can do this for the offer as well. The average of the bid and offer is SPX mid-market.

Futures contracts, based on no-arbitrage pricing, have a fair value based on what you gain or give up by owning the future instead of the cash basket. Since you save the interest on the cash it would cost to own the basket but forgo any dividends you would have received that offset the drop in the shares when they are paid the futures are valued as

SPX + (interest – dividends)

The interest and dividends are estimated from today until the futures’ expiry date.

Assume:

SPX cash mid market: 7,800

RFR = 3.5%

annual dividend yield =1.5%

t = 1 year

The 1-year future fair value is approximately 7800e(3.5% – 1.5%)*1 = 7957.57

The basis between the future and the cash index was known as the EFP (“exchange for physical”). In our example, it equals 157.57

If the future was trading 7977.57, we’d say “the futures are trading over”. In this case, it would be 20 points or about 25 bps rich as a percent of the cash index.

This is an arbitrageable difference as a trading firm can sell the future, buy all the components, and book a theoretical profit that will become a real profit if their carry assumptions hold true. As a market-maker on the floor, I was not involved in index arbitrage this directly (although most of my clerking experience was much closer to these strategies).

S&P Futures and Fair Value. | The Blue Collar Investor
You have probably seen this kind of graphic on CNBC

The futures are more liquid than the cash index so the basket price lags. If systematic bullish news hits the tape, you cross the tiny bid/ask on the futures. In fact, the ES or e-mini future actually leads the big SPUs, so ES was used for computing the premium or discount.

All of the market-makers had a cell on the spreadsheet on their handheld tablets that showed the futures’ premium/discount to fair value (FV = cash index + EFP). Those premiums or discounts were typically small, 10 bps or less, as the market bounced around. But huge news could catapult the futures 300 bps before much of the basket could blink. As you can imagine, index arbs get very uneven fills on the baskets they try to execute to close the gap. That’s because everyone making markets has not only pulled their stock offers, but lifting resting customer offers, often several levels through the NBBO that was posted a split second ago.

You can beta-weight the “amount over” or premium the futures are trading to estimate a new fair value for the stock. If the stock you trade has a 1.2 beta, then you might think its new fair value is 3.6% higher than the pre-news price if the futures are trading at a 3% premium. Of course, beta is just a statistical quantity so you have a confidence interval around the beta, which is pretty much life as a market maker. Futures are trading 300 bps over, what’s your 2-way on XYZ stock? Maybe I’m “2% over” bid, offered “4% over”. Then, you’d need to be quick to remember what bids and offers on the option chain are resting that you should race to lift (all the calls should increase by their delta * your beta-weighted change in fair value) and hit (all the puts will decrease by the same factor and this is not even considering the effect of gamma).

Today, all the quotes being streamed by traders will get pulled while they simultaneously blast all the resting orders in a flash, but this process happened a bit more at human speed 30 years ago (although you still weren’t gonna beat the floor…the competition for the resting orders would be market-maker on market-maker).

I’m mostly sharing this because it’s fun for the stragglers who read this far into the masochists section, but it’s worth mentioning the context that made me think of it.

This phenomenon sometimes leads to anti-data or an interpretation of data that is exactly the opposite of the typical inference.

Why?

If a contract or security trades on the offer, the assumption is generally that “paper” (trader language for customers) is buying and the market-maker is selling. But in the example above, the market-makers are the aggressors because they know the market is much higher than last sale. They are lifting calls and hitting puts, so your read on flow sentiment is exactly the opposite of a normal market condition.

Just another example of reality having a surprising amount of detail.

Moontower #325

In this issue:

  • for when the world slows down
  • family investing webinar
  • futures premium cell

Friends,

I’m in Vegas with friends to see GnR this weekend. 36 hours of noise bookended by 1-hour flights of peaceful purification in the form of flight reading.

Most of what I read on the regular is work-related and instrumental. I save appreciative reading for when the world slows down. Flights are a reliable example of such a time.

Here is both a short article and a long one for when you’re in the mood to actually read every word.

The Real Meaning of the “Road Less Taken” | 2 min read

I love this message and agree with the author about it being the one we need. It’s also the message I closed last year’s trading bootcamp talk with, so I’m biased.

Letter from Wendell: How a poet and farmer and a guitarist and teacher connected across time and space 36 min read

A friend of my wife who reads my stuff sent her this to pass along to me. There is no summarizing an article like this. It would completely miss the point. That’s why I recommend reading it when you are in the mood to read every word. The friend has a good read on me (or my aspirational self anyway).

 


Money Angle

This summer I scheduled a lab to help kids in the Investment Beginnings Class buy their first portfolios.

I scheduled the lab during market hours so I could help with the execution. Between vacations and camps, a lot of kids were unable to make the lab, so it turned into rolling office hours. I ended up doing 8 sessions, sometimes just with one student on Zoom (one of the kids called from sleepaway camp!)

Parents sat in on some of these. Word spread, so I ended up getting a bunch of students who didn’t come to the 5-week course, but the parents are friends, so I put together a super abbreviated version of the course.

Actually, it wasn’t really a review because that’s hard to have people engage with if they never went through the course in the first place. It evolved into a spiel that overlapped with key ideas from the course. I’ve done it a bunch of times now so it’s more like a 2-hour seminar.

The whole experience makes me wanna try something.

Let’s do a “family investing webinar” where a kid and their parent(s)/guardian(s) join a Zoom call that I host.

2 hours in one shot is too long. We’ll break it up into 2 1-hour sessions.

There won’t be any charge for it, but I will limit it to paid subs to this newsletter.

 

Money Angle For Masochists

First, a heads-up that Thursday’s using tick data for spot-vol correlations was enough masochism that I was told to back off:

On to today’s masochism…

The futures-premium cell

Back in my SIG days, every trader had a cell on their spreadsheet showing the SPUs (the name for SPX futures back in the day; it’s derived from the September symbol) premium or discount to the “cash”.

A little background

The cash is the SPX index price based on its components’ spot prices. You can compute the index value from the stock bids, and that would be the “SPX cash bid”. You can do this for the offer as well. The average of the bid and offer is SPX mid-market.

Futures contracts, based on no-arbitrage pricing, have a fair value based on what you gain or give up by owning the future instead of the cash basket. Since you save the interest on the cash it would cost to own the basket but forgo any dividends you would have received that offset the drop in the shares when they are paid the futures are valued as

SPX + (interest – dividends)

The interest and dividends are estimated from today until the futures’ expiry date.

Assume:

SPX cash mid market: 7,800

RFR = 3.5%

annual dividend yield =1.5%

t = 1 year

The 1-year future fair value is approximately 7800e(3.5% – 1.5%)*1 = 7957.57

The basis between the future and the cash index was known as the EFP (“exchange for physical”). In our example, it equals 157.57

If the future was trading 7977.57, we’d say “the futures are trading over”. In this case, it would be 20 points or about 25 bps rich as a percent of the cash index.

This is an arbitrageable difference as a trading firm can sell the future, buy all the components, and book a theoretical profit that will become a real profit if their carry assumptions hold true. As a market-maker on the floor, I was not involved in index arbitrage this directly (although most of my clerking experience was much closer to these strategies).

S&P Futures and Fair Value. | The Blue Collar Investor
You have probably seen this kind of graphic on CNBC

The futures are more liquid than the cash index so the basket price lags. If systematic bullish news hits the tape, you cross the tiny bid/ask on the futures. In fact, the ES or e-mini future actually leads the big SPUs, so ES was used for computing the premium or discount.

All of the market-makers had a cell on the spreadsheet on their handheld tablets that showed the futures’ premium/discount to fair value (FV = cash index + EFP). Those premiums or discounts were typically small, 10 bps or less, as the market bounced around. But huge news could catapult the futures 300 bps before much of the basket could blink. As you can imagine, index arbs get very uneven fills on the baskets they try to execute to close the gap. That’s because everyone making markets has not only pulled their stock offers, but lifting resting customer offers, often several levels through the NBBO that was posted a split second ago.

You can beta-weight the “amount over” or premium the futures are trading to estimate a new fair value for the stock. If the stock you trade has a 1.2 beta, then you might think its new fair value is 3.6% higher than the pre-news price if the futures are trading at a 3% premium. Of course, beta is just a statistical quantity so you have a confidence interval around the beta, which is pretty much life as a market maker. Futures are trading 300 bps over, what’s your 2-way on XYZ stock? Maybe I’m “2% over” bid, offered “4% over”. Then, you’d need to be quick to remember what bids and offers on the option chain are resting that you should race to lift (all the calls should increase by their delta * your beta-weighted change in fair value) and hit (all the puts will decrease by the same factor and this is not even considering the effect of gamma).

Today, all the quotes being streamed by traders will get pulled while they simultaneously blast all the resting orders in a flash, but this process happened a bit more at human speed 30 years ago (although you still weren’t gonna beat the floor…the competition for the resting orders would be market-maker on market-maker).

I’m mostly sharing this because it’s fun for the stragglers who read this far into the masochists section, but it’s worth mentioning the context that made me think of it.

This phenomenon sometimes leads to anti-data or an interpretation of data that is exactly the opposite of the typical inference.

Why?

If a contract or security trades on the offer, the assumption is generally that “paper” (trader language for customers) is buying and the market-maker is selling. But in the example above, the market-makers are the aggressors because they know the market is much higher than last sale. They are lifting calls and hitting puts, so your read on flow sentiment is exactly the opposite of a normal market condition.

Just another example of reality having a surprising amount of detail.

 


Moontower Weekly Recap

“why should I learn this?”

I usually have a concrete plan in advance of writing, but today’s letter is totally spontaneous, other than knowing that something would be published. I wrote it all, but found myself stuck on a title. I hope it will make sense by the time you’re done.

So…one of my oldest friends, someone I consider family really, is enjoying a sabbatical year. He and his wife crashed with us for 10 days. It’s one of those slices of time that you know before it’s even over that you will have nostalgia for.

We didn’t do anything overly special, although it was a great catalyst to convene with the rest of our Bay Area college crew over the weekend. The week was something in between a staycation and just a far more elevated (ie joyful) routine. I would work during the day while they went on an excursion, but we’d make sure to go to the gym or at very least a walk together daily, and the evenings were filled with good food (his baked ziti is in contention for my electric chair meal which I don’t anticipate needing unless they start rounding up the dorks) and games (Scrabble and Decrypto mostly) with the whole family.

One of my favorite parts of it was to have them around in such an informal way, just like family crashing. The kids would hang around and be part of the discussion like these were the aunties and uncles they normally see. Everyone actually gets to know each other. My kids get to see that their dad’s friends are weird just like their dad is. It’s funny for them to see them process it, but I hope it models to them what friendship is, even if they do think we’re old aliens. I’m also happy to see my friends know my kids. To see their different personalities and for them to be more than just names.

Personally, the week has been so much fun, especially because my friend and I have always had a nerd bond. He’s much more educated and technical than I (this sabbatical very likely ends with him working for Jane Street or Waymo), so I also just have permission to geek out without worrying that I’ll be talking to myself soon. I also convinced him, although it didn’t require much, to show me the game Factorio. This confirmed what I expected…I can’t introduce THAT into my life. Diet soda is addicting enough.

A nice bonus feature of the week was getting to chat about various math topics and education broadly because we were both trying to help Zak with his Math Academy lessons by trying to break down the concepts in intuitive ways. He’s currently jamming on logs and exponentials, which I’ll come back to in a moment. But I want to share a nice analogy first. When Zak hits a wall, he’ll do that thing all kids do when they feel frustrated. “Why do I need to learn this, I’m never going to need it.”

A reflexive, true, and entirely uncompelling response to such pleas is “Actually, you might. It depends what you end up doing for a living.” But kids think the future is as distant as the afterlife, so the argument for doing homework is as convincing as telling them they’ll burn in hell for punching their brother. My friend used an analogy that meets Zak on his terms. “Why do you do pushups or lift weights? You’re never going to do a pushup on the court.”

My mother shared this exact point of view when I was growing up. It’s training. Let’s be honest, when it comes to actual application, the most useful classes you take in school are home ec and typing. And while I certainly have many gripes with the non-useful stuff they teach, there’s stuff that you will not use but counts as training like math and critical reading. Numeracy and literacy. Even if their utility were diminished, they make for a richer interior existence, allowing you to be amused and intrigued by the world for free. Anyway, a nerd writing on the internet comes off as one-note at best and self-flattering at worst when carrying a flag for hokey ideas like doing your times tables, so I’ll stop there on all that.

Back to the log stuff real quick. I sometimes wonder if it’s such a challenging topic because our minds struggle naturally with non-linearity or if we should actually just learn it earlier. It seemed helpful for Zak to realize that all logs are is another rung on the ladder of basic math operations.

We start with addition.

Subtraction is the inverse of addition. It “undoes” addition.

We then move up to multiplication. Multiplication is repeated addition. 4 tires, 40 times is the same as 4+4+4…(repeated 37 more times) which is the same as a Nascar race.

Division undoes multiplication by repeating subtraction. 40 divided by 8 is how many times can I take away 8 from 40.

Then we move up to exponents. Exponents are repeated multiplication of the same number.

Logs undo exponents by repeatedly dividing by the same number (ie the base).

In infographic form:

moontower

Zak was struggling with understanding natural log. I took a stab at it from the compounding angle during a car ride last week.

“If you invest $100 and receive back $110 in a year what interest rate did you receive?”

10%

“What if I told you you compounded semi-annually…do you think the interest rate that got you to $110 is greater than or less than 10%?”

Less than.

“Good. Forget the calculator, let’s just guess and say the semiannual compounding at 9.8% got us to $110. What if we compound daily, is the rate greater than or less than 9.8%?”

Less than.

“Now imagine we keep shortening the interval from daily compounding to minute-compouding to seconds to nanoseconds. We can shorten the interval until it gets close to zero without touching zero. Later in calculus you’re learn that this is a useful trick where you approach zero but don’t touch it. It’s called a limit. If you shorten the interval until the limit, almost zero, we call that continuous compounding. The natural log gives you the rate if we assume continuous compounding. So in the case of our investment, we can compute the continuous rate by taking the natural log of 1.1 because our return was $110/$100.

LN(1.1) is about 9.5% going off memory and represents the continuously compouded rate that would give a total return of 10%.”

Then I did that thing he hates which is try to give him more than he asked for.

“You know how if you double your money that’s a 100% return. Well, the continuous compounded rate comes from taking LN(2). Before we compute that, do you think the continuous compound rate is going to be less or greater than 100% if we doubled our money?”

Less than 100%

“Exactly. LN(2) is about 72%. The cool thing about logs is you can simply divide the already compounded return by the number of years to get an annual compounding rate. So if you double your money in 10 years, the annual rate is 7.2%

That’s where the rule of 72 comes from!

It’s just inverting the logic. If you continuously compound at 10% per year it takes ~7 years to have a continuously compounded return of 70%, which corresponds to doubling your money.”


e (Euler’s constant)

We talked just a little about e.

If you continuously compound at 100% for 1 year, you end up with e, or about 2.718x what you started with.

e1 = 2.718

Undo it:

ln(2.7128) = ln(e) = 1 = 100%

If it takes 10 years for your money to grow to 2.718, then you are continuously compounding at 10% or 100% / 10

Contrast this with solving for the annual compounded return where you compute:

2.7181/10 – 1 = 10.5% annual compouding

We just did ln(2.7128)/10 = 10% continuous compounding

Continuous return in finance

2 properties make log returns convenient for financial math.

A) Logreturns are linearly proportional to time making them easier to manipulate.

The wealth equation is:

W = ert where r is the logreturn. Note its proportionality to t.

but for annual compounding:

W = (1+r)t where r is a discretely compounded return.

In discrete land, we need the t root to back out the rate.

2.7181/10 – 1 = 10.5% annual compouding

In continuous terms,

ln(2.718)/10 = 10% continuous compouding

 

Logs linearize the time component of returns so you can aggregate them as time-weighted sums or take simple averages to get the mean logreturn. Discrete compounding requires multiplying returns then taking the geometric mean (t root)

By example, suppose you grow by 10% in year 1, 26 % in year 2, and 0% in year 3. What’s average total growth?

If those are log returns, the average annual log return is simply 12%

If they are annual returns, we must compute:

(1.1 * 1.26 * 1.00 )1/3 – 1 = 11.49%

B) Log returns are normally distributed, making them a convenient measure of return for statistical modeling.

 


Learning Math in the AI Era

A fun thing my friend told me was that the natural log function was discovered and the term coined before e, which is surprising because I think of the natural log as just a logarithm with e as a base. But it had relevance before e!

I needed Claude to spell it out:

The natural log arrived through the hyperbola. Saint-Vincent’s equal-area observation was around 1647, Sarasa spelled out “this area behaves like a logarithm” in 1649, and Mercator published the series and the name logarithmus naturalis in 1668. So by the late 1660s people had the natural log as a working object — they could compute it, they had its series, they knew it was the log that came out of the 1/x area with no choice of base.

What they did not have was e as a recognized number. Nobody in that group pointed at 2.718… and said “there’s a special constant here, and it’s the base of this thing.” Napier’s 1614 construction secretly ran on 1/e, and the hyperbola log secretly had base e, but “secretly” is the operative word.

e gets isolated as a number of interest later and from the side door — Jacob Bernoulli, around 1683, poking at compound interest and the limit of (1 + 1/n)ⁿ. He found the value sat somewhere between 2 and 3 but didn’t tie it cleanly to the logarithm or give it a name. The unification is Euler’s, in the early-to-mid 1700s: he named it e, nailed it as the base of the natural log, established eˣ and ln as inverses, and made the whole structure look inevitable rather than accidental.

Charlie Munger was on a podcast shortly before his death and he spoke of how a solid understanding of grade-school and high school math basics was fundamental to thinking. He had a highly utilitarian perspective rather than an academic one.

If we combine the natural log story with Munger’s perspective, I think we land at an interesting idea. A math history approach to the basics.

In elementary school, the focus should certainly be operations. There’s a grammar to math that complements the many abstractions of counting, which is what I think you’re ultimately learning. But by late middle school, we should include an appeal to stories, history, mystery, and pragmatism by personalizing the context of the math we learn. To put a student in the shoes of someone trying to solve a problem for the first time in history with the tools that were available at the time. Obviously, asking students to do what geniuses did is not the goal. But AI would be an amazing tool for placing the student in an RPG where you drip as much information as they need to get to the next step within an appropriate level of difficulty for the individual.

It wouldn’t be a substitute for instrumental math education but a way to deepen our relationship with the fundamental concepts Munger thinks we could benefit from deepening. And it’s not limited to math. It’s more like STEM History 101. Science ed seems to have a bit more focus on the individual scientists and stories of discovery, but many of these figures are fascinating eccentrics if not crazies. It’s a colorful way to captivate.

I admit it’s less than a half-baked idea, so it’s really more of a “here’s a side-project that could be cool, feel free to run with it” but I do keep coming back to it as something I’d like to see.

I also want to take a moment to repeat myself — AI is a tireless tutor. A gift to the curious.

This investor has been live-tweeting his own learning arc:

It’s a great example of the similar projects I’ve been doing for self-help.

I’ve unpaywalled the below article Socrates 2026: How to Use Highlights.

It’s stuffed to the gills with things I think are fun and can hopefully help you help yourself.

Socrates 2026: how to use highlights

Socrates 2026: how to use highlights

Follows from Part 1: uncovering the laws of nature


Friends,

In Part 1, I teased that Geoffrey West’s Scale was a perfect surface to show how you can learn as you’ve always wanted. Or needed but didn’t know it.

Plan

  1. Cover how I used LLM to self-teach, which you can use for learning or re-learning anything.
  2. Cement and practice our understanding of power functions

🧠For those familiar with learning science, you will recognize several techniques, but I’ll label them as they appear.

How this all started

When I read a physical book, I will usually take a screenshot and then OCR the page to keep a digital excerpt. This is ok if there aren’t a lot of excerpts or highlights to preserve. I quickly realized I was going to need the Kindle version of Scale as the highlights and their accompanying inconvenience were piling up fast. I snagged it on Libby (this is your library’s digital loan service). There was no waitlist. Yet another reminder that there is so much joy available for free.

We must talk about highlighting.

The naive understanding of highlighting is that by taking the effort to trace a 25% opacity yellow film over words, you have learned something. By now you know this is item #77 on the list of self-deceptions. Still we carry on because it’s a cheap option. Somewhere in the recesses of your dopamine-addled mind (remember dopamine is the “seeking” chemical), you expect the highlights you stash like old coax cables will find fresh life when recombined with technology. Don’t be hard on yourself, an impotent but aspirational habit ranks less than wisdom but higher than apathy.

But it turns out this self-deception call option may finally have a payoff.

As I was marking up Scale, I had no guilt about not internalizing what I was reading in the moment. My plan was to export all the highlights like I usually do.

⚠️Kindle formats usually limit your exports to 10% of the book for copyright reasons. It’s a bit messy, but once I think I’m about there, I export those highlights to a file, then delete them in the Kindle app. For Scale, this happened when I was about half finished with the book. I could then start highlighting from zero for the second half.

This time, instead of just storing them, I was going to give them to a Claude project to seeding a “curriculum” so I could learn in a way that only comes from practice, not recitation. Reviewing your notes/highlights lets you cram for a test, but it’s not the kind of learning you can call on for invention.

Socrates 2026

Let’s rewind for a moment.

A few months back I bombed a Jane Street interview question I found online. This wouldn’t normally bother me as I’m far past the time in my life where my self-esteem teeters on an illusion of cleverness. But it was a question I felt I should know how to answer as opposed to the corpus of questions from which I wouldn’t even know where to begin.

You can see my write-up about it here: turning a Jane Street interview question blunder to a lesson

I realized that I couldn’t answer the question because I didn’t fully appreciate that variance is the spread between the expectation of a square and the square of an expectation.

Adjacent thoughts

  1. That variance is always non-negative is a demonstration of Jensen’s inequality operating on a function that takes the sum of squared deviations.
  2. Variance can actually be a little easier to appreciate as an instance of covariance between a random variable and itself!

My knowledge of variance was vague and formulaic. My knowledge of many things is like that. I don’t find that comforting, just a practical necessity in a limited life. But part of life’s pleasures is the freedom to NOT 80/20 something if doing so bugs you.

Alas, this one bugged me and the cost to fix it is lower now since LLM’s can be used as tireless tutors whose judgement of our faculties presents no threat.

I opened a chat and asked it to teach me Socratically, one small question at a time. When I run out of time or get tired, I know I can pick up where I left off. Or a little bit before that, since I usually need to insert before the point where I got tired since that point coincides with the material that made you take a break, so you don’t quite “own” it.

This is a snapshot of where I am in my Variance progression where I derive every formula from the already intuitive definition of “sum of squared deviations”:

 

As the learning progresses, you build more cases. With practice, you see that the key to all the derivations is that you are building on things you already know:

  1. The FOIL method from algebra
  2. PEMDAS from arithmetic
  3. The substitution that comes from seeing expectation or E[X] as nothing but a weighted average which means it’s equivalent to x_bar when each sample has equal weight

It’s hard to see this without practice.

When I revisit some of the derivations, I sometimes get stuck again but I know I’m screwing up one of these 3 foundational elements. That’s pretty crazy. It’s a gap in something I thought I knew cold, but the diagnosis is far more apparent because of how I’ve structured the learning in cahoots with Claude.

This is a timely place to name a few learning science techniques at play (see the appendix for more on these):

  • deliberate practice — deriving every variance formula from the definition of “sum of squared deviations,” over and over, refining each pass
  • desirable difficulty — doing the algebra by hand and taking pictures of the scratch work instead of watching it get done
  • spaced repetition — returning to the same derivation threads over days and weeks, not one sitting
  • expert guidance — Claude posing the next question and catching my errors with numerical counterexamples
  • layering skills — building each new case on FOIL, PEMDAS, and E[X]-as-weighted-average, things I already own
  • expertise reversal effect — starting with scaffolded one-question-at-a-time prompts rather than open-ended problems
  • consolidation — having Claude summarize what stuck, weighted to my actual gaps and the spots I tripped

The entire process is infused with the “generation effect” which takes advantage of our ability to remember something far better when you produce the answer yourself than when you read it.

And finally, every topic is a branch of an overarching commitment to interleaving. Power functions are mixed into a learning practice that includes other topics I want a closer look at. The approach makes affordances for both variety and synergy.

Before getting back to Scale and power functions, I have one more remark on this whole personal project I’ve donned Socrates 2026.

I really want AI companies to launch a Native Ink Surface with a submit button. Math derivation, music notation, art. All of these would be far less painful with a stylus. Is this too much to ask:


Automaticity

As I was reading and highlighting Scale, I strained to interpret the exponents. That means there are gaps in my understanding. Simple as that. These aren’t new concepts, but it’s clear I need some mental Dap if I want “automaticity”.

Paraphrasing Math Academy:

Automaticity is the ability to recall foundational math facts instantly and accurately from long-term memory, requiring zero conscious effort or working memory….

Automaticity is the prerequisite to true computational fluency. Once low-level skills (arithmetic, exponent rules, trig identities) are automated, recall becomes effortless, allowing your brain to focus entirely on higher-level problem solving and critical thinking.

We’ve been taught to think of tests (ie retrieval practice) as how you check whether you learned something. But it’s actually how you learn in the first place because it’s “doing”. To learn in a durable way is to “do”.

The highlights I collected became the raw material for Claude to design questions. But AI is obviously capable of far more than regurgitation, distillation and re-shuffling. It constructs sensible questions that arise from the text but not directly addressed. It can order the questions so they build gradually. It can relate material across domains. This is a gift to a learner.

Injecting a thought

AI cannot motivate you. It cannot inspire you. AI offers an unbundling of the tutor, not a replacement. The role of humans in the learning loop is going to grow, which might be a contrarian position. Think of coaches. Some are exceptional because they are masters of the Xs and Os. Some are exceptional because of their ability to lead and communicate. These are squishy. The squishy things will not rise in relative importance. They are important and AI doesn’t change that either way. It’s that AI will put a spotlight on the fact that there will be relatively higher yields to focus on the squishy. Whether we will or not (and be able to judge the delta) is an open question. A topic for another day perhaps, but I’m betting on this with my time.

There are no shortcuts. If you want automaticity, you gotta hit the gym.

The reps

We’re working with y = xᵃ throughout. The exponent a is the only thing carrying information about the relationship.

Warmup.

y = x². If x doubles, what happens to y?

POLL

y = x². If x doubles, what happens to y?

Goes up by 2
Goes up by a factor of 4
Goes up by a factor of 8
Stays the same
57 VOTES · · SHOW RESULTS

 

POLL

The rule: multiply x by some factor F, multiply y by F to the exponent. Here F is 2 and the exponent is 2, so y goes up by 2² = 4. Same law, y = x². If x triples?

Goes up b 3
Goes up by 6
Goes up by a factor of 9
Goes up by a factor of 27
38 VOTES · · SHOW RESULTS
POLL

In the last question, the exponent is fixed. You just swap the multiplier. Now a square root. y = √x, which is y = x¹ᐟ². If x quadruples, what happens to y?

Quadruples
Doubles
Goes up by a factor of 8
Halves
35 VOTES · · SHOW RESULTS

That question should feel familiar. Option prices follow a square root relationship with respect to time.. Doubling the time to expiry only multiplies a straddle by √2, while quadrupling it doubles the straddle.

POLL

Kleiber’s law. Metabolic rate scales as mass³ᐟ⁴. A mammal’s mass doubles. Its metabolic rate goes up by:

Exactly 2x (it doubles)
More than 2x
Less than 2x
It halves
30 VOTES · · SHOW RESULTS

The exponent is less than 1, so y grows slower than x. Less than double. This is the whole idea of sublinear scaling and economy of scale. Double the animal and it needs about 68% more energy, not 100% more.

POLL

Is 2³ᐟ⁴ the same as 2³ / 2⁴?

Yes, both equal 0.5
No, they’re different operations
Yes, both equal about 1.68
They’re both undefined
27 VOTES · · SHOW RESULTS

 

A fraction in the exponent is one number, not a division. 2³ᐟ⁴ means “take the fourth root of 2, then cube it,” which is about 1.68. Meanwhile 2³ / 2⁴ = 2³⁻⁴ = 2⁻¹ = 0.5. Dividing powers subtracts exponents.

POLL

What is 2⁻¹ᐟ⁴?

1/16
About .84
-1.19
-16
25 VOTES · · SHOW RESULTS

 

A negative exponent is always a reciprocal. Compute the positive version, then flip.

POLL

City infrastructure scales as population⁰·⁸⁵. A city’s population doubles. Total road length multiplies by roughly:

2.0
1.8
1.4
.85
22 VOTES · · SHOW RESULTS

2⁰·⁸⁵ sits between 2⁰·⁵ ≈ 1.41 and 2¹ = 2, closer to 2 because the exponent is close to 1. About 1.8. Roads go up 80% when the city doubles. Sublinear again, which means per person, road length actually falls.

POLL

City wages and output scale as population¹·¹⁵. Population doubles. Total wages multiply by roughly:

2.15
2.0
2.2
4.0
23 VOTES · · SHOW RESULTS

 

2¹·¹⁵ ≈ 2.22. Careful here. The move is NOT “2 plus 0.15.” It’s 2¹ × 2⁰·¹⁵ = 2 × 1.11 ≈ 2.22. Superlinear. Bigger city, disproportionately more output per person. Also disproportionately more crime and disease. Good and the bad scale together.

POLL

Strength scales as weight²ᐟ³. A horse weighs 8x what a small dog weighs. Per pound of body weight, the horse is:

Stronger than the dog
Exactly as strong per pound
Half as strong per pound
A quarter as strong per pound
22 VOTES · · SHOW RESULTS

 

Horse is 8²ᐟ³ = 4x stronger in total, but 8x heavier. So per pound it’s 4/8 = half as strong. This was Galileo’s observation. A small dog can carry two or three dogs on its back. A horse can’t carry even one. Strength grows like area, weight grows like volume, and volume outruns area as things get bigger.

Why Godzilla can’t exist

Strength scales with cross-sectional area, not size. This is why lumber is sold as a “2×4.” The two-by-four inches of cross-section is what bears the load. Double every dimension of a beam and its strength goes up 4x, because area scales with length squared.

Mass scales with volume, which is length cubed. Double every dimension and the thing weighs 8x more.

Scale a creature up and its weight (volume, 8x) outruns its strength (cross section, 4x) with every doubling. At Godzilla’s size, the legs would have to support a mass that has exploded as the cube of height while the bones holding it up only got stronger as the square. He’d snap under his own weight before he took a step. Same reason an ant can carry many times its body weight and an elephant can barely carry its own.

The toolkit

You build the knowledge, check yourself on new questions, come back another day, see how much ground you gave back. It’s 2 steps forward, 1 step back. Eventually you earn the consolidated reference and a sense that you have earned the shortcuts.

For y = xᵃ:

  1. Multiply x by F, and y multiplies by Fᵃ. This is scale invariance.
  2. For every order of magnitude in x, y changes by a orders of magnitude. This is the log-log slope reading.
  3. Double x, and y changes by a factor of 2ᵃ. This is the doubling sentence, the one West uses constantly. It’s natural for us to think of scaling with respect to doubling.
  4. Per unit of x, the quantity scales as xᵃ⁻¹. This is economy of scale versus increasing returns.

 

Interpreting the power

  • The sign tells you direction.
  • The magnitude tells you speed.
  • The distance from 1 tells you how it compares to a plain linear relationship.

Three worked slopes, read in orders of magnitude

y = x¹ᐟ², slope one half. y grows at half the rate of x. x goes up two orders of magnitude, y goes up one. Or: to get y up one order of magnitude, x has to move two.

y = x¹ᐟ⁴, slope one quarter. Even more damped. x times 10,000 (four orders of magnitude), y only times 10 (one order). This is per-cell metabolism, which scales as mass⁻¹ᐟ⁴.

y = x³ᐟ⁴, slope three-quarters. Kleiber. x times 10,000, y times 1,000. Three orders of magnitude of metabolism per four orders of magnitude of mass. The “3 to 4 ratio in powers of ten”.

The economy-of-scale shortcut

If the total scales as xᵃ, then per unit of x it scales as xᵃ⁻¹. Subtract 1 from the exponent and you have the per-person, per-pound, per-cell law. The sign of a−1 is the whole story:

  • a > 1: per-unit grows. Increasing returns.
  • a = 1: per-unit flat. Constant returns.
  • a < 1: per-unit shrinks. Economy of scale.

Cities have two exponents that mirror each other around 1. Physical stuff like roads and cables scales at 0.85, so per person it gets cheaper as the city grows. Social stuff like wages and patents scales at 1.15, so per person it grows.

Power law versus exponential: application to tail probability

Power law: the variable is in the base. It cares about ratios. Multiply the input, multiply the output, and the multiplier is the same no matter where you started.

Exponential: the variable is in the exponent. It cares about differences. Add to the input, multiply the output.

CAGR is a familiar exponential. Consider 10% CAGR. Going from year 5 to year 10 does not multiply your balance by the same factor as going from year 30 to year 60, even though both double the time. The multiplier depends on where you are, not on the ratio.

That base-versus-exponent distinction isn’t just about growth over time. It also governs how the probability of a large moves. The mean-standard deviation framework we are so familiar with is Gaussian “mediocrastian” math. But we know empirically that tails reside in “extremistan”. You can’t have a 10-sigma move every few decades. The bell curve is just a misspecified description of returns.

The fat tail in returns is more of a power law in the size of the move.

That’s a mouthful. Let’s make it easier.

Fix a horizon, say one day. Walk out from the average toward bigger and bigger moves and ask how fast the probability drops.

Gaussian answer: probability shrinks like e−x². Not just exponential, but exponential in the square of the move. Each additional unit of move costs you more probability than the last. A few units out and it’s effectively zero.

Power-law answer: probability shrinks like x^(−α), some fixed power of the move size. Double the move and you divide the probability by a constant factor (2^α), and you keep dividing by that same factor forever. There’s no cliff like the shoulder of a Gaussian curve.

The contrast is entirely about what sits where in the decay formula:

  • Gaussian: the move x is up in the exponent (e−x²). Move in the exponent means probability responds to differences in move size.
  • Power law: the move x is in the base (x−α). Move in the base means probability responds to ratios of move size.

Base means ratios and slow.

Exponent means differences and fast.

A bell curve places the move in the exponent and squares it, so its tail vanishes. A power law leaves the move in the base, extending its probability further out in the wings.

Handy intuition (and trivia!)

 

Wrapping up

To take the message of this 2-part series seriously means it’s unlikely that simply reading it imparted knowledge that you magically internalized (assuming you’re not multiple standard deviations up the IQ curve).

Instead, I hope you can employ AI tools to learn as you need. At your pace, and until your satisfaction. All those notes and highlights were not the learning itself but the fuel for powering a custom learning engine directed to your own goals and interests.

I of course hope that scaling laws, exponents, and variance were a desirable canvas to demonstrate the learning process, but they were not the point themselves. You can use any book as a starting point to go deeper. The combination of disaggregated training knowledge embedded in LLMs with narrower, concentrated material from an author or group of authors promises the best of 2 different advances against our humble ignorance.

Feel free to feed this post into your favorite agent to seed your own learning quest.