“end behavior”

Alex is an options trader you should follow in case he ever tweets a lot. Because he doesn’t, when he posted the question below a year ago, it got few responses. I took the liberty of posting it myself this week.

Article content

This was fun because it led to a lot of discussion on the timeline and DMs. I was told it sparked a bunch of quant debate on one trader’s desk.

The most popular answer, which was still less than 1/3 of the responses, was the correct answer.

Why?

The maximum value of a put is the strike. The maximum value of a call is the stock price.

Straddle is C + P so $100+$100 = $200

Notice how this means all call spreads go to zero since the calls are worth the same — the stock price. All put spreads go to their max value— the distance between strikes because the puts themselves are worth the strikes.

Logic for delta:

Delta is the change in option price per change in stock.

But the put’s strike is fixed, so the value of the put doesn’t depend on the stock price. The put has zero delta. It’s always worth $100. Which means it has no gamma either 🙂

The call is $100 because the max value of the call is the stock price. The call value moves 1-to-1 with the stock, so it has a delta of 1 or 100%

The max value of a straddle is therefore the stock price plus the strike price.

If you sell the straddle or either option at max value and hedge on its delta one time (this is known as a static hedge in contrast to dynamic hedging where you would rebalance as your hedge ratio changes), you cannot lose. It is that simple fact of arbitrage that makes it the upper bound.

To address the second most popular response in the poll, those who said the straddle is $100 (wrong) and has a 1.00 delta (correct), we will demonstrate why this is incorrect.

What’s your p/l if you sell 1 straddle at $100 and buy 100 shares against, if the stock goes to $300?

The straddle will be worth $400, so you lose $300 but make $200 on your long share.

Hmm, maybe I’m just underhedged. Fine, what if I hedge on a 200 delta?

In that case, you actually make money; you win $400 on your 2 shares more than offsetting the $300 straddle loss. But what if the stock went to zero?

Your straddle p/l is unchanged, but you lost $200 on the long stock position. Arbitrage max value means you cannot lose if you sell at that price. Since we found a losing scenario, the price is not the maximum arbitrage bound. If you sell the straddle at $200 and buy a single share of stock, there’s no scenario where you lose. It is the lowest straddle value for which this no-lose scenario is true, thus it’s the arbitrage bound.

Of course, this is but a toy problem where the call and put go to their maximum values because it’s a degenerate case of infinite time or vol. But learning how a function (an option price is just a function) behaves by observing its boundaries is good for intuition. You did this in 9th grade. Khan Academy can jog your memory:

Article content

In the real world, you can fleetingly find options that trade beyond their arbitrage values:

Article content

Earlier in the week, @DeepDishEnjoyer aka p4 wrote a thread about a dividend mispricing.

It led to some back and forth with passersbys who use options but appear to have large gaps in the fundamentals.

Between the maximum value poll and p4’s dividend lesson, it’s worth saying it:

In a proper option education, you spend a lot of time on arbitrage relationships, cost of carry, and synthetics before you ever hear the word “volatility”.

I didn’t study formal math but I imagine there’s a lot in common with the process of proofs. Arbitrages rest heavily on assumptions. So to understand the relationships, you are forced into an intimate familiarity with the assumptions. And in the extremes of everything, it’s the failure to examine assumptions that leads to being blindsided. But also, when things get extreme, to go on the attack means asking yourself, “Who’s on autopilot? Is this price resting on a stale assumption?” The arbitrage relationships give you the highest conceptual ROI that derivatives offer, you never learn the most useful thing derivatives can teach…passage over the “bridge of asses”.

If you want to see more examples of why option basics are so key to understanding assumptions and opportunities when things get weird:

blindsided

My cousin Nicole just had her first child in the past year and is now compiling resources for new parents in light of where her attention has obviously been.

In our family chat, she asked:

When you became a new parent, what blindsided you the most?

I’ll share my answer, which I qualify with both the awareness that having a child is a gamble on many levels and that conception itself is a miracle and should never be taken for granted. What was I blindsided by?

[9:16 AM, 9/16/2026] Kris Abdelmessih: that they were gonna be so awesome

[9:17 AM, 9/16/2026] Kris Abdelmessih: that last one is important when we live in times where people have less kids and talk about it as though it’s a chore (it is) but not the amazing upside

The most transcendent single moment of my life thus far was to hear my son’s voice the day he came into the world. I don’t know if that’s every parent’s experience, but immediately I felt the joy and clarity of purpose. Before that cry, I don’t think I would have said I had no purpose, but the moment revealed that I didn’t believe I did. The speed and intensity of this rush of belief was a novel feeling. Thus, irreparably blindsided.

Share your own answers in the comments and I’ll share them with Nicole. Thank you!


I offered a couple of less serious answers to her question as well.

Article content

On that last one, I wrote about that 2 years ago in our minds love to betray us:

Article content

When I was at the Sphere with my family over Spring Break, I wouldn’t ride the long, exposed escalators. I took the elevator where I found the rest of the scaredy-cats.

If there’s anything good about having a phobia, it’s empathy for the range of what can go on in people’s minds and bodies.

I’m watching this poor guy thinking, don’t do it man, it’s not worth and all he wants is a glimpse:

Man crawls through his phobia of heights to get a view of the Atlantic Ocean from the edge of a cliff. 😅

Article content

1:44 PM · Sep 12, 2026 · 36.9M Views


2.99K Replies · 5.3K Reposts · 131K Likes

The comment section understands, and based on the number of “likes”, many others do too.

Article content

Anyway, I blame my kids for my embarrassment when I go to the Sphere to see Metallica with a group of guys next month and have to explain that I’ll meet them at the seats.

Moontower #328

In this issue:

  • blindsided
  • “end behavior”

Friends,

Blindsided

My cousin Nicole just had her first child in the past year and is now compiling resources for new parents in light of where her attention has obviously been.

In our family chat, she asked:

When you became a new parent, what blindsided you the most?

I’ll share my answer, which I qualify with both the awareness that having a child is a gamble on many levels and that conception itself is a miracle and should never be taken for granted. What was I blindsided by?

[9:16 AM, 9/16/2026] Kris Abdelmessih: that they were gonna be so awesome

[9:17 AM, 9/16/2026] Kris Abdelmessih: that last one is important when we live in times where people have less kids and talk about it as though it’s a chore (it is) but not the amazing upside

The most transcendent single moment of my life thus far was to hear my son’s voice the day he came into the world. I don’t know if that’s every parent’s experience, but immediately I felt the joy and clarity of purpose. Before that cry, I don’t think I would have said I had no purpose, but the moment revealed that I didn’t believe I did. The speed and intensity of this rush of belief was a novel feeling. Thus, irreparably blindsided.

Share your own answers in the comments and I’ll share them with Nicole. Thank you!


I offered a couple of less serious answers to her question as well.

Article content

On that last one, I wrote about that 2 years ago in our minds love to betray us:

Article content

When I was at the Sphere with my family over Spring Break, I wouldn’t ride the long, exposed escalators. I took the elevator where I found the rest of the scaredy-cats.

If there’s anything good about having a phobia, it’s empathy for the range of what can go on in people’s minds and bodies.

I’m watching this poor guy thinking, don’t do it man, it’s not worth and all he wants is a glimpse:

Man crawls through his phobia of heights to get a view of the Atlantic Ocean from the edge of a cliff. 😅

Article content

1:44 PM · Sep 12, 2026 · 36.9M Views


2.99K Replies · 5.3K Reposts · 131K Likes

The comment section understands, and based on the number of “likes”, many others do too.

Article content

Anyway, I blame my kids for my embarrassment when I go to the Sphere to see Metallica with a group of guys next month and have to explain that I’ll meet them at the seats.


Money Angle

A couple of Option Trench episodes to share:

📺Terminal vs Path-Dependent Value Explained Using Collars | 39 min

📺The (Not So) Efficient Market Hypothesis? | 59 min

The first one will be useful for anyone wanting to learn more about option collars, which I’ve been writing a lot about. A video might be a gentler format so check that out.

The second one applies to investing broadly. I also use the Paradox of Provable Alpha at the end to answer a good question Erik asks.


Money Angle For Masochists

Alex is an options trader you should follow in case he ever tweets a lot. Because he doesn’t, when he posted the question below a year ago, it got few responses. I took the liberty of posting it myself this week.

Article content

This was fun because it led to a lot of discussion on the timeline and DMs. I was told it sparked a bunch of quant debate on one trader’s desk.

The most popular answer, which was still less than 1/3 of the responses, was the correct answer.

Why?

The maximum value of a put is the strike. The maximum value of a call is the stock price.

Straddle is C + P so $100+$100 = $200

Notice how this means all call spreads go to zero since the calls are worth the same — the stock price. All put spreads go to their max value— the distance between strikes because the puts themselves are worth the strikes.

Logic for delta:

Delta is the change in option price per change in stock.

But the put’s strike is fixed, so the value of the put doesn’t depend on the stock price. The put has zero delta. It’s always worth $100. Which means it has no gamma either 🙂

The call is $100 because the max value of the call is the stock price. The call value moves 1-to-1 with the stock, so it has a delta of 1 or 100%

The max value of a straddle is therefore the stock price plus the strike price.

If you sell the straddle or either option at max value and hedge on its delta one time (this is known as a static hedge in contrast to dynamic hedging where you would rebalance as your hedge ratio changes), you cannot lose. It is that simple fact of arbitrage that makes it the upper bound.

To address the second most popular response in the poll, those who said the straddle is $100 (wrong) and has a 1.00 delta (correct), we will demonstrate why this is incorrect.

What’s your p/l if you sell 1 straddle at $100 and buy 100 shares against, if the stock goes to $300?

The straddle will be worth $400, so you lose $300 but make $200 on your long share.

Hmm, maybe I’m just underhedged. Fine, what if I hedge on a 200 delta?

In that case, you actually make money; you win $400 on your 2 shares more than offsetting the $300 straddle loss. But what if the stock went to zero?

Your straddle p/l is unchanged, but you lost $200 on the long stock position. Arbitrage max value means you cannot lose if you sell at that price. Since we found a losing scenario, the price is not the maximum arbitrage bound. If you sell the straddle at $200 and buy a single share of stock, there’s no scenario where you lose. It is the lowest straddle value for which this no-lose scenario is true, thus it’s the arbitrage bound.

Of course, this is but a toy problem where the call and put go to their maximum values because it’s a degenerate case of infinite time or vol. But learning how a function (an option price is just a function) behaves by observing its boundaries is good for intuition. You did this in 9th grade. Khan Academy can jog your memory:

Article content

In the real world, you can fleetingly find options that trade beyond their arbitrage values:

Article content

Earlier in the week, @DeepDishEnjoyer aka p4 wrote a thread about a dividend mispricing.

It led to some back and forth with passersbys who use options but appear to have large gaps in the fundamentals.

Between the maximum value poll and p4’s dividend lesson, it’s worth saying it:

In a proper option education, you spend a lot of time on arbitrage relationships, cost of carry, and synthetics before you ever hear the word “volatility”.

I didn’t study formal math but I imagine there’s a lot in common with the process of proofs. Arbitrages rest heavily on assumptions. So to understand the relationships, you are forced into an intimate familiarity with the assumptions. And in the extremes of everything, it’s the failure to examine assumptions that leads to being blindsided. But also, when things get extreme, to go on the attack means asking yourself, “Who’s on autopilot? Is this price resting on a stale assumption?” The arbitrage relationships give you the highest conceptual ROI that derivatives offer, you never learn the most useful thing derivatives can teach…passage over the “bridge of asses”.

If you want to see more examples of why option basics are so key to understanding assumptions and opportunities when things get weird:

From My Actual Life

I leave you with another pic from our family chat where my wife posted a photo of where she was walking.

Article content

I don’t know how many Egyptian Arabic speakers we got in the crowd but “shib-shib” is like a slipper. Mom’s weapon of choice. Apparently this is a broader thing:

Mothers and shoes

Stay groovy

☮️


Moontower Weekly Recap

Variance & Covariance Cheat Sheet

Variance & Covariance Cheat Sheet

Starting points (where every derivation begins)

Everything below is derived from these definitions. They’re the raw material — average squared deviation for variance, average product of deviations for covariance. When a derivation feels stuck, come back here and plug in.

Variance — average squared deviation from the mean
Var(X) = E[(X μ)2]    μ = E[X]
Covariance — average product of deviations
Cov(X, Y) = E[(X μX)(Y μY)]
Sum of squared deviations (the un-averaged version)
SS = Σ (xi μ)2    Var = SSn

Variance is just SS divided by n (or n−1 for a sample). Same object, before you average.

The move in every derivation: plug into one of these, expand the square or product (pure algebra), apply E using linearity, then recognize the Var/Cov patterns that fall out. The computational forms below (E[X2] (E[X])2, E[XY] E[X]E[Y]) are results of doing this, not starting points.


The identities

Variance from the definition
Var(X) = E[(X μ)2] = E[X2] (E[X])2

Average of the squares minus the square of the average. Worth showing where that second form comes from, since every later grind reuses this exact collapse. Start from the deviation definition and expand the square:

1n Σ(xi x)2 = 1n Σ(xi2 2xix + x2)

Average term by term. The key is that x is a constant (already computed), so it pulls out of the sums:

  • First term: (1/nxi2 = E[X2]
  • Middle term: (1/n)Σ(2xix) = 2x · (1/nxi = 2x · x = 2(E[X])2
  • Last term: (1/nx2 = x2 = (E[X])2 (averaging a constant returns the constant)

Put them together — and notice the last term carries a coefficient of 1, not 2:

E[X2] 2(E[X])2 + (E[X])2 = E[X2] (E[X])2

The 2 and +1 combine to 1. That collapse — middle and last terms both becoming (E[X])2 and partially cancelling — is the same move behind every Var/Cov identity on this sheet.

Covariance from the definition
Cov(X, Y) = E[(X μX)(Y μY)] = E[XY] E[X]E[Y]

Average of the products minus the product of the averages.

Variance is covariance with itself
Cov(X, X) = Var(X)
Scaling rule for variance
Var(aX) = a2 · Var(X)

Constants pull out as their square.

Scaling rule for covariance
Cov(aX, bY) = ab · Cov(X, Y)
Shifting rule for covariance
Cov(X + c, Y) = Cov(X, Y)

Adding a constant doesn’t change covariance.

Covariance with a constant is zero
Cov(X, c) = 0

Constants don’t co-vary.

Variance of a sum
Var(X + Y) = Var(X) + Var(Y) + 2 · Cov(X, Y)
Variance of a weighted sum (the workhorse)
Var(aX + bY) = a2 Var(X) + b2 Var(Y) + 2ab · Cov(X, Y)

Here a and b are the amounts held of each asset. They’re portfolio weights when they sum to 1. Var(X+Y) above is just this formula with a = b = 1 — one unit of each, no weighting lever. The weights are what turn a raw sum into a portfolio.

Worked example. Two assets: σX = 20%, σY = 10%, ρ = 0.3. Equal weights a = b = 0.5.
  • Var(X) = 0.04,   Var(Y) = 0.01
  • Cov(X, Y) = ρ · σX · σY = 0.3 · 0.2 · 0.1 = 0.006
  • Var(P) = 0.25·0.04 + 0.25·0.01 + 2·0.5·0.5·0.006 = 0.01 + 0.0025 + 0.003 = 0.0155
  • σP = √0.0155 ≈ 12.4%
Compare to the naive weighted-average vol, 0.5·20% + 0.5·10% = 15%. The cross term (with ρ < 1) is what pulls portfolio vol below the average of the two vols. That gap is the diversification benefit.
Variance of a difference (spread variance / pair-trading formula)
Var(X Y) = Var(X) + Var(Y) 2 · Cov(X, Y)

Same as Var(X+Y) but cross term flips sign. When X and Y are highly correlated, spread variance is small — the math behind why pair trades work.

Bilinearity of covariance
Cov(X, A + B) = Cov(X, A) + Cov(X, B)

Same in the first slot by symmetry.

Where it’s used. This is the move that lets you compute an asset’s covariance with a whole portfolio without re-deriving anything. Say a portfolio P = 0.5A + 0.5B and you want how asset A co-moves with the portfolio it sits in:
Cov(A, P) = Cov(A, 0.5A + 0.5B) = 0.5 Var(A) + 0.5 Cov(A, B)
Distribute across the sum, pull the weights out. That number — an asset’s covariance with its own portfolio — is its marginal contribution to portfolio risk, and it’s exactly what you FOIL out when you expand Var(w1X1 + … + wnXn) into the full covariance matrix. Bilinearity is the engine under every portfolio-variance calculation.
Variance of a binomial
Var(H) = np(1p)   where H = Σ Xi

Derived in two steps, both from scratch.

Step 1 — variance of a single flip. One flip X is 1 with probability p, 0 with probability (1p). Mean is E[X] = p. Plug into the squared-deviation definition — deviations are (1p) for heads and (0p) = p for tails, each weighted by its probability:

Var(X) = p(1p)2 + (1p)p2

Factor out p(1p): the bracket is (1p) + p = 1, so

Var(X) = p(1p)

(Peaks at p = 0.5, value 0.25 — the fair coin is the most uncertain, most variance per flip.)

Step 2 — n flips. Write H as a sum of n independent single flips, H = X1 + … + Xn. Variance of a sum adds the pairwise Cov terms, but independent flips have Cov(Xi, Xj) = 0, so every cross term drops. The n identical variances just add:

Var(H) = Σ Var(Xi) = n · p(1p)

The np(1p) isn’t handed to you — it falls out of one Bernoulli’s p(1p) times n, because independence kills the covariances.

Standard deviation scaling
StDev(aX) = |a| · StDev(X)
Correlation definition
ρ = Cov(X, Y)σX · σY

When ρ = 1: Cov(X, Y)2 = Var(X) · Var(Y).


The derivation recipe

Every identity in this neighborhood comes out of the same five moves. When you see a Var or Cov of something built from linear combinations of random variables, this is the procedure.

  1. Plug into the definition. Use Var(Z) = E[Z2] (E[Z])2 or Cov(X, Y) = E[XY] E[X]E[Y] depending on what you’re computing.
  2. Expand squares and products. Pure algebra on the random variables. FOIL out any binomials. No expectations yet.
  3. Apply E using linearity. Distribute E across sums, pull constants out of expectations. This is the step that does the most work. Always handle linearity first when you have the chance — squaring is not linear, so you simplify E first and let the square wrap what’s left.
  4. Group matching terms. Line up the things that share factors (a2 terms together, ab terms together, b2 terms together, etc.).
  5. Factor and recognize. Pull out shared factors and spot the patterns: (E[X2] (E[X])2) is Var(X), and (E[XY] E[X]E[Y]) is Cov(X, Y).

The reason this recipe always closes: variances and covariances are quadratic in the underlying random variables, so expanding any square or product of linear combinations only generates more variances and covariances. Step 5 is recognition, not computation. There’s nowhere else for the algebra to land.


Two applications

Interview problem: E[H · T] for n coin flips

Flip a fair coin n = 100 times. H = heads, T = tails. Find E[H · T]. Worked slowly, because the one-line answer hides about six moves.

Step 1 — first reach, and why it fails. The instinct is E[H · T] = E[H] · E[T] = 50 · 50 = 2,500. But splitting a product of expectations like that is only legal when the two variables are independent. Check the precondition: H + T = 100, so knowing H pins down T exactly. Not independent. The naive split is off by a correction.

Step 2 — name the correction. That correction is what covariance is. Rearranging Cov(X, Y) = E[XY] E[X]E[Y]:

E[H · T] = E[H] · E[T] + Cov(H, T)
Independent → Cov = 0 → naive split exact. Locked → Cov ≠ 0 → you need the term.

Step 3 — get the sign first. H + T = 100, so when H is above its mean, T is forced below. They move opposite, always. So Cov(H, T) is negative, and the true answer lands below 2,500.

Step 4 — compute Cov(H, T) via substitution. Since T = 100 H, write Cov(H, T) = Cov(H, 100 H) and split with bilinearity:

Cov(H, 100 H) = Cov(H, 100) + Cov(H, H)
First term is covariance with a constant → 0. Second term, pull out the 1 (scaling rule, ab = 1 · (1) = 1):
= 0 Cov(H, H) = Var(H)
So Cov(H, T) = Var(H). Now it’s earned, not asserted.

Step 5 — Var(H) is the binomial variance. H is the count of heads in n flips, so Var(H) = np(1p) = 100 · 0.5 · 0.5 = 25.

Step 6 — land it.

E[H · T] = 2,500 25 = 2,475

Where n(n1) comes from. Keep everything in symbols instead of plugging in. E[H] = np and E[T] = n(1p), so E[H]·E[T] = n2p(1p), and Var(H) = np(1p). Then:

E[H · T] = n2p(1p) np(1p) = p(1p)[n2 n] = n(n1)p(1p)
The n2 is the naive product, the n is the variance shortfall, and factoring out p(1p) leaves n(n1). Sanity check: 100 · 99 · 0.25 = 2,475. ✓

The through-line: the answer falls short of E[H]·E[T] by exactly Var(H), because Cov(H, T) = Var(H) whenever H and T sum to a constant.

Two-stock equal-weight portfolio with equal variances σ2 and correlation ρ

Start from the workhorse with a = b = 0.5:

Var(P) = 0.25 Var(X) + 0.25 Var(Y) + 2(0.5)(0.5) Cov(X, Y)

Impose equal variances Var(X) = Var(Y) = σ2, and write the cross term with correlation, Cov(X, Y) = ρσ2 (since Cov = ρ · σX · σY and both vols are σ):

Var(P) = 0.25σ2 + 0.25σ2 + 0.5ρσ2 = 0.5σ2 + 0.5ρσ2

Factor out 0.5σ2:

Var(P) =  σ2(1 + ρ)2
Why this is the instructive form. Weights are fixed (50/50) and both vols are fixed (σ). The only thing left moving is ρ. So the entire diversification effect is carried by the single factor (1 + ρ)/2 — a clean dial from 0 to 1 that multiplies the single-name variance. Sweep ρ and read what correlation actually does:
ρ Var(P) σP (vol) vs. one stock
+1σ2σno benefit — identical names
+0.50.75σ20.87σ13% vol cut
00.5σ20.71σ29% vol cut (the √½ case)
−0.50.25σ20.5σhalf the vol
−100risk fully cancels
The variance scales linearly in ρ, but the thing you feel — vol, σP = σ√((1+ρ)/2) — scales as the square root, so the first chunk of decorrelation buys more than the last. Going from ρ = 1 to ρ = 0.5 already takes 13% off your vol. You do not need negative correlation to diversify; anything below +1 helps. Negative correlation is just the strong form, and ρ = −1 is the perfect hedge where the two positions cancel outright.

The whole two-name diversification story lives in that (1 + ρ)/2 factor. Same vols, same weights, and correlation alone moves you from “no benefit” to “risk gone.”

Two-stock unequal-weight portfolio with equal variances σ2 and correlation ρ
Var(P) = σ2 [1 2w(1w)(1ρ)]

Diversification benefit is the product of a weight piece (2w(1w), maxed at w = 0.5) and a correlation piece (1ρ). Need both to get benefit. With equal variances, equal weighting is optimal — any tilt from 50/50 sacrifices diversification.

Two-stock unequal-variance portfolio → inverse-variance weighting

Now drop the equal-variance assumption. Keep σX2 and σY2 separate. Weights w and (1w):

Var(P) = w2σX2 + (1w)2σY2 + 2w(1w)ρσXσY

Minimize over w. Var(P) is an upward parabola in w (positive coefficient on w2), so the critical point is the min. Differentiate term by term and set to zero:

2wσX2 2(1wY2 + 2(12w)ρσXσY = 0

Divide by 2, expand, collect the w terms on the left and constants on the right, factor w out:

wX2 + σY2 2ρσXσY) = σY2 ρσXσY
w* =  σY2 ρσXσYσX2 + σY2 2ρσXσY

The denominator is Var(X Y) — the spread variance from earlier. The numerator is σY2 Cov(X, Y).

The payoff — set ρ = 0 (independent names):
w* = σY2σX2 + σY2 = 1/σX21/σX2 + 1/σY2
That’s inverse-variance weighting: each asset’s weight is its inverse variance over the sum of inverse variances. The quieter asset gets more money. This is the result behind Kalman filters, weighted least squares, and meta-analysis — anywhere you optimally combine noisy estimates, you weight by precision (1/variance). It’s also the “optimal” cousin of the inverse-vol risk-parity heuristic, which ignores correlations.

Intuition — hold σX fixed at 20% (σX2 = 0.04), turn the σY knob:

σY σY2 w* on X
000
10%0.010.20
20%0.040.50
40%0.160.80
→ 1

X’s weight is driven by Y’s variance, not its own. The noisier the alternative, the more you pile into X. Three anchors: σY2 = 0 → w* = 0 (Y is riskless, hold only Y); σY2 = σX2w* = 0.5 (equal variances recover equal weighting); σY2 → ∞ → w* → 1 (Y is pure noise, flee into X). The cleanest limit: if σX2 = 0, then w* = 1 — a riskless X takes the whole book. Precision is just quietness, and you trust the quiet estimate more.

Three-variable portfolio variance → why the matrix shows up

Same Form A grind, one more variable. Expand (aX + bY + cZ)2, apply E, subtract the squared-mean term. Every squared term becomes a variance, every cross term a covariance:

Var(aX + bY + cZ) = a2Var(X) + b2Var(Y) + c2Var(Z) + 2ab Cov(X,Y) + 2ac Cov(X,Z) + 2bc Cov(Y,Z)

Counting the terms. For n assets you always get:

  • n variance terms — one per asset (the ai2 Var pieces)
  • nC2 = n(n−1)/2 covariance pairs — one per distinct pair

Total = n + nC2. For n = 3: 3 + 3 = 6. For n = 100: 100 variances + 4,950 covariance pairs. The cross terms grow as n2, which is exactly why nobody writes portfolio variance longhand past n = 3 — you switch to the matrix form.

The double-sum / matrix form. Organize every term into a grid indexed by asset pairs. With weights wi and returns ri:
Var(P) = Σi Σj wi wj Cov(ri, rj) = wΣw
Reading the double sum: the outer Σ over i and the inner Σ over j together form every ordered pair (i, j). For each pair you drop in one term, wiwjCov(ri, rj), and add them all up. For n = 3 that’s 3 × 3 = 9 cells:
X (j=1) Y (j=2) Z (j=3)
X (i=1)a2Var(X)ab Cov(X,Y)ac Cov(X,Z)
Y (i=2)ab Cov(X,Y)b2Var(Y)bc Cov(Y,Z)
Z (i=3)ac Cov(X,Z)bc Cov(Y,Z)c2Var(Z)
Sum all nine cells and you get the six-term formula above. Two things to see:
  • Diagonal (i = j, shaded): Cov(ri, ri) = Var(ri), so the diagonal is the n variance terms.
  • Off-diagonal (ij): each unordered pair appears twice — cell (X,Y) and cell (Y,X) are identical — and those two copies are exactly where the factor of 2 on each covariance comes from. You never write the 2 by hand; the grid double-counts it for you.
So Σ is the covariance matrix: variances down the diagonal, covariances off it. wΣw just says “sweep every cell of the grid, weight it, sum it.” The n + nC2 count is the matrix — diagonal plus (doubled) upper triangle. You’ve already discovered why the matrix form is inevitable; it’s just bookkeeping for the term explosion.
Figure — the double sum is the matrix: one sweep, three views
Var(P) = Σi Σj wiwj Cov(ri, rj) an instruction for sweeping a grid: for every row i, for every column j, add that cell inner Σ over j → picks the column X (j=1) Y (j=2) Z (j=3) outer Σ over i → picks the row X (i=1) Y (i=2) Z (i=3) w1² Var(X) w1w2 Cov(X,Y) w1w3 Cov(X,Z) w2w1 Cov(X,Y) w2² Var(Y) w2w3 Cov(Y,Z) w3w1 Cov(X,Z) w3w2 Cov(Y,Z) w3² Var(Z) Diagonal (i = j): Cov(ri, ri) = Var(ri) — the 3 variance terms. Twin cells (i,j) & (j,i): identical — every covariance is visited twice. That double-count is the ×2 you wrote by hand. The grid supplies it for free. add all 9 cells = wT Σ w Σ is the covariance matrix: variances on the diagonal, covariances off it. n assets = the same sweep on an n×n grid. Nothing new happens; the grid just grows.
Figure — three forms of the same variance, and where the 2s go
① ALGEBRAIC — you write the 2s by hand a²Var(X) + b²Var(Y) + c²Var(Z) 2ab Cov(X,Y) + 2ac Cov(X,Z) + 2bc Cov(Y,Z) ② MATRIX — the 2s vanish into the symmetry Var(P) = wT Σ w,   where Σ = Var(X) Cov(X,Y) Cov(X,Z) Cov(X,Y) Var(Y) Cov(Y,Z) Cov(X,Z) Cov(Y,Z) Var(Z) Each covariance sits in two matched-color cells (mirrored across the diagonal). Summed, the two cells are the 2 from Stage 1. Variances (diagonal, purple) have no mirror — that’s why they’re never doubled. ③ DOUBLE SUM — the matrix written in math Var(P) = Σ n i=1 Σ n j=1 wiwj Cov(ri, rj) Cov(ri, ri) = Var(ri) — covariance with itself is just variance, so the diagonal needs no special case. Both sums run 1 to n, so the sweep is n² terms. 10 assets → 100 terms, not 20. That quadratic blow-up is exactly why the compact matrix form earns its keep.
Notation you’ll see in practice: wΣw. You’ll run into this constantly in risk models, optimizers, and quant papers. It’s the same portfolio variance, packaged as a matrix operation. Reading it piece by piece:
  • w — the weight vector, weights stacked in a column.
  • w — “w transpose,” the same weights laid flat as a row. Transpose just tips a column over into a row.
  • Σ — the covariance matrix (the grid above). Watch out: this capital-sigma is the matrix, not a summation sign. Variances on the diagonal, covariances off it.
So wΣw is row-of-weights × matrix × column-of-weights, which multiplies out to a single number — the portfolio variance. It’s identical to the double sum, and in a spreadsheet it’s literally =MMULT(MMULT(TRANSPOSE(w), Σ), w).

Why bother, when the double sum already shows everything? Three practical reasons, none of them “it’s more correct.” It doesn’t grow — three symbols whether n is 2 or 2,000, where the double sum for 500 names is 250,000 terms. It’s how software actually computes portfolio variance (one fast matrix op). And optimization only speaks matrix: the minimum-variance weights come out as Σ−11 normalized, and the inverse Σ−1 has no double-sum spelling. The two-asset inverse-variance weighting derived above is Σ−1 for n = 2 — the matrix form is how that generalizes. For understanding, the double sum is enough; this is the notation for doing things with it.


The ladder

Each rung is built from the one before it — the definition first, then the algebra of scaling and adding, then portfolios, then the jump to the matrix. Nothing is assumed that wasn’t derived earlier.

1.  Variance as average squared deviation
2.  Var(X) = E[X2] (E[X])2 — from the definition
3.  Cov(X, Y) = E[XY] E[X]E[Y] — from the definition
4.  Var(aX) = a2·Var(X)
5.  Cov(aX, bY) = ab·Cov(X, Y)
6.  Cov(X + c, Y) = Cov(X, Y) and Cov(X, c) = 0
7.  Var(X + Y) = Var(X) + Var(Y) + 2·Cov(X, Y)
8.  Var(aX + bY) — the full weighted-sum workhorse
9.  Two-stock equal-weight portfolio variance → the (1+ρ)/2 diversification factor
10. Two-stock unequal-weight portfolio variance (equal variances)
11. Var(X Y) — spread variance / pair-trading formula
12. Variance of a binomial = np(1p)
13. Two-stock unequal-variance portfolio → inverse-variance weighting
14. Var(aX + bY + cZ) — three variables, and the n + nC2 term count
15. General n-asset portfolio variance — the wΣw matrix form

Is this one lesson in a math course?

No. This would be roughly half a semester of a first probability course, or a full chapter and a half of a more applied stats book.

Rough mapping to a standard curriculum:

  • Variance from the definition, E[X2] identity: one lecture, plus a problem set
  • Covariance and the product identity: one lecture
  • Scaling rules, bilinearity, variance of a sum: one to two lectures
  • Portfolio variance, weighted sums, correlation: one lecture in the probability course, or the opening week of a finance/portfolio-theory course
  • Binomial variance, applications: another lecture or two

So this is the equivalent of maybe four to six lectures of material, plus the problem sets that go with them. The reason it feels like a lot is that this sheet does the whole pipeline — derivation, intuition, numerical examples, applications — for each piece, instead of showing a formula and moving on.

The trade-off is real: this is slower but produces durable understanding. A typical math course shows you Var(aX + bY) on day one, leaves the “why it’s that and not something else” fuzzy, and you pattern-match for the rest of the semester. Done this way, when σ2 · (1+ρ)/2 turns up in a textbook two years later, you see the bilinearity FOIL behind it instead of recognizing a memorized formula.

That’s the trade. Slower, but it sticks.

why i’m not more bearish equities compared to 6 months ago

Programming Note: I normally publish Munchies on Wednesday and the paid post on Thursday, but this week I will post both a day early, as they are market-related and I want to get them out ahead of the Fed meeting.


Friends,

I’ll start with updating some broad market observations I laid out in March and then square that with what option surfaces are telling us in the context of the Fed meeting and beyond. The Fed meeting is a highly skewed event with a 25 bp hike more than about 90% priced in.

The probability of Fed target rate between 3.75% and 4% has shot from 35% to 90% in under 3 weeks

Recent context:

The Jackson Hole keynote was on 8/28/26. The probability of a rate hike shot from 35% to 57% in one day.

The 10-year yield has rallied from 4.67% on 8/26 to 4.96% as I write on 9/14.

In that same window, the SPX is down a mere 50 bps and the NDX is basically unched.

Let’s rewind to my March post trading is like a sudoku puzzle with prices as the given numbers. I compared earnings yields to bond yields to get an equity risk premium. This comparison in the modern era of massive budget deficits, where a large “G” in the Kalecki-Levy world tells us that money ends up as private sector nominal income by identity*, never flatters equities. We can reconcile skinny equity risk premiums the same we reconcile low earnings yield any stock…there’s an expectation of growth.

*While this is reductionist in the sense that the flow doesn’t definitionally have to end up there, it also happens to be where it has ended up.

If we are desensitized to the low absolute level of equity risk premiums, it is because it has been easy to presume nominal growth. We have relatively low unemployment, technology companies growing quickly despite massive scale, and, I almost forgot, A GOVERNMENT THAT IS EXISTENTIALLY POT-COMMITTED TO DEFICIT SPENDING.

“But, we have a Republican president.”

Ok, define Republican. Because if you look closely at our fiscal…oh never mind. Nobody cares. We’re well past fiscal discipline as a political delimiter.

The more money there is, the more surface area there is for theft/grift/graft.

[While this is now quite obvious and exploited by both right and left, Don’s brazenness feels like it’s giving everyone permission. He is a populist wrestling everyone’s birthright to cheat away from the cloaked political elites who once monopolized it. That sentence is the dress…whether you see blue or gold is up to you.]

Alrighty then, where were we? Ahh, yes, nominal growth as a given, because our society is a passenger in a global economic trolley problem. Fantastic. Instead of fretting over the level of the equity risk premium, we can look at the past 6 months to consider the change or, in a sudoku-esque way, solve for what needs to happen for the risk premium to not get worse.

Thus far, higher energy prices and higher bond yields haven’t produced the equity decline in the scenario I outlined. My hunch is the market got more expensive on a relative basis, but let’s investigate.

First, the updated numbers.

Market quotes below are September 14 intraday observations, taken at different times. The equity calculations use the same snapshot as the charts.

WTI oil is up 18.5% since the March post (March price from 3/27/26)

RBOB gasoline is up ~28%

IEF on a div-adjusted basis is down 1.9%, about half what the duration (which is a snapshot like delta) expects because you picked up over 200 bps of carried interest (ie those divs) in the meantime.

Equity valuation as the missing Sudoku number

In late March, I used an equity earnings yield of roughly 5% vs 4.4% ten-year Treasury yield. This represented a 60 bps nominal equity risk premium and ~300 bps in real terms. My downside scenario considers what might be expected if the Treasury yield reached 5.5% and equities needed to offer another 50 bps above that. The target nominal earnings yield would rise to 6%, requiring roughly 17% equity decline if earnings stayed flat.

On partial probabilities

My thinking was on a subset of causes for yields to rise. I was thinking about the inflation channel via energy price pass-through in the event that futures prices rolled up to war-bolstered petroleum spot prices. This is supply-shock inflation, but there is also demand-pull inflation. Yields can also increase because of concerns about sovereign creditworthiness. Asset prices are complex because they embed expectations about many variables, some of which reinforce each other and some offset. Arrows in every direction. Complex portfolios can be constructed to isolate bets on conditional or partial probabilities.

An example from sports betting:

Say you bet $900 to win $100 on the Seahawks not winning the Super Bowl, and $100 to win $600 on them winning the NFC. If they don’t win the NFC, the bets cancel. If they reach the Super Bowl, you make $700 if they lose and lose $300 if they win. You’ve constructed a conditional bet: Seattle to lose the Super Bowl, provided they get there.

The prices imply a 10% chance of winning the Super Bowl and a 14.3% chance of getting there. Divide those and you get a 70% chance of winning conditional on getting there. Your combined position bets against that 70%.

The way I’m reasoning in a macro way is implicitly partial. All relative value bets have this property, but be aware that if you bet on cross-asset, you are not truly isolating bets as cleanly as the Seahawks example. You are hand-waving all the other ways a bond yield, oil price, or stock earnings can change relative to one another. There’s no equivalent to “If they don’t win the NFC, the bets cancel” because the relationships in assets are not deterministic in the same way that winning the Super Bowl encompasses “winning the NFC”.

This undermines all the logic of my trade ideas to the extent that it sets up lots of ways for the market to creatively “middle” you, just like the bettor who lays off a sports bet and gets careless about that half a point. But since all relative value trading deserves the error bars I’m dancing between, I feel a bit better having disclaimed them. Confidence sells, but when it comes to markets it’s craven. Something to keep in mind when you’re listening to your next podcast.

Of course, if earnings growth increased fast enough, equities don’t need to decline at all to maintain the same equity risk premium to bond yields.

That leaves us with two questions:

  • What spread do today’s earnings estimates offer?
  • How much more earnings would restore the 50 bp equity risk premium?

At SPX 7,637.79, the approximately $397 of earnings expected over the next twelve months gives us a 5.20% earnings yield. Against a 4.96% Treasury yield, that’s only a 24 bp spread.

To get 50 bps over Treasuries, we need a 5.46% earnings yield:

7,637.79 × 5.46% = $417 of annual EPS.

How close are we to earning that much?

The first-half figures total approximately $181 per share, up 39% from the same quarters in 2025. Second-half estimates total approximately $183. The forecast calls for roughly maintaining the first-half earnings level through the rest of 2026, which represents 26% growth over the second half of 2025. Historically high, but a slower pace vs the first half’s increase over the same period a year earlier.

For the full calendar years, consensus is $362 in 2026 and $417 in 2027. That’s another 15% growth after this year’s expected 32% increase. Together, those forecasts take annual earnings from $275 in 2025 to $417 in 2027—52% growth in two years. These figures come from the same FactSet earnings series.

Back to the valuation. The calendar-2027 estimate already gets us almost exactly to our $417 target and assumes EPS growth of 15%, much more reasonable compared to historical changes and far slower growth than 2026 experienced.

Despite the rise in yields and inflation, consensus equity pricing offers an earnings path that supports approximately the original 50 bp spread benchmark. It requires delivering the rest of 2026 plus “only” 15% subsequent growth next year.

The equity risk premium has been skinny, but EPS growth has delivered such that those risk premiums aren’t actually shrinking despite the continued outperformance of equities vs bonds. I wouldn’t call that bullish exactly, but it’s not bearish versus where we were 6 months ago. When I first started this investigation with the high energy prices and rising yields, I expected the “missing Sudoku number” of equity risk premium to look even more paltry, but sparkling earnings have bailed the multiples out.

Stay groovy

☮️


End Notes

Gasoline futures and CPI

Persistently high fuel prices burden consumers and businesses. But keeping the price high doesn’t mean inflation stays high. If gasoline rises from $2 to $3 and stays there, year-over-year inflation is positive until the comparison catches up. Comparing $3 with $3 gives zero inflation.

The spread benchmark

“Equity risk premium” here is shorthand for earnings yield minus the nominal ten-year Treasury yield. It is a valuation comparison, not a complete estimate of equities’ expected excess return. Earnings are not contractual interest payments and are not necessarily distributed to shareholders.

Comparing that earnings yield with the roughly 2% TIPS yield gives a 300 bp difference. That is a comparison with a real bond yield, not a separately calculated real equity risk premium. My mental shorthand for real equity returns is that they typically realize between 300 and 600 bps over the risk-free rate. Equity valuation is on the high side (ie low real earnings yield over CPI), but it has been for a while, as it has been priced for growth which has been repeatedly confirmed.

Market snapshot and forward calculation

The calculations hold SPX at 7,637.79 and the ten-year Treasury yield at 4.96%, using September 14 intraday observations from MarketWatch and Trading Economics. These are not closing prices.

Tracking 2026

FactSet lists Q1 2026 EPS of $80.99 and Q2 EPS of $100.28. Q1 is shown as actual; Q2 remains marked as estimated. The $181.27 first-half total therefore should not be described as entirely finalized.

Index composition

The S&P 500’s membership changes. Depending on how a historical series is constructed, year-over-year earnings changes can reflect additions, deletions and changes in index representation, as well as growth within businesses.

Comparing each period’s membership differs from comparing a fixed set of companies across both periods. The chart therefore describes the published index earnings series, not a constant-company measure of organic growth. No adjustment for composition has been made.

after this post you will be sizing bets in your head

One of the most important concepts in risk-taking is bet sizing. Which is unfortunate because people are quite bad at it, while the effort to be way above average is quite low.

A jarring and famous demonstration of this is the Haghani-Dewey Coin Flipping study, which showed how even college grads with business, economic, and technical backgrounds incinerated their capital or massively underperformed the expected profits presented to them by a game they knew was rigged in their favor.

You can read my synopsis in Bet Sizing Is Not Intuitive.

For a binary wager (ie win or lose), if you know the payoffs and the probability of winning, both of which were known to the participants, the solution is to use the Kelly Criterion.

The tragedy is that it is incredibly simple to compute and applies to many conventional gambles and decisions (the examples in the quiz will span various life situations!).

If something is both easy and widely relevant, it should be common knowledge. So let’s fix that today. I’ll show you how easy it is to use, and you’ll forever be able to do it in your head.

First, a succinct definition:

Kelly is the bet size, as a fraction of bankroll, that maximizes the long-run compounded growth rate of your wealth. It’s a mathematical solution to bet size that doesn’t seek to maximize expected profit per trial, but the size that optimally balances compounding rate and survival.

If you want to go deep on this, see Moontowerquant’s Kelly Criterion Resources, but today’s focus is on getting straight to usability.

We will use this formulation of Kelly because it’s general:

f* = p − q/b

where:

p = probability of winning

q = 1−p or probability of losing

b = the odds you’re getting → what you win divided by what you risk

The easiest way to learn it is just jump right in with a few worked examples:

Fair coin wager (even odds style bet)

p =50%

q= 50%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 50% – 50% / 1 = 0 → bet nothing, zero edge

Coin biased in your favor (even odds style bet)

p =60%

q= 40%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 60% – 40% / 1 = .20 → bet 20% of your bankroll

Roll a 6 on a die (underdog bet where you get odds)

p =1/6

q= 5/6

b =8 (for a $1 bet, you either lose a $1 or make an $8 profit)

f* = 1/6 – (5/6) / 8

f* =8/48 – 5/48 = 3/48 → bet 6.25% of your bankroll

If f* is 0 or negative, you have no edge, so not betting is prescribed

Sports moneyline (betting as a favorite where you lay odds)

A −200 favorite. You risk $2 to win $1, and the line implies 2/3, but you think it’s closer to 3 in 4.

p = 75%
q = 25%
b = 0.5 (getting 50% return on the amount you risk)

f* = 75% − 25% / 0.5 = 75% − 50% = 25% → bet 25% of your bankroll

Wait a minute, these are large bets?!!

If these bet sizes seem surprisingly large for the given advantages, then your senses are well-tuned. For most people, “full” Kelly is too big!

Kelly maximizes long-run growth on the assumption your probability is correct. Well, it probably isn’t because the world is messy. We can inject some humility by using a fraction of Kelly:

  • “Half Kelly” gives up about a quarter of the growth rate and roughly halves the drawdowns. If you invert that, you see that the Kelly scaling law means as you bet bigger, you get diminishing returns per unit of risk. Extrapolating that logic, betting more than “full Kelly” is incinerating compounded wealth even if the individual bet has positive EV.
  • “Quarter Kelly” or less is far more common in practice.

Please don’t let the equation scare you, it’s intuitive and easy to remember

Look at the equation again:

f* = p − q/b

It’s just “how often you win” minus “how often you lose.” It’s just that the second term incorporates the payoff. The loss term gets divided by b, which represents the return you collect when you’re right.

  • When b = 1, you’re getting even money. A 100% return. Dividing by 1 leaves q alone, and the equation collapses to pure hit rate: p − q. That’s the coin case where you bet $1 to make $1.
  • When b > 1, you’re getting long odds. The division shrinks the loss term. This is why the die works. You lose 5 out of 6 rolls. Straight subtraction says you’re down 66 cents on the dollar and should never play, but you’re paid 8-to-1, so that 5/6 becomes 5/48, and suddenly the 1/6 win percentage is the bigger number. Long odds forgive a bad hit rate.
  • When b < 1, you’re laying odds. Now division stretches the loss term. The moneyline: you only lose a quarter of the time, but at −200 each loss costs you two units to earn back one, so that 25% loss percentage behaves like 50%. Being right three times out of four barely clears the bar. Lay enough odds and even a very good record is a losing proposition.

The graphic shows how you can think of the odds (the denominator) as shrinking or inflating q, as you collapse your thinking to a comparison of p vs q.

A word on b

b trips people up because “odds” is loaded gambler jargon. A wider interpretation of b is that it’s a percent return.

It’s what you make divided by what you risk. Even money is b = 1: risk a dollar, make a dollar. That’s a 100% return on the amount at stake. 3-to-1 is b = 3, a 300% return. Laying −200 is b = 0.5 because if you risk two to make one, it’s a 50% return.

[Return is a profit, while multiples don’t subtract your initial risk. It’s the difference between “I 2x’d my money” vs “I made 100%” or “I 10x’d my money” vs “I made 900%”. The percent return is the multiple minus one because we subtract our initial risk.]

The reason it’s a return and not just “the odds” is that Kelly assumes a loss wipes out the whole stake. The denominator is always the same number: everything you put up. b is comparable across a coin, a die, and a moneyline because it’s the return on risk, always measured against a total loss.

b = 1 is a natural reference point. At 100% return, a win exactly cancels a loss, so you need to win more than half the time. The breakeven hit rate changes with b.

Set f* = 0 and you get p = 1/(1+b).

Read the table as a menu of the hit rates you’re allowed to have. At b = 24 you can be wrong 24 times out of 25 and still be flat. At b = 0.25 you can be right four out of five and still be flat

Applying to real life: when is Kelly the right tool?

Kelly needs a few inputs: a bankroll, a payoff you know, and a probability estimate.

Which of these is a Kelly problem?

  1. A prediction market contract trading at 30¢. You think it’s worth 45¢.
  2. You’re all-in-or-fold on the river with a read that you’re good 40% of the time, getting 3-to-1 from the pot.
  3. How much of your 401(k) to put in equities.
  4. Writing checks as an angel investor across 30 startups.
  5. Buying a weekly call on a biotech ahead of an FDA decision date.
  6. Whether to take the new job.
  7. Your buddy offers you 5-to-1 that it rains in Oakland tomorrow. The forecast says 30%.
  8. Buying homeowners insurance. The premium is clearly more than the expected loss — that’s how the insurer stays in business.
  9. Your neighbor doesn’t carry homeowners coverage. She banks the premium instead.
  10. Your auto policy offers a $500 deductible or a $2,500 deductible, for a $340/yr discount on the premium.
  11. You’ve got vested startup options. Exercising costs $40k out of pocket in strike, and you think there’s maybe a 15% chance the company gets somewhere that makes them worth $1M.
  12. A merger arb spread. Target’s at $46, deal price is $50, and it trades back to $38 if the deal breaks. You think it closes 90% of the time.
  13. Your agency spends 20 hours of unbilled time on a speculative pitch. You win about a quarter of them, and a win is worth 80 billable hours.

Solutions to Kelly Problems

  1. Yes. Cleanest case there is. Binary, known payoff, and the price provides b directly. Risk 30¢ to make 70¢, so b = 2.33. f* = 45% − 55%/2.33 = 21%.
  2. Yes. This is the canonical one. p = 40%, b = 3, f* = 40% − 60%/3 = 20% of your stack. The wrinkle is that in poker your stack isn’t really your bankroll. There’s a whole literature on pros using Kelly for bankroll management across sessions rather than for a single river decision.
  3. No. Not this version of it. Stock returns aren’t generally binary so there’s no p, q, or b. There’s a continuous analog called Merton’s Share, which is similarly rooted in reward vs variance. What Gamblers Can Teach the Buy-and-Hold Crowd can get you started.
  4. Sort of. The structure is right: repeated, roughly binary, long odds. The problem is that p is a guess and b is a bigger guess, and Kelly is violently sensitive to overestimating your edge. Garbage in, garbage out.
  5. Approximately. If you treat it as approve/reject it’s binary enough to size with. If the expiry aligns with the date such that you are betting strictly on the terminal intrinsic value, the option piece will inherit the binary modeling you imposed on the stock.
  6. No. The variables are too opaque.
  7. Yes. p = 30%, b = 5, f* = 30% − 70%/5 = 16%. Note, you’ll lose this bet more than twice as often as you win it, so your most likely scenario is losing 16%. You can shrink the Kelly fraction if this makes you uncomfortable.
  8. Wrong side of the equation. Run f* on this and you get a negative number, because you’re buying a negative-EV bet.
  9. Yes. It’s the same policy, so notice that the bet only exists on the insurer’s side of it! Every year your neighbor doesn’t buy, she collects a premium and writes a tail. Rebuild cost $500k, premium $3,000, call it a 1-in-500 chance of a total loss.

    p = 99.8%
    q = 0.2%
    b = 3,000 / 500,000 = 0.006 (risk $500k to win $3,000)

    f* = 99.8% − 0.2%/0.006 = 99.8% − 33.3% = 66.5%

    Positive, as expected since insurers price premiums well above fair value. In this case, her bet size is the $500k house. If the house is most of her net worth, she’s at 100% on a bet capped at 66%. Rather than overbet, she should buy the policy. If she’s worth $5M, she’s betting 10% when she’s allowed 66%, which puts her near quarter Kelly (~14%), and she could skip the insurance. If she’s worth $1mm, it’s a 50% bet, which is more than half Kelly. I’d say take the insurance but I’m a wimp. There are other considerations (would she have the liquidity to rebuild the home or maybe taking the insurance with a high deductible is a better fit), but just doing this exercise gives you a sense of how risky or conservative your choices are relative to the bet share that maximizes long-term wealth.

  10. Yes. Raising the deductible is like you writing a $2,000 policy and collecting $340 a year for it. You’re the insurer again, so work out what you need to believe. Take the high deductible and save $340. Have a claim, and you’re out $2,000 more, but you already banked the $340, so the loss is $1,660.

    b = 340 / 1,660 = 0.205

    f* = p − q/0.205 = p − 4.88q

    Set that to zero, and you get p = 4.88q, which, with p + q = 1, means q = 17%. Your breakeven is a claim every 5.9 years. Anything less frequent and you’re the one with the edge.

    Let’s say real-world collision frequency is more like 6%. So p = 94%:

    f* = 94% − 6%/0.205 = 94% − 29.3% = 64.7%

    Which says risk at most ~65% of your bankroll. The risk here is $1,660. That clears as long as you have about $2,600 in liquid savings, which is to say the sizing check is trivially satisfied for almost everyone so you should generally opt for the higher deductible. For quarter Kelly, we’d need savings of $1,660/(.647 * .25) = $10,262.

  11. Approximately. It’s not truly binary, but if you frame it in a way where you are comfortable with the no consolation prize of a medium outcome, you can see it as paying $40k for a 15% shot at $1M. b = 24, so f* = 15% − 85%/24 = 11.5% of your liquid net worth. This is quite sensitive to your estimate of p of course.
  12. Yes, a classic example of binary-type risk in markets. You risk $8 to make $4, so b = 0.5 — you’re laying odds, same as the moneyline. f* = 90% − 10%/0.5 = 70%. That number is only as good as the 90%. Revise p to 75% and f* is 25%.
  13. Yes, in a subtle way! Your bankroll is capacity, not cash. b = 80/20 = 4, so f* = 25% − 75%/4 = 6.25%. Twenty hours has to be 6% of what you’re working with, which means you can’t run this pitch out of a 100-hour month. If you’re the manager, you can put it in dollar terms by converting to wages.

Finally, I strongly recommend William Poundstone’s book Fortune’s Formula: The Untold Story of the Scientific Betting System That Beat the Casinos and Wall Street

Description:

In 1956, two Bell Labs scientists discovered the scientific formula for getting rich. One was mathematician Claude Shannon, neurotic father of our digital age, whose genius is ranked with Einstein’s. The other was John L. Kelly Jr., a Texas-born, gun-toting physicist. Together they applied the science of information theory—the basis of computers and the Internet—to the problem of making as much money as possible, as fast as possible.

Shannon and MIT mathematician Edward O. Thorp took the “Kelly formula” to Las Vegas. It worked. They realized that there was even more money to be made in the stock market. Thorp used the Kelly system with his phenomenally successful hedge fund, Princeton-Newport Partners. Shannon became a successful investor, too, topping even Warren Buffett’s rate of return. Fortune’s Formula traces how the Kelly formula sparked controversy even as it made fortunes at racetracks, casinos, and trading desks. It reveals the dark side of this alluring scheme, which is founded on exploiting an insider’s edge.

Shannon believed it was possible for a smart investor to beat the market—and William Poundstone’s Fortune’s Formula will convince you that he was right.

And this is from my notes, Insights From Fortune’s Formula:

a gripping narrative full of 20th century trivia that ties together the birth of information theory, some of the greatest scientific minds of the 1900s, the rise of quantitative finance, and the role of organized crime. These topics come alive in a fresh, memorable way when discovered through the lens of its colorful characters.

It chronicles the history of the efficient market hypothesis (MIT, U Chicago, Paul Samuelson). You can organize its conclusion around this excerpt:

There is much truth in the efficient market hypothesis. The controversy has always been over just how far the claim can be pressed. Asking whether markets are efficient is like asking whether the world is round. The best way to answer depends on the expectations and sophistication of the questioner. If someone is asking whether the world is round or flat, as fifteenth-century Europeans might have asked, then “round” is a better answer. If someone knows that and is asking whether the earth is a geometrically perfect sphere, the answer is no.

Moontower #327

In this issue:

  • slow burn
  • mental bet sizing for everyone
  • sizing for real-life decisions

Friends,

Riddle me this:

Reading moontower makes me feel…

⏹️Smart

⏹️Stupid

(read the article on substack to actually submit an answer)

Over the years, I’ve administered 2 reader polls to take the pulse of why you bother reading this newsletter. Both times I’ve asked the above question but with an open-ended fill-in-the-blank template. The answers come back in all shapes and sizes, but most of them can be boiled down to “smart” or “stupid”.

My reactions to the “stupid” responses:

My writing has a style that can be, at times, counterproductive. It’s often unfiltered and parenthetical. My wife constantly tells me that if I write more clearly, I’d have a bigger audience, and she’s, of course, right. The writing is not optimized for clarity. Scott H Young is my canonical example of such writing. I’m a big fan, but my lack of self-control forbids this common-sense approach.

My writing is aspirationally in the spirit of a Pixar movie where it serves 2 masters: the child and the adult who brought them to the theater. I’m trying to teach relatively basic or intermediate concepts while leaving Easter eggs for readers who arrive with more finance/trading context. This will leave the more basic readers feeling like they are missing a joke.

Fellow trader and writer RobotJames, once told me that Moontower is a “slow burn”. There’s no single banger that you could send to someone and say “this is why you should read this”, but reading it week in and week out leads to something more than the sum of its parts. My lament over never having a truly massive hit notwithstanding, I will not complain about the “slow burn” status. Many of my favorite things often took time to warm up to. Acquired tastes are rewards for crossing rugged terrain. (Except when they’re status cosplays. Our wants are even opaque to ourselves. There’s a thin line between appreciation and snobbery and its placement is up to the onlookers, not the admirer.)

To feel stupid when reading moontower would be natural. There’s technical or domain-specific minutiae mixed in with the meat and potatoes. This is further exacerbated by the cohort of readers who only show up for the personal sections and whose interests lie especially far from any of the material I cover. But if I read the blog of a doctor friend because it’s a way to follow along their life, the object-level material is gonna make me feel stupid.

And of course sometimes you feel stupid because I fail. The sequence of words fails to unlock the concept. Today I’m actively trying not to f this up. I want to teach you a neat topic. My goal is to really cement it for you. To make it accessible in real life for you in the same way multiplication tables are at your disposal.

If you don’t regularly read Money Angle, maybe see if today works for you. It’s long, but that’s because we are going to drill the concept. If you succeed in walking away with a new mental tool to practice, then the ROI on the reading time will look like a steal. If you’re not interested, then I’ll channel one of my college roommates whenever I tried to bail on hitting the gym, ”That’s cool man, I’ll just take you off my list of successful people for today” and leave you one of the OG memes that deserves annual visitation.

 

Commencing demystifications now…

Money Angle

One of the most important concepts in risk-taking is bet sizing. Which is unfortunate because people are quite bad at it, while the effort to be way above average is quite low.

A jarring and famous demonstration of this is the Haghani-Dewey Coin Flipping study, which showed how even college grads with business, economic, and technical backgrounds incinerated their capital or massively underperformed the expected profits presented to them by a game they knew was rigged in their favor.

You can read my synopsis in Bet Sizing Is Not Intuitive.

For a binary wager (ie win or lose), if you know the payoffs and the probability of winning, both of which were known to the participants, the solution is to use the Kelly Criterion.

The tragedy is that it is incredibly simple to compute and applies to many conventional gambles and decisions (the examples in the quiz will span various life situations!).

If something is both easy and widely relevant, it should be common knowledge. So let’s fix that today. I’ll show you how easy it is to use, and you’ll forever be able to do it in your head.

First, a succinct definition:

Kelly is the bet size, as a fraction of bankroll, that maximizes the long-run compounded growth rate of your wealth. It’s a mathematical solution to bet size that doesn’t seek to maximize expected profit per trial, but the size that optimally balances compounding rate and survival.

If you want to go deep on this, see Moontowerquant’s Kelly Criterion Resources, but today’s focus is on getting straight to usability.

We will use this formulation of Kelly because it’s general:

f* = p − q/b

where:

p = probability of winning

q = 1−p or probability of losing

b = the odds you’re getting → what you win divided by what you risk

The easiest way to learn it is just jump right in with a few worked examples:

Fair coin wager (even odds style bet)

p =50%

q= 50%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 50% – 50% / 1 = 0 → bet nothing, zero edge

Coin biased in your favor (even odds style bet)

p =60%

q= 40%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 60% – 40% / 1 = .20 → bet 20% of your bankroll

Roll a 6 on a die (underdog bet where you get odds)

p =1/6

q= 5/6

b =8 (for a $1 bet, you either lose a $1 or make an $8 profit)

f* = 1/6 – (5/6) / 8

f* =8/48 – 5/48 = 3/48 → bet 6.25% of your bankroll

If f* is 0 or negative, you have no edge, so not betting is prescribed

Sports moneyline (betting as a favorite where you lay odds)

A −200 favorite. You risk $2 to win $1, and the line implies 2/3, but you think it’s closer to 3 in 4.

p = 75%
q = 25%
b = 0.5 (getting 50% return on the amount you risk)

f* = 75% − 25% / 0.5 = 75% − 50% = 25% → bet 25% of your bankroll

Wait a minute, these are large bets?!!

If these bet sizes seem surprisingly large for the given advantages, then your senses are well-tuned. For most people, “full” Kelly is too big!

Kelly maximizes long-run growth on the assumption your probability is correct. Well, it probably isn’t because the world is messy. We can inject some humility by using a fraction of Kelly:

  • “Half Kelly” gives up about a quarter of the growth rate and roughly halves the drawdowns. If you invert that, you see that the Kelly scaling law means as you bet bigger, you get diminishing returns per unit of risk. Extrapolating that logic, betting more than “full Kelly” is incinerating compounded wealth even if the individual bet has positive EV.
  • “Quarter Kelly” or less is far more common in practice.

Please don’t let the equation scare you, it’s intuitive and easy to remember

Look at the equation again:

f* = p − q/b

It’s just “how often you win” minus “how often you lose.” It’s just that the second term incorporates the payoff. The loss term gets divided by b, which represents the return you collect when you’re right.

  • When b = 1, you’re getting even money. A 100% return. Dividing by 1 leaves q alone, and the equation collapses to pure hit rate: p − q. That’s the coin case where you bet $1 to make $1.
  • When b > 1, you’re getting long odds. The division shrinks the loss term. This is why the die works. You lose 5 out of 6 rolls. Straight subtraction says you’re down 66 cents on the dollar and should never play, but you’re paid 8-to-1, so that 5/6 becomes 5/48, and suddenly the 1/6 win percentage is the bigger number. Long odds forgive a bad hit rate.
  • When b < 1, you’re laying odds. Now division stretches the loss term. The moneyline: you only lose a quarter of the time, but at −200 each loss costs you two units to earn back one, so that 25% loss percentage behaves like 50%. Being right three times out of four barely clears the bar. Lay enough odds and even a very good record is a losing proposition.

The graphic shows how you can think of the odds (the denominator) as shrinking or inflating q, as you collapse your thinking to a comparison of p vs q.

A word on b

b trips people up because “odds” is loaded gambler jargon. A wider interpretation of b is that it’s a percent return.

It’s what you make divided by what you risk. Even money is b = 1: risk a dollar, make a dollar. That’s a 100% return on the amount at stake. 3-to-1 is b = 3, a 300% return. Laying −200 is b = 0.5 because if you risk two to make one, it’s a 50% return.

[Return is a profit, while multiples don’t subtract your initial risk. It’s the difference between “I 2x’d my money” vs “I made 100%” or “I 10x’d my money” vs “I made 900%”. The percent return is the multiple minus one because we subtract our initial risk.]

The reason it’s a return and not just “the odds” is that Kelly assumes a loss wipes out the whole stake. The denominator is always the same number: everything you put up. b is comparable across a coin, a die, and a moneyline because it’s the return on risk, always measured against a total loss.

b = 1 is a natural reference point. At 100% return, a win exactly cancels a loss, so you need to win more than half the time. The breakeven hit rate changes with b.

Set f* = 0 and you get p = 1/(1+b).

Read the table as a menu of the hit rates you’re allowed to have. At b = 24 you can be wrong 24 times out of 25 and still be flat. At b = 0.25 you can be right four out of five and still be flat

Applying to real life: when is Kelly the right tool?

Kelly needs a few inputs: a bankroll, a payoff you know, and a probability estimate.

Which of these is a Kelly problem?

  1. A prediction market contract trading at 30¢. You think it’s worth 45¢.
  2. You’re all-in-or-fold on the river with a read that you’re good 40% of the time, getting 3-to-1 from the pot.
  3. How much of your 401(k) to put in equities.
  4. Writing checks as an angel investor across 30 startups.
  5. Buying a weekly call on a biotech ahead of an FDA decision date.
  6. Whether to take the new job.
  7. Your buddy offers you 5-to-1 that it rains in Oakland tomorrow. The forecast says 30%.
  8. Buying homeowners insurance. The premium is clearly more than the expected loss — that’s how the insurer stays in business.
  9. Your neighbor doesn’t carry homeowners coverage. She banks the premium instead.
  10. Your auto policy offers a $500 deductible or a $2,500 deductible, for a $340/yr discount on the premium.
  11. You’ve got vested startup options. Exercising costs $40k out of pocket in strike, and you think there’s maybe a 15% chance the company gets somewhere that makes them worth $1M.
  12. A merger arb spread. Target’s at $46, deal price is $50, and it trades back to $38 if the deal breaks. You think it closes 90% of the time.
  13. Your agency spends 20 hours of unbilled time on a speculative pitch. You win about a quarter of them, and a win is worth 80 billable hours.

Money Angle For Masochists

Solutions to Kelly Problems

  1. Yes. Cleanest case there is. Binary, known payoff, and the price provides b directly. Risk 30¢ to make 70¢, so b = 2.33. f* = 45% − 55%/2.33 = 21%.
  2. Yes. This is the canonical one. p = 40%, b = 3, f* = 40% − 60%/3 = 20% of your stack. The wrinkle is that in poker your stack isn’t really your bankroll. There’s a whole literature on pros using Kelly for bankroll management across sessions rather than for a single river decision.
  3. No. Not this version of it. Stock returns aren’t generally binary so there’s no p, q, or b. There’s a continuous analog called Merton’s Share, which is similarly rooted in reward vs variance. What Gamblers Can Teach the Buy-and-Hold Crowd can get you started.
  4. Sort of. The structure is right: repeated, roughly binary, long odds. The problem is that p is a guess and b is a bigger guess, and Kelly is violently sensitive to overestimating your edge. Garbage in, garbage out.
  5. Approximately. If you treat it as approve/reject it’s binary enough to size with. If the expiry aligns with the date such that you are betting strictly on the terminal intrinsic value, the option piece will inherit the binary modeling you imposed on the stock.
  6. No. The variables are too opaque.
  7. Yes. p = 30%, b = 5, f* = 30% − 70%/5 = 16%. Note, you’ll lose this bet more than twice as often as you win it, so your most likely scenario is losing 16%. You can shrink the Kelly fraction if this makes you uncomfortable.
  8. Wrong side of the equation. Run f* on this and you get a negative number, because you’re buying a negative-EV bet.
  9. Yes. It’s the same policy, so notice that the bet only exists on the insurer’s side of it! Every year your neighbor doesn’t buy, she collects a premium and writes a tail. Rebuild cost $500k, premium $3,000, call it a 1-in-500 chance of a total loss.

    p = 99.8%
    q = 0.2%
    b = 3,000 / 500,000 = 0.006 (risk $500k to win $3,000)

    f* = 99.8% − 0.2%/0.006 = 99.8% − 33.3% = 66.5%

    Positive, as expected since insurers price premiums well above fair value. In this case, her bet size is the $500k house. If the house is most of her net worth, she’s at 100% on a bet capped at 66%. Rather than overbet, she should buy the policy. If she’s worth $5M, she’s betting 10% when she’s allowed 66%, which puts her near quarter Kelly (~14%), and she could skip the insurance. If she’s worth $1mm, it’s a 50% bet, which is more than half Kelly. I’d say take the insurance but I’m a wimp. There are other considerations (would she have the liquidity to rebuild the home or maybe taking the insurance with a high deductible is a better fit), but just doing this exercise gives you a sense of how risky or conservative your choices are relative to the bet share that maximizes long-term wealth.

  10. Yes. Raising the deductible is like you writing a $2,000 policy and collecting $340 a year for it. You’re the insurer again, so work out what you need to believe. Take the high deductible and save $340. Have a claim, and you’re out $2,000 more, but you already banked the $340, so the loss is $1,660.

    b = 340 / 1,660 = 0.205

    f* = p − q/0.205 = p − 4.88q

    Set that to zero, and you get p = 4.88q, which, with p + q = 1, means q = 17%. Your breakeven is a claim every 5.9 years. Anything less frequent and you’re the one with the edge.

    Let’s say real-world collision frequency is more like 6%. So p = 94%:

    f* = 94% − 6%/0.205 = 94% − 29.3% = 64.7%

    Which says risk at most ~65% of your bankroll. The risk here is $1,660. That clears as long as you have about $2,600 in liquid savings, which is to say the sizing check is trivially satisfied for almost everyone so you should generally opt for the higher deductible. For quarter Kelly, we’d need savings of $1,660/(.647 * .25) = $10,262.

  11. Approximately. It’s not truly binary, but if you frame it in a way where you are comfortable with the no consolation prize of a medium outcome, you can see it as paying $40k for a 15% shot at $1M. b = 24, so f* = 15% − 85%/24 = 11.5% of your liquid net worth. This is quite sensitive to your estimate of p of course.
  12. Yes, a classic example of binary-type risk in markets. You risk $8 to make $4, so b = 0.5 — you’re laying odds, same as the moneyline. f* = 90% − 10%/0.5 = 70%. That number is only as good as the 90%. Revise p to 75% and f* is 25%.
  13. Yes, in a subtle way! Your bankroll is capacity, not cash. b = 80/20 = 4, so f* = 25% − 75%/4 = 6.25%. Twenty hours has to be 6% of what you’re working with, which means you can’t run this pitch out of a 100-hour month. If you’re the manager, you can put it in dollar terms by converting to wages.

Finally, I strongly recommend William Poundstone’s book Fortune’s Formula: The Untold Story of the Scientific Betting System That Beat the Casinos and Wall Street

Description:

In 1956, two Bell Labs scientists discovered the scientific formula for getting rich. One was mathematician Claude Shannon, neurotic father of our digital age, whose genius is ranked with Einstein’s. The other was John L. Kelly Jr., a Texas-born, gun-toting physicist. Together they applied the science of information theory—the basis of computers and the Internet—to the problem of making as much money as possible, as fast as possible.

Shannon and MIT mathematician Edward O. Thorp took the “Kelly formula” to Las Vegas. It worked. They realized that there was even more money to be made in the stock market. Thorp used the Kelly system with his phenomenally successful hedge fund, Princeton-Newport Partners. Shannon became a successful investor, too, topping even Warren Buffett’s rate of return. Fortune’s Formula traces how the Kelly formula sparked controversy even as it made fortunes at racetracks, casinos, and trading desks. It reveals the dark side of this alluring scheme, which is founded on exploiting an insider’s edge.

Shannon believed it was possible for a smart investor to beat the market—and William Poundstone’s Fortune’s Formula will convince you that he was right.

And this is from my notes, Insights From Fortune’s Formula:

a gripping narrative full of 20th century trivia that ties together the birth of information theory, some of the greatest scientific minds of the 1900s, the rise of quantitative finance, and the role of organized crime. These topics come alive in a fresh, memorable way when discovered through the lens of its colorful characters.

It chronicles the history of the efficient market hypothesis (MIT, U Chicago, Paul Samuelson). You can organize its conclusion around this excerpt:

There is much truth in the efficient market hypothesis. The controversy has always been over just how far the claim can be pressed. Asking whether markets are efficient is like asking whether the world is round. The best way to answer depends on the expectations and sophistication of the questioner. If someone is asking whether the world is round or flat, as fifteenth-century Europeans might have asked, then “round” is a better answer. If someone knows that and is asking whether the earth is a geometrically perfect sphere, the answer is no.

Stay groovy

☮️


Moontower Weekly Recap

Moontower #326

In this issue:

  • energy as the ultimate currency
  • “middle manager” fiction

Friends,

A year ago I wrote a piece called Capitalism Is a Temporary Condition. Buried in it was a thought:

An investor in an age of acceleration watches as commodity prices, real things, hover in familiar territory while “future cash flows discounted” reward attention — in some cases because of optimistic stories, in some cases because the company exchanges cash for BTC (an act which adds no economic value), and in some cases for nostalgic lolz.

Today, it’s embarrassing to think of actual value. The most basic form is stored energy**. An obvious example is oil. Slightly more abstract —a building is a collection of atoms that took energy to arrange and provides utility. It’s proof of work. Same with a good reputation.

The ** footnote:

I increasingly think that a durable concept of a risk-free or least-risky rate is more likely to come from Vaclav Smil than U Chicago.

I knew about Smil from a profile piece that celebrated his suffer-no-fools rigor in making sense of the world.

I’m finally getting around to How the World Really Works, and I dig how quickly it’s zeroing in on what I anticipated to be the biggest muscle movement of all. I’ll let it unfold as he does the book’s intro.

Smil asks us to imagine an alien probe watching Earth, programmed to wake up whenever something important changes in the way energy is captured or used.

For a very long time, it sleeps.

Then life figures out photosynthesis: sunlight can be captured and stored as chemical energy. Much later, humans learn to control fire. Then agriculture. Draft animals. Wind and water.

The gaps are enormous. Billions of years. Millions. Thousands.

Then they start collapsing.

Coal gives humans access to immense stores of ancient solar energy. You can see a parallel to wealth in that phrase alone. Steam will turn that energy into mechanical work, replacing what he calls prime mover energy (human and animal labor). Oil makes enormous amounts of energy portable while electricity makes it transmissible and almost infinitely adaptable.

Smil puts numbers on this.

In 1800, the world had access to roughly 0.05 gigajoules of useful energy per person each year. By 1900 it was 2.7. By 1950, about 10. By 2000, 28. By 2020, roughly 34.

In a little over two centuries, useful energy per person increased nearly 700-fold.

The average person today commands roughly the energy equivalent of 60 adults working continuously, day and night. In affluent countries, the equivalent is closer to 200–240 permanent laborers.

And the averages hide enormous differences.

Measured in primary energy, someone in one of the world’s poorest countries may use less than 10 GJ a year. India is around 20–30. The global average is roughly 75–80. China is above 100. Western Europe and Japan generally sit around 120–160. The United States is closer to 250–280.

It sounds a lot like how GDP per capita is distributed. We measure GDP in dollars, but money can be printed, thus distorting its exchange rate versus the stored value of work. In that sense, energy may be civilization’s only uncorruptible currency.

If we continue pulling on the thread, wealth might be thought of as an account from which we can spend to hold entropy at bay.

  • A 72-degree house in Scottsdale.
  • Mango salad in during a Minnesota January.
  • Clean water pumped up the Hollywood Hills.
  • A sterile operating room.
  • Skyscrapers defying gravity.
  • Data centers (I couldn’t resist).

Much of the AI commotion revolves around what it means for humans to directly turn electricity into intelligence. That the product is intelligence which can recursively accelerate knowledge instinctively unsettles us because it violates our sensibilities around balance. Like energy is somehow not being conserved, leading to the type of divergence we associate with chain reactions.

I’m actually struggling to put my finger on the right analogy, which validates another point Smil makes. He argues that at precisely the moment when we have become capable of commanding extraordinary quantities of energy, most of us have become almost completely detached from how any of it happens.

Smil calls this our “comprehension deficit.”

Modernity is a black box. Light comes from a switch, and meat comes from a carton lined with a maxi pad. It’s all so effortless. Students of economics read I, Pencil to appreciate how capitalism’s profit-motive and competition actually lead to complex chains of cooperation. The side effect of the story is a (demoralizing?) inference that the world is so invisibly complicated now that you should feel good in your choice of major because it’s strategic to be a symbol-pusher since details are hopelessly infinite. (Sorry, did I go to far here? I’m an econ major, so it’s kind of like Chappelle using the n-word. Kinda? Nevermind.)

Smil’s punchy take on specialization:

You could meet real Renaissance men on Florence’s Piazza Signoria in 1500, but not for too long after that…By the middle of the 18th century, Diderot and d’Alembert could still assemble a group capable of summing up much of their era’s knowledge in the Encyclopédie…In 1872, a century after the appearance of the last volume of the Encyclopédie, any collection of knowledge had to resort to the superficial treatment of a rapidly expanding range of topics…today, it is impossible to sum up our understanding even within narrowly circumscribed specialties…experts in particle physics would find it very hard to understand even the first page of a new research paper in viral immunology …Highly specialized branches of modern science have become so arcane” that many practitioners must train into their thirties “in order to join the new priesthood.

All I’ve done is read the introduction to the book, but the “comprehension deficit”, which he’s fixin’ to remedy with respect to how the world works (at least from a physics/chemistry macro point of view), is just a fascinating idea of its own. Before the printing process and hyper-connectedness, it was hard to learn much of what was a relatively small body of knowledge, but since then, even though you could learn more in a lifetime, it was a smaller percent of what could be known.

But that I can sit here in an air-conditioned house on a 90-degree day typing contemplations about abstractions like how wealth is really the temporary abatement of nature’s randomness and how the accounting of such wealth should be denominated in units of work, which is traditionally measured by money and its continuously leaking exchange rate to said work is itself a wild collective achievement fueled by eons of carbon remnants from lives I never knew.

I guess I feel like I owe it to the universe to be curious about its ways even if it’s a Sisyphean fantasy to close the comprehension deficit. Looking forward to continuing:

Article content

A thought worth sharing more widely from hemispheres this week’s post about discretionary trading:

In this AI era, brains and consciousness are major topics du jour. I increasingly see us as technology centaurs with both a meat and silicon brain mediating structured and unstructured data, which in turn can be private or public.

On the private side, I have Claude connected to:

  • my personal knowledge management or PKM (Notion)
  • MCP to moontower.ai data
  • git repos across work and personal
  • email (which means it can access all my writing)
  • project management software (Linear)
  • Google drive
  • Google calendar

And of course, the entire world of public info lives in the model’s training data and ability to go online.

You can be taking a walk with voice mode on spec’ing a prototype that will connect to data with nearly any context you want to provide from your private knowledge, the internet, or what’s in your head that moment.

The future is here for anyone who feels haunted by dozens of ideas a day that they previously couldn’t or wouldn’t act on.

And a discretionary trader is nobody if not someone who starts 50 sentences a day with “I wonder what happens if…”

Related:

Article content

Money Angle

I ask for your indulgence, for what follows is an informal word-wall connecting conversations I’ve had with some W2 friends.

Pick your favorite gospel on disruption. Schumpeter’s creative destruction, Christensen’s innovator’s dilemma. Most companies, sometimes even industries, will eventually recognize themselves as a melting ice cube.

In a great talk back in 2009, media investor Peter Chernin warned:

One of the things I always used to say to the people who work for me is that you can’t protect your business, and your job is not to protect your business. Your job isn’t to protect. Your job is to maximize your business at any given time, but your real job is to grow new businesses faster than the old ones decay.

He uses record labels as the cautionary example. They set out to protect their business, but “to the degree you’re going to try and do it you’re going to get killed because technology is going to liberate audiences in such a way that they’re going to get what they want regardless.”

It feels like there are only bad choices for companies that are now run-off businesses. Re-investment feels too speculative, but panic will preclude any chance of a soft or at least dignified landing. Navigating these moments well is a sign of grace, but if grace were easy we’d call it something else.

Instead we get to witness corporate pathology.

The ice cube is melting. But it started as a very big ice cube. Big enough to bridge a middle-aged middle-manager’s soft landing into an early retirement if they can claim a larger share of a shrinking cube. Welcome to corporate hunger games on steroids.

The middle manager is a risk-averse incrementalist. That’s WHY they are middle managers. They will do no such thing as grow new businesses faster than the old ones decay. They need an accomplice to secure their spot on the life raft. This accomplice comes from an unexpected place. The partners.

You see, entrepreneurs have deranged brain chemistry. My friend Jason Buck and I were drinking coffee in my backyard on Wednesday when he described an entrepreneur as someone who works 100 hours a week for themselves to not have to work an hour for someone else.

The entrepreneur, the founder, the partner is a delusional optimist. A melting ice cube can be refrozen if you just find the right segment to sell to, the one that somehow eluded you when things were going well. The middle manager latches on to this hope.

He crafts a business plan to do things “radically” differently. Of course, radical in this context is a mere gesture in comparison to the extinction-level shift in the business environment. The partners, never ones to back down from a fight, support this can-do attitude from the manager who stepped up.

The manager, energized and enabled, moves to consolidate influence. By promoting their plan, any plan, doomed as it is, they portray unsupportive colleagues as quitters by virtue of their dissent.

“If you’re not on board with my [dumb] plan, you clearly hate the company and are not a team player.”

There’s no serious appraisal of the dissenting argument’s merit. And that’s probably because the dissenting argument is “Umm, we’re f’d, so we should focus on retention, not growth, because [insert analysis that actually makes sense]”. Any analysis that maximizes EV in a losing game will be hard to sell against a bad strategy employed with hope. We gamble to get back to even and we’re risk-averse when we’re ahead.

Our middle-manager doesn’t just knife out the dissenters. They fully larp optimism with expansion. They hire loyalists, veterans of this obsolete-but-not-yet-obviously-so strategy. The loyalists are relieved. They, themselves, were cooked seeing the same writing-on-the-wall, but they just got thrown a lifeline by the last firm running the old plays. The only plays they’re familiar with. The reunion with their middle-manager buddy is as predictable as you expect, complete with brown liquor and nostalgia for those nights during training when they were chasing tail in Murray Hill and scarfing khati roll together if they struck out that night. Ah, Indian food before bed. To be young again.

The loyalist cluster is pragmatic. Every year of health insurance and private school tuition extends the polyester harmony that is their home life. But for the middle manager, the loyalist cluster is strategic. He’s like an old piece of tile covering himself in linoleum. It’s all wrong but a bigger nuisance to remove. Become hard to kill by entangling yourself deeper in a hierarchy that you created and from which you are a convenient buffer between partners and minions whose names they never want to know.

The misalignment, politics, and the waste of human life force in the name of self-preservation of an artificial environment (that’s a big one to unpack, but y’all might have to come have coffee or maybe something a bit stiffer for that convo) are regrettable.

But you know what I find worse?

That our middle-manager protagonist was not as cunning as he seems. That his primary offense was stupidity. That he sincerely mistook the situation for something he could fix instead of a noble impulse towards calculated self-preservation.

Stupidity is uncivilized. It’s unpredictable. “Say what you want about the tenets of national socialism, at least it’s an ethos. These guys are nihilists.” That’s how I feel about our manager if he is stupid. That’s a barbarian. It might be unpleasant, but you can reason with the merely deceptive.

I’m not sure if our middle-manager character (whose likeness to any real-life individual is purely coincidental) is sincere and stupid or just trying to survive, but those accomplices on high, blinded by desperate optimism, channeled the old guy at the club instead of having the courage to reinvent or the grace to just go home.

Then again, being who they are is what got them to giant ice cube status in the first place.

Money Angle For Masochists

A recommendation

Besides writing, part of my role here is to find good stuff to share with you.

One of my favorite online discoveries this year has been the X account of @SowingAlphaSeed. It’s an anonymous account that goes by “Farmer”.

The bio says:

Retail investor w/ a stretch goal of 2 Sharpe + 30% CAGR. Please DM me if you have fund recommendations. Live portfolio and track record on my website.

The main draw is that this account has steadily been learning and doing “in public”. It’s a constant source of interesting and smart ideas. It’s also focused on what it says in the bio, so the percentage of useful posts to total posts is extremely high.

Tip of the hat to Farmer (who I never met and don’t know).


An observation:

1-year collars in MU are close to the cheapest they’ve in the past year. If you’re bullish you can budget a trade that can offer multiples of return per dollar bet. If MU already made you rich, the cost to hedge is actuarially small.

Then again, you didn’t get rich by hedging amirite

This week:

Article content

Zooming in:

Article content
moontower.ai
Article content
moontower.ai

Moontower Weekly Recap

the “middle manager”

I ask for your indulgence, for what follows is an informal word-wall connecting conversations I’ve had with some W2 friends.

Pick your favorite gospel on disruption. Schumpeter’s creative destruction, Christensen’s innovator’s dilemma. Most companies, sometimes even industries, will eventually recognize themselves as a melting ice cube.

In a great talk back in 2009, media investor Peter Chernin warned:

One of the things I always used to say to the people who work for me is that you can’t protect your business, and your job is not to protect your business. Your job isn’t to protect. Your job is to maximize your business at any given time, but your real job is to grow new businesses faster than the old ones decay.

He uses record labels as the cautionary example. They set out to protect their business, but “to the degree you’re going to try and do it you’re going to get killed because technology is going to liberate audiences in such a way that they’re going to get what they want regardless.”

It feels like there are only bad choices for companies that are now run-off businesses. Re-investment feels too speculative, but panic will preclude any chance of a soft or at least dignified landing. Navigating these moments well is a sign of grace, but if grace were easy we’d call it something else.

Instead we get to witness corporate pathology.

The ice cube is melting. But it started as a very big ice cube. Big enough to bridge a middle-aged middle-manager’s soft landing into an early retirement if they can claim a larger share of a shrinking cube. Welcome to corporate hunger games on steroids.

The middle manager is a risk-averse incrementalist. That’s WHY they are middle managers. They will do no such thing as grow new businesses faster than the old ones decay. They need an accomplice to secure their spot on the life raft. This accomplice comes from an unexpected place. The partners.

You see, entrepreneurs have deranged brain chemistry. My friend Jason Buck and I were drinking coffee in my backyard on Wednesday when he described an entrepreneur as someone who works 100 hours a week for themselves to not have to work an hour for someone else.

The entrepreneur, the founder, the partner is a delusional optimist. A melting ice cube can be refrozen if you just find the right segment to sell to, the one that somehow eluded you when things were going well. The middle manager latches on to this hope.

He crafts a business plan to do things “radically” differently. Of course, radical in this context is a mere gesture in comparison to the extinction-level shift in the business environment. The partners, never ones to back down from a fight, support this can-do attitude from the manager who stepped up.

The manager, energized and enabled, moves to consolidate influence. By promoting their plan, any plan, doomed as it is, they portray unsupportive colleagues as quitters by virtue of their dissent.

“If you’re not on board with my [dumb] plan, you clearly hate the company and are not a team player.”

There’s no serious appraisal of the dissenting argument’s merit. And that’s probably because the dissenting argument is “Umm, we’re f’d, so we should focus on retention, not growth, because [insert analysis that actually makes sense]”. Any analysis that maximizes EV in a losing game will be hard to sell against a bad strategy employed with hope. We gamble to get back to even and we’re risk-averse when we’re ahead.

Our middle-manager doesn’t just knife out the dissenters. They fully larp optimism with expansion. They hire loyalists, veterans of this obsolete-but-not-yet-obviously-so strategy. The loyalists are relieved. They, themselves, were cooked seeing the same writing-on-the-wall, but they just got thrown a lifeline by the last firm running the old plays. The only plays they’re familiar with. The reunion with their middle-manager buddy is as predictable as you expect, complete with brown liquor and nostalgia for those nights during training when they were chasing tail in Murray Hill and scarfing khati roll together if they struck out that night. Ah, Indian food before bed. To be young again.

The loyalist cluster is pragmatic. Every year of health insurance and private school tuition extends the polyester harmony that is their home life. But for the middle manager, the loyalist cluster is strategic. He’s like an old piece of tile covering himself in linoleum. It’s all wrong but a bigger nuisance to remove. Become hard to kill by entangling yourself deeper in a hierarchy that you created and from which you are a convenient buffer between partners and minions whose names they never want to know.

The misalignment, politics, and the waste of human life force in the name of self-preservation of an artificial environment (that’s a big one to unpack, but y’all might have to come have coffee or maybe something a bit stiffer for that convo) are regrettable.

But you know what I find worse?

That our middle-manager protagonist was not as cunning as he seems. That his primary offense was stupidity. That he sincerely mistook the situation for something he could fix instead of a noble impulse towards calculated self-preservation.

Stupidity is uncivilized. It’s unpredictable. “Say what you want about the tenets of national socialism, at least it’s an ethos. These guys are nihilists.” That’s how I feel about our manager if he is stupid. That’s a barbarian. It might be unpleasant, but you can reason with the merely deceptive.

I’m not sure if our middle-manager character (whose likeness to any real-life individual is purely coincidental) is sincere and stupid or just trying to survive, but those accomplices on high, blinded by desperate optimism, channeled the old guy at the club instead of having the courage to reinvent or the grace to just go home.

Then again, being who they are is what got them to giant ice cube status in the first place.

energy is the ultimate currency

A year ago I wrote a piece called Capitalism Is a Temporary Condition. Buried in it was a thought:

An investor in an age of acceleration watches as commodity prices, real things, hover in familiar territory while “future cash flows discounted” reward attention — in some cases because of optimistic stories, in some cases because the company exchanges cash for BTC (an act which adds no economic value), and in some cases for nostalgic lolz.

Today, it’s embarrassing to think of actual value. The most basic form is stored energy**. An obvious example is oil. Slightly more abstract —a building is a collection of atoms that took energy to arrange and provides utility. It’s proof of work. Same with a good reputation.

The ** footnote:

I increasingly think that a durable concept of a risk-free or least-risky rate is more likely to come from Vaclav Smil than U Chicago.

I knew about Smil from a profile piece that celebrated his suffer-no-fools rigor in making sense of the world.

I’m finally getting around to How the World Really Works, and I dig how quickly it’s zeroing in on what I anticipated to be the biggest muscle movement of all. I’ll let it unfold as he does the book’s intro.

Smil asks us to imagine an alien probe watching Earth, programmed to wake up whenever something important changes in the way energy is captured or used.

For a very long time, it sleeps.

Then life figures out photosynthesis: sunlight can be captured and stored as chemical energy. Much later, humans learn to control fire. Then agriculture. Draft animals. Wind and water.

The gaps are enormous. Billions of years. Millions. Thousands.

Then they start collapsing.

Coal gives humans access to immense stores of ancient solar energy. You can see a parallel to wealth in that phrase alone. Steam will turn that energy into mechanical work, replacing what he calls prime mover energy (human and animal labor). Oil makes enormous amounts of energy portable while electricity makes it transmissible and almost infinitely adaptable.

Smil puts numbers on this.

In 1800, the world had access to roughly 0.05 gigajoules of useful energy per person each year. By 1900 it was 2.7. By 1950, about 10. By 2000, 28. By 2020, roughly 34.

In a little over two centuries, useful energy per person increased nearly 700-fold.

The average person today commands roughly the energy equivalent of 60 adults working continuously, day and night. In affluent countries, the equivalent is closer to 200–240 permanent laborers.

And the averages hide enormous differences.

Measured in primary energy, someone in one of the world’s poorest countries may use less than 10 GJ a year. India is around 20–30. The global average is roughly 75–80. China is above 100. Western Europe and Japan generally sit around 120–160. The United States is closer to 250–280.

It sounds a lot like how GDP per capita is distributed. We measure GDP in dollars, but money can be printed, thus distorting its exchange rate versus the stored value of work. In that sense, energy may be civilization’s only uncorruptible currency.

If we continue pulling on the thread, wealth might be thought of as an account from which we can spend to hold entropy at bay.

  • A 72-degree house in Scottsdale.
  • Mango salad in during a Minnesota January.
  • Clean water pumped up the Hollywood Hills.
  • A sterile operating room.
  • Skyscrapers defying gravity.
  • Data centers (I couldn’t resist).

Much of the AI commotion revolves around what it means for humans to directly turn electricity into intelligence. That the product is intelligence which can recursively accelerate knowledge instinctively unsettles us because it violates our sensibilities around balance. Like energy is somehow not being conserved, leading to the type of divergence we associate with chain reactions.

I’m actually struggling to put my finger on the right analogy, which validates another point Smil makes. He argues that at precisely the moment when we have become capable of commanding extraordinary quantities of energy, most of us have become almost completely detached from how any of it happens.

Smil calls this our “comprehension deficit.”

Modernity is a black box. Light comes from a switch, and meat comes from a carton lined with a maxi pad. It’s all so effortless. Students of economics read I, Pencil to appreciate how capitalism’s profit-motive and competition actually lead to complex chains of cooperation. The side effect of the story is a (demoralizing?) inference that the world is so invisibly complicated now that you should feel good in your choice of major because it’s strategic to be a symbol-pusher since details are hopelessly infinite. (Sorry, did I go to far here? I’m an econ major, so it’s kind of like Chappelle using the n-word. Kinda? Nevermind.)

Smil’s punchy take on specialization:

You could meet real Renaissance men on Florence’s Piazza Signoria in 1500, but not for too long after that…By the middle of the 18th century, Diderot and d’Alembert could still assemble a group capable of summing up much of their era’s knowledge in the Encyclopédie…In 1872, a century after the appearance of the last volume of the Encyclopédie, any collection of knowledge had to resort to the superficial treatment of a rapidly expanding range of topics…today, it is impossible to sum up our understanding even within narrowly circumscribed specialties…experts in particle physics would find it very hard to understand even the first page of a new research paper in viral immunology …Highly specialized branches of modern science have become so arcane” that many practitioners must train into their thirties “in order to join the new priesthood.

All I’ve done is read the introduction to the book, but the “comprehension deficit”, which he’s fixin’ to remedy with respect to how the world works (at least from a physics/chemistry macro point of view), is just a fascinating idea of its own. Before the printing process and hyper-connectedness, it was hard to learn much of what was a relatively small body of knowledge, but since then, even though you could learn more in a lifetime, it was a smaller percent of what could be known.

But that I can sit here in an air-conditioned house on a 90-degree day typing contemplations about abstractions like how wealth is really the temporary abatement of nature’s randomness and how the accounting of such wealth should be denominated in units of work, which is traditionally measured by money and its continuously leaking exchange rate to said work is itself a wild collective achievement fueled by eons of carbon remnants from lives I never knew.

I guess I feel like I owe it to the universe to be curious about its ways even if it’s a Sisyphean fantasy to close the comprehension deficit. Looking forward to continuing: