Kelly Criterion — Cheatsheet Derivation

Kelly Criterion — Cheatsheet Derivation


Step 1 One Period Expectancy

E = pB − q

where p = win probability, q = 1−p = loss probability, B = net odds (win B per unit staked, lose 1).

Why: Weighted average of outcomes. Win B with probability p, lose 1 with probability q.


Step 2 Per-Flip Wealth Multipliers

Bet fraction f of current wealth W:

  • Win: W₁ = W(1+Bf)
  • Lose: W₁ = W(1−f)

Why: You keep the unbet portion W(1−f) regardless. On a win you collect B times your stake Wf on top. On a loss your stake Wf is gone.


Step 3 Wealth After n Flips

After h wins and (n−h) losses:

Wₙ = W₀(1+Bf)h(1−f)n−h

Why: The flips compound — each one rescales whatever the previous left. That means multiply, not add. Order doesn’t matter, only h and n−h.


Step 4 Per-Flip Growth Rate G

Take the nth root of total growth to extract the per-period rate:

G = (Wₙ/W₀)1/n = (1+Bf)h/n · (1−f)(n−h)/n

As n → ∞, law of large numbers: h/n → p, (n−h)/n → q.

G = (1+Bf)p(1−f)q

Why nth root: Same logic as extracting r from (1+r)n = total growth. Geometric mean, not arithmetic, because the process is multiplicative.


Step 5 Take ln Before Differentiating

Define g = ln(G). Since ln is monotonically increasing, maximizing g gives the same f* as maximizing G.

Apply two log rules:

  • ln(AB) = ln(A) + ln(B)  →  product becomes sum
  • ln(Ap) = p·ln(A)  →  exponent drops to coefficient
g = p·ln(1+Bf) + q·ln(1−f)

Why: Differentiating a product of powers is a mess. A sum of logs is trivial. Valid because ln is monotone — same maximum, easier math.


Step 6 Differentiate and Set to Zero

Rule: ddx[ln(x)] = 1x. Chain rule: multiply by derivative of the inside.

  • ddf[p·ln(1+Bf)] = pB1+Bf  ← chain rule gives B from inside (1+Bf)
  • ddf[q·ln(1−f)] = −q1−f  ← chain rule gives −1 from inside (1−f)

Set dg/df = 0:

pB1+Bf − q1−f = 0

Step 7 Solve for f*

Cross-multiply:

pB(1−f) = q(1+Bf)

Expand:

pB − pBf = q + qBf

Collect f terms:

pB − q = pBf + qBf = Bf(p+q)

Since p+q = 1:

f* = pB−qB = p − qB

The Answer

f* = pB − qB

Read as: edge / odds

  • Numerator pB−q is your expected profit per unit bet
  • Denominator B scales it by the odds

Special case B=1 (even money): f* = p−q

Your optimal bet equals your raw edge.


Key Insights

What Why it matters
Multiplicative wealth function One bad bet can’t be offset by other bets — sizing matters
Geometric mean not arithmetic Compounding processes need per-period rates, not averages
ln transform Turns product into sum without moving the maximum
Chain rule on ln(1−f) The −1 derivative is what creates a finite optimum
p+q=1 The simplification that closes the algebra cleanly

“end behavior”

Alex is an options trader you should follow in case he ever tweets a lot. Because he doesn’t, when he posted the question below a year ago, it got few responses. I took the liberty of posting it myself this week.

Article content

This was fun because it led to a lot of discussion on the timeline and DMs. I was told it sparked a bunch of quant debate on one trader’s desk.

The most popular answer, which was still less than 1/3 of the responses, was the correct answer.

Why?

The maximum value of a put is the strike. The maximum value of a call is the stock price.

Straddle is C + P so $100+$100 = $200

Notice how this means all call spreads go to zero since the calls are worth the same — the stock price. All put spreads go to their max value— the distance between strikes because the puts themselves are worth the strikes.

Logic for delta:

Delta is the change in option price per change in stock.

But the put’s strike is fixed, so the value of the put doesn’t depend on the stock price. The put has zero delta. It’s always worth $100. Which means it has no gamma either 🙂

The call is $100 because the max value of the call is the stock price. The call value moves 1-to-1 with the stock, so it has a delta of 1 or 100%

The max value of a straddle is therefore the stock price plus the strike price.

If you sell the straddle or either option at max value and hedge on its delta one time (this is known as a static hedge in contrast to dynamic hedging where you would rebalance as your hedge ratio changes), you cannot lose. It is that simple fact of arbitrage that makes it the upper bound.

To address the second most popular response in the poll, those who said the straddle is $100 (wrong) and has a 1.00 delta (correct), we will demonstrate why this is incorrect.

What’s your p/l if you sell 1 straddle at $100 and buy 100 shares against, if the stock goes to $300?

The straddle will be worth $400, so you lose $300 but make $200 on your long share.

Hmm, maybe I’m just underhedged. Fine, what if I hedge on a 200 delta?

In that case, you actually make money; you win $400 on your 2 shares more than offsetting the $300 straddle loss. But what if the stock went to zero?

Your straddle p/l is unchanged, but you lost $200 on the long stock position. Arbitrage max value means you cannot lose if you sell at that price. Since we found a losing scenario, the price is not the maximum arbitrage bound. If you sell the straddle at $200 and buy a single share of stock, there’s no scenario where you lose. It is the lowest straddle value for which this no-lose scenario is true, thus it’s the arbitrage bound.

Of course, this is but a toy problem where the call and put go to their maximum values because it’s a degenerate case of infinite time or vol. But learning how a function (an option price is just a function) behaves by observing its boundaries is good for intuition. You did this in 9th grade. Khan Academy can jog your memory:

Article content

In the real world, you can fleetingly find options that trade beyond their arbitrage values:

Article content

Earlier in the week, @DeepDishEnjoyer aka p4 wrote a thread about a dividend mispricing.

It led to some back and forth with passersbys who use options but appear to have large gaps in the fundamentals.

Between the maximum value poll and p4’s dividend lesson, it’s worth saying it:

In a proper option education, you spend a lot of time on arbitrage relationships, cost of carry, and synthetics before you ever hear the word “volatility”.

I didn’t study formal math but I imagine there’s a lot in common with the process of proofs. Arbitrages rest heavily on assumptions. So to understand the relationships, you are forced into an intimate familiarity with the assumptions. And in the extremes of everything, it’s the failure to examine assumptions that leads to being blindsided. But also, when things get extreme, to go on the attack means asking yourself, “Who’s on autopilot? Is this price resting on a stale assumption?” The arbitrage relationships give you the highest conceptual ROI that derivatives offer, you never learn the most useful thing derivatives can teach…passage over the “bridge of asses”.

If you want to see more examples of why option basics are so key to understanding assumptions and opportunities when things get weird:

Variance & Covariance Cheat Sheet

Variance & Covariance Cheat Sheet

Starting points (where every derivation begins)

Everything below is derived from these definitions. They’re the raw material — average squared deviation for variance, average product of deviations for covariance. When a derivation feels stuck, come back here and plug in.

Variance — average squared deviation from the mean
Var(X) = E[(X − μ)2]    μ = E[X]
Covariance — average product of deviations
Cov(X, Y) = E[(X − μX)(Y − μY)]
Sum of squared deviations (the un-averaged version)
SS = Σ (xi − μ)2    Var = SSn

Variance is just SS divided by n (or n−1 for a sample). Same object, before you average.

The move in every derivation: plug into one of these, expand the square or product (pure algebra), apply E using linearity, then recognize the Var/Cov patterns that fall out. The computational forms below (E[X2] − (E[X])2, E[XY] − E[X]E[Y]) are results of doing this, not starting points.


The identities

Variance from the definition
Var(X) = E[(X − μ)2] = E[X2] − (E[X])2

Average of the squares minus the square of the average. Worth showing where that second form comes from, since every later grind reuses this exact collapse. Start from the deviation definition and expand the square:

1n Σ(xi − x)2 = 1n Σ(xi2 − 2xix + x2)

Average term by term. The key is that x is a constant (already computed), so it pulls out of the sums:

  • First term: (1/n)Σxi2 = E[X2]
  • Middle term: (1/n)Σ(−2xix) = −2x · (1/n)Σxi = −2x · x = −2(E[X])2
  • Last term: (1/n)Σx2 = x2 = (E[X])2 (averaging a constant returns the constant)

Put them together — and notice the last term carries a coefficient of 1, not 2:

E[X2] − 2(E[X])2 + (E[X])2 = E[X2] − (E[X])2

The −2 and +1 combine to −1. That collapse — middle and last terms both becoming (E[X])2 and partially cancelling — is the same move behind every Var/Cov identity on this sheet.

Covariance from the definition
Cov(X, Y) = E[(X − μX)(Y − μY)] = E[XY] − E[X]E[Y]

Average of the products minus the product of the averages.

Variance is covariance with itself
Cov(X, X) = Var(X)
Scaling rule for variance
Var(aX) = a2 · Var(X)

Constants pull out as their square.

Scaling rule for covariance
Cov(aX, bY) = ab · Cov(X, Y)
Shifting rule for covariance
Cov(X + c, Y) = Cov(X, Y)

Adding a constant doesn’t change covariance.

Covariance with a constant is zero
Cov(X, c) = 0

Constants don’t co-vary.

Variance of a sum
Var(X + Y) = Var(X) + Var(Y) + 2 · Cov(X, Y)
Variance of a weighted sum (the workhorse)
Var(aX + bY) = a2 Var(X) + b2 Var(Y) + 2ab · Cov(X, Y)

Here a and b are the amounts held of each asset. They’re portfolio weights when they sum to 1. Var(X+Y) above is just this formula with a = b = 1 — one unit of each, no weighting lever. The weights are what turn a raw sum into a portfolio.

Worked example. Two assets: σX = 20%, σY = 10%, ρ = 0.3. Equal weights a = b = 0.5.
  • Var(X) = 0.04,   Var(Y) = 0.01
  • Cov(X, Y) = ρ · σX · σY = 0.3 · 0.2 · 0.1 = 0.006
  • Var(P) = 0.25·0.04 + 0.25·0.01 + 2·0.5·0.5·0.006 = 0.01 + 0.0025 + 0.003 = 0.0155
  • σP = √0.0155 ≈ 12.4%
Compare to the naive weighted-average vol, 0.5·20% + 0.5·10% = 15%. The cross term (with ρ < 1) is what pulls portfolio vol below the average of the two vols. That gap is the diversification benefit.
Variance of a difference (spread variance / pair-trading formula)
Var(X − Y) = Var(X) + Var(Y) − 2 · Cov(X, Y)

Same as Var(X+Y) but cross term flips sign. When X and Y are highly correlated, spread variance is small — the math behind why pair trades work.

Bilinearity of covariance
Cov(X, A + B) = Cov(X, A) + Cov(X, B)

Same in the first slot by symmetry.

Where it’s used. This is the move that lets you compute an asset’s covariance with a whole portfolio without re-deriving anything. Say a portfolio P = 0.5A + 0.5B and you want how asset A co-moves with the portfolio it sits in:
Cov(A, P) = Cov(A, 0.5A + 0.5B) = 0.5 Var(A) + 0.5 Cov(A, B)
Distribute across the sum, pull the weights out. That number — an asset’s covariance with its own portfolio — is its marginal contribution to portfolio risk, and it’s exactly what you FOIL out when you expand Var(w1X1 + … + wnXn) into the full covariance matrix. Bilinearity is the engine under every portfolio-variance calculation.
Variance of a binomial
Var(H) = np(1−p)   where H = Σ Xi

Derived in two steps, both from scratch.

Step 1 — variance of a single flip. One flip X is 1 with probability p, 0 with probability (1−p). Mean is E[X] = p. Plug into the squared-deviation definition — deviations are (1−p) for heads and (0−p) = −p for tails, each weighted by its probability:

Var(X) = p(1−p)2 + (1−p)p2

Factor out p(1−p): the bracket is (1−p) + p = 1, so

Var(X) = p(1−p)

(Peaks at p = 0.5, value 0.25 — the fair coin is the most uncertain, most variance per flip.)

Step 2 — n flips. Write H as a sum of n independent single flips, H = X1 + … + Xn. Variance of a sum adds the pairwise Cov terms, but independent flips have Cov(Xi, Xj) = 0, so every cross term drops. The n identical variances just add:

Var(H) = Σ Var(Xi) = n · p(1−p)

The np(1−p) isn’t handed to you — it falls out of one Bernoulli’s p(1−p) times n, because independence kills the covariances.

Standard deviation scaling
StDev(aX) = |a| · StDev(X)
Correlation definition
ρ = Cov(X, Y)σX · σY

When ρ = 1: Cov(X, Y)2 = Var(X) · Var(Y).


The derivation recipe

Every identity in this neighborhood comes out of the same five moves. When you see a Var or Cov of something built from linear combinations of random variables, this is the procedure.

  1. Plug into the definition. Use Var(Z) = E[Z2] − (E[Z])2 or Cov(X, Y) = E[XY] − E[X]E[Y] depending on what you’re computing.
  2. Expand squares and products. Pure algebra on the random variables. FOIL out any binomials. No expectations yet.
  3. Apply E using linearity. Distribute E across sums, pull constants out of expectations. This is the step that does the most work. Always handle linearity first when you have the chance — squaring is not linear, so you simplify E first and let the square wrap what’s left.
  4. Group matching terms. Line up the things that share factors (a2 terms together, ab terms together, b2 terms together, etc.).
  5. Factor and recognize. Pull out shared factors and spot the patterns: (E[X2] − (E[X])2) is Var(X), and (E[XY] − E[X]E[Y]) is Cov(X, Y).

The reason this recipe always closes: variances and covariances are quadratic in the underlying random variables, so expanding any square or product of linear combinations only generates more variances and covariances. Step 5 is recognition, not computation. There’s nowhere else for the algebra to land.


Two applications

Interview problem: E[H · T] for n coin flips

Flip a fair coin n = 100 times. H = heads, T = tails. Find E[H · T]. Worked slowly, because the one-line answer hides about six moves.

Step 1 — first reach, and why it fails. The instinct is E[H · T] = E[H] · E[T] = 50 · 50 = 2,500. But splitting a product of expectations like that is only legal when the two variables are independent. Check the precondition: H + T = 100, so knowing H pins down T exactly. Not independent. The naive split is off by a correction.

Step 2 — name the correction. That correction is what covariance is. Rearranging Cov(X, Y) = E[XY] − E[X]E[Y]:

E[H · T] = E[H] · E[T] + Cov(H, T)
Independent → Cov = 0 → naive split exact. Locked → Cov ≠ 0 → you need the term.

Step 3 — get the sign first. H + T = 100, so when H is above its mean, T is forced below. They move opposite, always. So Cov(H, T) is negative, and the true answer lands below 2,500.

Step 4 — compute Cov(H, T) via substitution. Since T = 100 − H, write Cov(H, T) = Cov(H, 100 − H) and split with bilinearity:

Cov(H, 100 − H) = Cov(H, 100) + Cov(H, −H)
First term is covariance with a constant → 0. Second term, pull out the −1 (scaling rule, ab = 1 · (−1) = −1):
= 0 − Cov(H, H) = −Var(H)
So Cov(H, T) = −Var(H). Now it’s earned, not asserted.

Step 5 — Var(H) is the binomial variance. H is the count of heads in n flips, so Var(H) = np(1−p) = 100 · 0.5 · 0.5 = 25.

Step 6 — land it.

E[H · T] = 2,500 − 25 = 2,475

Where n(n−1) comes from. Keep everything in symbols instead of plugging in. E[H] = np and E[T] = n(1−p), so E[H]·E[T] = n2p(1−p), and Var(H) = np(1−p). Then:

E[H · T] = n2p(1−p) − np(1−p) = p(1−p)[n2 − n] = n(n−1)p(1−p)
The n2 is the naive product, the −n is the variance shortfall, and factoring out p(1−p) leaves n(n−1). Sanity check: 100 · 99 · 0.25 = 2,475. ✓

The through-line: the answer falls short of E[H]·E[T] by exactly Var(H), because Cov(H, T) = −Var(H) whenever H and T sum to a constant.

Two-stock equal-weight portfolio with equal variances σ2 and correlation ρ

Start from the workhorse with a = b = 0.5:

Var(P) = 0.25 Var(X) + 0.25 Var(Y) + 2(0.5)(0.5) Cov(X, Y)

Impose equal variances Var(X) = Var(Y) = σ2, and write the cross term with correlation, Cov(X, Y) = ρσ2 (since Cov = ρ · σX · σY and both vols are σ):

Var(P) = 0.25σ2 + 0.25σ2 + 0.5ρσ2 = 0.5σ2 + 0.5ρσ2

Factor out 0.5σ2:

Var(P) =  σ2(1 + ρ)2
Why this is the instructive form. Weights are fixed (50/50) and both vols are fixed (σ). The only thing left moving is ρ. So the entire diversification effect is carried by the single factor (1 + ρ)/2 — a clean dial from 0 to 1 that multiplies the single-name variance. Sweep ρ and read what correlation actually does:
ρ Var(P) σP (vol) vs. one stock
+1σ2σno benefit — identical names
+0.50.75σ20.87σ13% vol cut
00.5σ20.71σ29% vol cut (the √½ case)
−0.50.25σ20.5σhalf the vol
−100risk fully cancels
The variance scales linearly in ρ, but the thing you feel — vol, σP = σ√((1+ρ)/2) — scales as the square root, so the first chunk of decorrelation buys more than the last. Going from ρ = 1 to ρ = 0.5 already takes 13% off your vol. You do not need negative correlation to diversify; anything below +1 helps. Negative correlation is just the strong form, and ρ = −1 is the perfect hedge where the two positions cancel outright.

The whole two-name diversification story lives in that (1 + ρ)/2 factor. Same vols, same weights, and correlation alone moves you from “no benefit” to “risk gone.”

Two-stock unequal-weight portfolio with equal variances σ2 and correlation ρ
Var(P) = σ2 [1 − 2w(1−w)(1−ρ)]

Diversification benefit is the product of a weight piece (2w(1−w), maxed at w = 0.5) and a correlation piece (1−ρ). Need both to get benefit. With equal variances, equal weighting is optimal — any tilt from 50/50 sacrifices diversification.

Two-stock unequal-variance portfolio → inverse-variance weighting

Now drop the equal-variance assumption. Keep σX2 and σY2 separate. Weights w and (1−w):

Var(P) = w2σX2 + (1−w)2σY2 + 2w(1−w)ρσXσY

Minimize over w. Var(P) is an upward parabola in w (positive coefficient on w2), so the critical point is the min. Differentiate term by term and set to zero:

2wσX2 − 2(1−w)σY2 + 2(1−2w)ρσXσY = 0

Divide by 2, expand, collect the w terms on the left and constants on the right, factor w out:

w(σX2 + σY2 − 2ρσXσY) = σY2 − ρσXσY
w* =  σY2 − ρσXσYσX2 + σY2 − 2ρσXσY

The denominator is Var(X − Y) — the spread variance from earlier. The numerator is σY2 − Cov(X, Y).

The payoff — set ρ = 0 (independent names):
w* = σY2σX2 + σY2 = 1/σX21/σX2 + 1/σY2
That’s inverse-variance weighting: each asset’s weight is its inverse variance over the sum of inverse variances. The quieter asset gets more money. This is the result behind Kalman filters, weighted least squares, and meta-analysis — anywhere you optimally combine noisy estimates, you weight by precision (1/variance). It’s also the “optimal” cousin of the inverse-vol risk-parity heuristic, which ignores correlations.

Intuition — hold σX fixed at 20% (σX2 = 0.04), turn the σY knob:

σY σY2 w* on X
000
10%0.010.20
20%0.040.50
40%0.160.80
∞∞→ 1

X’s weight is driven by Y’s variance, not its own. The noisier the alternative, the more you pile into X. Three anchors: σY2 = 0 → w* = 0 (Y is riskless, hold only Y); σY2 = σX2 → w* = 0.5 (equal variances recover equal weighting); σY2 → ∞ → w* → 1 (Y is pure noise, flee into X). The cleanest limit: if σX2 = 0, then w* = 1 — a riskless X takes the whole book. Precision is just quietness, and you trust the quiet estimate more.

Three-variable portfolio variance → why the matrix shows up

Same Form A grind, one more variable. Expand (aX + bY + cZ)2, apply E, subtract the squared-mean term. Every squared term becomes a variance, every cross term a covariance:

Var(aX + bY + cZ) = a2Var(X) + b2Var(Y) + c2Var(Z) + 2ab Cov(X,Y) + 2ac Cov(X,Z) + 2bc Cov(Y,Z)

Counting the terms. For n assets you always get:

  • n variance terms — one per asset (the ai2 Var pieces)
  • nC2 = n(n−1)/2 covariance pairs — one per distinct pair

Total = n + nC2. For n = 3: 3 + 3 = 6. For n = 100: 100 variances + 4,950 covariance pairs. The cross terms grow as n2, which is exactly why nobody writes portfolio variance longhand past n = 3 — you switch to the matrix form.

The double-sum / matrix form. Organize every term into a grid indexed by asset pairs. With weights wi and returns ri:
Var(P) = Σi Σj wi wj Cov(ri, rj) = w⊤Σw
Reading the double sum: the outer Σ over i and the inner Σ over j together form every ordered pair (i, j). For each pair you drop in one term, wiwjCov(ri, rj), and add them all up. For n = 3 that’s 3 × 3 = 9 cells:
X (j=1) Y (j=2) Z (j=3)
X (i=1)a2Var(X)ab Cov(X,Y)ac Cov(X,Z)
Y (i=2)ab Cov(X,Y)b2Var(Y)bc Cov(Y,Z)
Z (i=3)ac Cov(X,Z)bc Cov(Y,Z)c2Var(Z)
Sum all nine cells and you get the six-term formula above. Two things to see:
  • Diagonal (i = j, shaded): Cov(ri, ri) = Var(ri), so the diagonal is the n variance terms.
  • Off-diagonal (i ≠ j): each unordered pair appears twice — cell (X,Y) and cell (Y,X) are identical — and those two copies are exactly where the factor of 2 on each covariance comes from. You never write the 2 by hand; the grid double-counts it for you.
So Σ is the covariance matrix: variances down the diagonal, covariances off it. w⊤Σw just says “sweep every cell of the grid, weight it, sum it.” The n + nC2 count is the matrix — diagonal plus (doubled) upper triangle. You’ve already discovered why the matrix form is inevitable; it’s just bookkeeping for the term explosion.
Figure — the double sum is the matrix: one sweep, three views
Var(P) = Σi Σj wiwj Cov(ri, rj) an instruction for sweeping a grid: for every row i, for every column j, add that cell inner Σ over j → picks the column X (j=1) Y (j=2) Z (j=3) outer Σ over i → picks the row X (i=1) Y (i=2) Z (i=3) w1² Var(X) w1w2 Cov(X,Y) w1w3 Cov(X,Z) w2w1 Cov(X,Y) w2² Var(Y) w2w3 Cov(Y,Z) w3w1 Cov(X,Z) w3w2 Cov(Y,Z) w3² Var(Z) Diagonal (i = j): Cov(ri, ri) = Var(ri) — the 3 variance terms. Twin cells (i,j) & (j,i): identical — every covariance is visited twice. That double-count is the ×2 you wrote by hand. The grid supplies it for free. add all 9 cells = wT Σ w Σ is the covariance matrix: variances on the diagonal, covariances off it. n assets = the same sweep on an n×n grid. Nothing new happens; the grid just grows.
Figure — three forms of the same variance, and where the 2s go
① ALGEBRAIC — you write the 2s by hand a²Var(X) + b²Var(Y) + c²Var(Z) 2ab Cov(X,Y) + 2ac Cov(X,Z) + 2bc Cov(Y,Z) ② MATRIX — the 2s vanish into the symmetry Var(P) = wT Σ w,   where Σ = Var(X) Cov(X,Y) Cov(X,Z) Cov(X,Y) Var(Y) Cov(Y,Z) Cov(X,Z) Cov(Y,Z) Var(Z) Each covariance sits in two matched-color cells (mirrored across the diagonal). Summed, the two cells are the 2 from Stage 1. Variances (diagonal, purple) have no mirror — that’s why they’re never doubled. ③ DOUBLE SUM — the matrix written in math Var(P) = Σ n i=1 Σ n j=1 wiwj Cov(ri, rj) Cov(ri, ri) = Var(ri) — covariance with itself is just variance, so the diagonal needs no special case. Both sums run 1 to n, so the sweep is n² terms. 10 assets → 100 terms, not 20. That quadratic blow-up is exactly why the compact matrix form earns its keep.
Notation you’ll see in practice: w⊤Σw. You’ll run into this constantly in risk models, optimizers, and quant papers. It’s the same portfolio variance, packaged as a matrix operation. Reading it piece by piece:
  • w — the weight vector, weights stacked in a column.
  • w⊤ — “w transpose,” the same weights laid flat as a row. Transpose just tips a column over into a row.
  • Σ — the covariance matrix (the grid above). Watch out: this capital-sigma is the matrix, not a summation sign. Variances on the diagonal, covariances off it.
So w⊤Σw is row-of-weights × matrix × column-of-weights, which multiplies out to a single number — the portfolio variance. It’s identical to the double sum, and in a spreadsheet it’s literally =MMULT(MMULT(TRANSPOSE(w), Σ), w).

Why bother, when the double sum already shows everything? Three practical reasons, none of them “it’s more correct.” It doesn’t grow — three symbols whether n is 2 or 2,000, where the double sum for 500 names is 250,000 terms. It’s how software actually computes portfolio variance (one fast matrix op). And optimization only speaks matrix: the minimum-variance weights come out as Σ−11 normalized, and the inverse Σ−1 has no double-sum spelling. The two-asset inverse-variance weighting derived above is Σ−1 for n = 2 — the matrix form is how that generalizes. For understanding, the double sum is enough; this is the notation for doing things with it.


The ladder

Each rung is built from the one before it — the definition first, then the algebra of scaling and adding, then portfolios, then the jump to the matrix. Nothing is assumed that wasn’t derived earlier.

1.  Variance as average squared deviation
2.  Var(X) = E[X2] − (E[X])2 — from the definition
3.  Cov(X, Y) = E[XY] − E[X]E[Y] — from the definition
4.  Var(aX) = a2·Var(X)
5.  Cov(aX, bY) = ab·Cov(X, Y)
6.  Cov(X + c, Y) = Cov(X, Y) and Cov(X, c) = 0
7.  Var(X + Y) = Var(X) + Var(Y) + 2·Cov(X, Y)
8.  Var(aX + bY) — the full weighted-sum workhorse
9.  Two-stock equal-weight portfolio variance → the (1+ρ)/2 diversification factor
10. Two-stock unequal-weight portfolio variance (equal variances)
11. Var(X − Y) — spread variance / pair-trading formula
12. Variance of a binomial = np(1−p)
13. Two-stock unequal-variance portfolio → inverse-variance weighting
14. Var(aX + bY + cZ) — three variables, and the n + nC2 term count
15. General n-asset portfolio variance — the w⊤Σw matrix form

Is this one lesson in a math course?

No. This would be roughly half a semester of a first probability course, or a full chapter and a half of a more applied stats book.

Rough mapping to a standard curriculum:

  • Variance from the definition, E[X2] identity: one lecture, plus a problem set
  • Covariance and the product identity: one lecture
  • Scaling rules, bilinearity, variance of a sum: one to two lectures
  • Portfolio variance, weighted sums, correlation: one lecture in the probability course, or the opening week of a finance/portfolio-theory course
  • Binomial variance, applications: another lecture or two

So this is the equivalent of maybe four to six lectures of material, plus the problem sets that go with them. The reason it feels like a lot is that this sheet does the whole pipeline — derivation, intuition, numerical examples, applications — for each piece, instead of showing a formula and moving on.

The trade-off is real: this is slower but produces durable understanding. A typical math course shows you Var(aX + bY) on day one, leaves the “why it’s that and not something else” fuzzy, and you pattern-match for the rest of the semester. Done this way, when σ2 · (1+ρ)/2 turns up in a textbook two years later, you see the bilinearity FOIL behind it instead of recognizing a memorized formula.

That’s the trade. Slower, but it sticks.

why i’m not more bearish equities compared to 6 months ago

Programming Note: I normally publish Munchies on Wednesday and the paid post on Thursday, but this week I will post both a day early, as they are market-related and I want to get them out ahead of the Fed meeting.


Friends,

I’ll start with updating some broad market observations I laid out in March and then square that with what option surfaces are telling us in the context of the Fed meeting and beyond. The Fed meeting is a highly skewed event with a 25 bp hike more than about 90% priced in.

The probability of Fed target rate between 3.75% and 4% has shot from 35% to 90% in under 3 weeks

Recent context:

The Jackson Hole keynote was on 8/28/26. The probability of a rate hike shot from 35% to 57% in one day.

The 10-year yield has rallied from 4.67% on 8/26 to 4.96% as I write on 9/14.

In that same window, the SPX is down a mere 50 bps and the NDX is basically unched.

Let’s rewind to my March post trading is like a sudoku puzzle with prices as the given numbers. I compared earnings yields to bond yields to get an equity risk premium. This comparison in the modern era of massive budget deficits, where a large “G” in the Kalecki-Levy world tells us that money ends up as private sector nominal income by identity*, never flatters equities. We can reconcile skinny equity risk premiums the same we reconcile low earnings yield any stock…there’s an expectation of growth.

*While this is reductionist in the sense that the flow doesn’t definitionally have to end up there, it also happens to be where it has ended up.

If we are desensitized to the low absolute level of equity risk premiums, it is because it has been easy to presume nominal growth. We have relatively low unemployment, technology companies growing quickly despite massive scale, and, I almost forgot, A GOVERNMENT THAT IS EXISTENTIALLY POT-COMMITTED TO DEFICIT SPENDING.

“But, we have a Republican president.”

Ok, define Republican. Because if you look closely at our fiscal…oh never mind. Nobody cares. We’re well past fiscal discipline as a political delimiter.

The more money there is, the more surface area there is for theft/grift/graft.

[While this is now quite obvious and exploited by both right and left, Don’s brazenness feels like it’s giving everyone permission. He is a populist wrestling everyone’s birthright to cheat away from the cloaked political elites who once monopolized it. That sentence is the dress…whether you see blue or gold is up to you.]

Alrighty then, where were we? Ahh, yes, nominal growth as a given, because our society is a passenger in a global economic trolley problem. Fantastic. Instead of fretting over the level of the equity risk premium, we can look at the past 6 months to consider the change or, in a sudoku-esque way, solve for what needs to happen for the risk premium to not get worse.

Thus far, higher energy prices and higher bond yields haven’t produced the equity decline in the scenario I outlined. My hunch is the market got more expensive on a relative basis, but let’s investigate.

First, the updated numbers.

Market quotes below are September 14 intraday observations, taken at different times. The equity calculations use the same snapshot as the charts.

WTI oil is up 18.5% since the March post (March price from 3/27/26)

RBOB gasoline is up ~28%

IEF on a div-adjusted basis is down 1.9%, about half what the duration (which is a snapshot like delta) expects because you picked up over 200 bps of carried interest (ie those divs) in the meantime.

Equity valuation as the missing Sudoku number

In late March, I used an equity earnings yield of roughly 5% vs 4.4% ten-year Treasury yield. This represented a 60 bps nominal equity risk premium and ~300 bps in real terms. My downside scenario considers what might be expected if the Treasury yield reached 5.5% and equities needed to offer another 50 bps above that. The target nominal earnings yield would rise to 6%, requiring roughly 17% equity decline if earnings stayed flat.

On partial probabilities

My thinking was on a subset of causes for yields to rise. I was thinking about the inflation channel via energy price pass-through in the event that futures prices rolled up to war-bolstered petroleum spot prices. This is supply-shock inflation, but there is also demand-pull inflation. Yields can also increase because of concerns about sovereign creditworthiness. Asset prices are complex because they embed expectations about many variables, some of which reinforce each other and some offset. Arrows in every direction. Complex portfolios can be constructed to isolate bets on conditional or partial probabilities.

An example from sports betting:

Say you bet $900 to win $100 on the Seahawks not winning the Super Bowl, and $100 to win $600 on them winning the NFC. If they don’t win the NFC, the bets cancel. If they reach the Super Bowl, you make $700 if they lose and lose $300 if they win. You’ve constructed a conditional bet: Seattle to lose the Super Bowl, provided they get there.

The prices imply a 10% chance of winning the Super Bowl and a 14.3% chance of getting there. Divide those and you get a 70% chance of winning conditional on getting there. Your combined position bets against that 70%.

The way I’m reasoning in a macro way is implicitly partial. All relative value bets have this property, but be aware that if you bet on cross-asset, you are not truly isolating bets as cleanly as the Seahawks example. You are hand-waving all the other ways a bond yield, oil price, or stock earnings can change relative to one another. There’s no equivalent to “If they don’t win the NFC, the bets cancel” because the relationships in assets are not deterministic in the same way that winning the Super Bowl encompasses “winning the NFC”.

This undermines all the logic of my trade ideas to the extent that it sets up lots of ways for the market to creatively “middle” you, just like the bettor who lays off a sports bet and gets careless about that half a point. But since all relative value trading deserves the error bars I’m dancing between, I feel a bit better having disclaimed them. Confidence sells, but when it comes to markets it’s craven. Something to keep in mind when you’re listening to your next podcast.

Of course, if earnings growth increased fast enough, equities don’t need to decline at all to maintain the same equity risk premium to bond yields.

That leaves us with two questions:

  • What spread do today’s earnings estimates offer?
  • How much more earnings would restore the 50 bp equity risk premium?

At SPX 7,637.79, the approximately $397 of earnings expected over the next twelve months gives us a 5.20% earnings yield. Against a 4.96% Treasury yield, that’s only a 24 bp spread.

To get 50 bps over Treasuries, we need a 5.46% earnings yield:

7,637.79 × 5.46% = $417 of annual EPS.

How close are we to earning that much?

The first-half figures total approximately $181 per share, up 39% from the same quarters in 2025. Second-half estimates total approximately $183. The forecast calls for roughly maintaining the first-half earnings level through the rest of 2026, which represents 26% growth over the second half of 2025. Historically high, but a slower pace vs the first half’s increase over the same period a year earlier.

For the full calendar years, consensus is $362 in 2026 and $417 in 2027. That’s another 15% growth after this year’s expected 32% increase. Together, those forecasts take annual earnings from $275 in 2025 to $417 in 2027—52% growth in two years. These figures come from the same FactSet earnings series.

Back to the valuation. The calendar-2027 estimate already gets us almost exactly to our $417 target and assumes EPS growth of 15%, much more reasonable compared to historical changes and far slower growth than 2026 experienced.

Despite the rise in yields and inflation, consensus equity pricing offers an earnings path that supports approximately the original 50 bp spread benchmark. It requires delivering the rest of 2026 plus “only” 15% subsequent growth next year.

The equity risk premium has been skinny, but EPS growth has delivered such that those risk premiums aren’t actually shrinking despite the continued outperformance of equities vs bonds. I wouldn’t call that bullish exactly, but it’s not bearish versus where we were 6 months ago. When I first started this investigation with the high energy prices and rising yields, I expected the “missing Sudoku number” of equity risk premium to look even more paltry, but sparkling earnings have bailed the multiples out.

Stay groovy

☮️


End Notes

Gasoline futures and CPI

Persistently high fuel prices burden consumers and businesses. But keeping the price high doesn’t mean inflation stays high. If gasoline rises from $2 to $3 and stays there, year-over-year inflation is positive until the comparison catches up. Comparing $3 with $3 gives zero inflation.

The spread benchmark

“Equity risk premium” here is shorthand for earnings yield minus the nominal ten-year Treasury yield. It is a valuation comparison, not a complete estimate of equities’ expected excess return. Earnings are not contractual interest payments and are not necessarily distributed to shareholders.

Comparing that earnings yield with the roughly 2% TIPS yield gives a 300 bp difference. That is a comparison with a real bond yield, not a separately calculated real equity risk premium. My mental shorthand for real equity returns is that they typically realize between 300 and 600 bps over the risk-free rate. Equity valuation is on the high side (ie low real earnings yield over CPI), but it has been for a while, as it has been priced for growth which has been repeatedly confirmed.

Market snapshot and forward calculation

The calculations hold SPX at 7,637.79 and the ten-year Treasury yield at 4.96%, using September 14 intraday observations from MarketWatch and Trading Economics. These are not closing prices.

Tracking 2026

FactSet lists Q1 2026 EPS of $80.99 and Q2 EPS of $100.28. Q1 is shown as actual; Q2 remains marked as estimated. The $181.27 first-half total therefore should not be described as entirely finalized.

Index composition

The S&P 500’s membership changes. Depending on how a historical series is constructed, year-over-year earnings changes can reflect additions, deletions and changes in index representation, as well as growth within businesses.

Comparing each period’s membership differs from comparing a fixed set of companies across both periods. The chart therefore describes the published index earnings series, not a constant-company measure of organic growth. No adjustment for composition has been made.

after this post you will be sizing bets in your head

One of the most important concepts in risk-taking is bet sizing. Which is unfortunate because people are quite bad at it, while the effort to be way above average is quite low.

A jarring and famous demonstration of this is the Haghani-Dewey Coin Flipping study, which showed how even college grads with business, economic, and technical backgrounds incinerated their capital or massively underperformed the expected profits presented to them by a game they knew was rigged in their favor.

You can read my synopsis in Bet Sizing Is Not Intuitive.

For a binary wager (ie win or lose), if you know the payoffs and the probability of winning, both of which were known to the participants, the solution is to use the Kelly Criterion.

The tragedy is that it is incredibly simple to compute and applies to many conventional gambles and decisions (the examples in the quiz will span various life situations!).

If something is both easy and widely relevant, it should be common knowledge. So let’s fix that today. I’ll show you how easy it is to use, and you’ll forever be able to do it in your head.

First, a succinct definition:

Kelly is the bet size, as a fraction of bankroll, that maximizes the long-run compounded growth rate of your wealth. It’s a mathematical solution to bet size that doesn’t seek to maximize expected profit per trial, but the size that optimally balances compounding rate and survival.

If you want to go deep on this, see Moontowerquant’s Kelly Criterion Resources, but today’s focus is on getting straight to usability.

We will use this formulation of Kelly because it’s general:

f* = p − q/b

where:

p = probability of winning

q = 1−p or probability of losing

b = the odds you’re getting → what you win divided by what you risk

The easiest way to learn it is just jump right in with a few worked examples:

Fair coin wager (even odds style bet)

p =50%

q= 50%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 50% – 50% / 1 = 0 → bet nothing, zero edge

Coin biased in your favor (even odds style bet)

p =60%

q= 40%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 60% – 40% / 1 = .20 → bet 20% of your bankroll

Roll a 6 on a die (underdog bet where you get odds)

p =1/6

q= 5/6

b =8 (for a $1 bet, you either lose a $1 or make an $8 profit)

f* = 1/6 – (5/6) / 8

f* =8/48 – 5/48 = 3/48 → bet 6.25% of your bankroll

If f* is 0 or negative, you have no edge, so not betting is prescribed

Sports moneyline (betting as a favorite where you lay odds)

A −200 favorite. You risk $2 to win $1, and the line implies 2/3, but you think it’s closer to 3 in 4.

p = 75%
q = 25%
b = 0.5 (getting 50% return on the amount you risk)

f* = 75% − 25% / 0.5 = 75% − 50% = 25% → bet 25% of your bankroll

Wait a minute, these are large bets?!!

If these bet sizes seem surprisingly large for the given advantages, then your senses are well-tuned. For most people, “full” Kelly is too big!

Kelly maximizes long-run growth on the assumption your probability is correct. Well, it probably isn’t because the world is messy. We can inject some humility by using a fraction of Kelly:

  • “Half Kelly” gives up about a quarter of the growth rate and roughly halves the drawdowns. If you invert that, you see that the Kelly scaling law means as you bet bigger, you get diminishing returns per unit of risk. Extrapolating that logic, betting more than “full Kelly” is incinerating compounded wealth even if the individual bet has positive EV.
  • “Quarter Kelly” or less is far more common in practice.

Please don’t let the equation scare you, it’s intuitive and easy to remember

Look at the equation again:

f* = p − q/b

It’s just “how often you win” minus “how often you lose.” It’s just that the second term incorporates the payoff. The loss term gets divided by b, which represents the return you collect when you’re right.

  • When b = 1, you’re getting even money. A 100% return. Dividing by 1 leaves q alone, and the equation collapses to pure hit rate: p − q. That’s the coin case where you bet $1 to make $1.
  • When b > 1, you’re getting long odds. The division shrinks the loss term. This is why the die works. You lose 5 out of 6 rolls. Straight subtraction says you’re down 66 cents on the dollar and should never play, but you’re paid 8-to-1, so that 5/6 becomes 5/48, and suddenly the 1/6 win percentage is the bigger number. Long odds forgive a bad hit rate.
  • When b < 1, you’re laying odds. Now division stretches the loss term. The moneyline: you only lose a quarter of the time, but at −200 each loss costs you two units to earn back one, so that 25% loss percentage behaves like 50%. Being right three times out of four barely clears the bar. Lay enough odds and even a very good record is a losing proposition.

The graphic shows how you can think of the odds (the denominator) as shrinking or inflating q, as you collapse your thinking to a comparison of p vs q.

A word on b

b trips people up because “odds” is loaded gambler jargon. A wider interpretation of b is that it’s a percent return.

It’s what you make divided by what you risk. Even money is b = 1: risk a dollar, make a dollar. That’s a 100% return on the amount at stake. 3-to-1 is b = 3, a 300% return. Laying −200 is b = 0.5 because if you risk two to make one, it’s a 50% return.

[Return is a profit, while multiples don’t subtract your initial risk. It’s the difference between “I 2x’d my money” vs “I made 100%” or “I 10x’d my money” vs “I made 900%”. The percent return is the multiple minus one because we subtract our initial risk.]

The reason it’s a return and not just “the odds” is that Kelly assumes a loss wipes out the whole stake. The denominator is always the same number: everything you put up. b is comparable across a coin, a die, and a moneyline because it’s the return on risk, always measured against a total loss.

b = 1 is a natural reference point. At 100% return, a win exactly cancels a loss, so you need to win more than half the time. The breakeven hit rate changes with b.

Set f* = 0 and you get p = 1/(1+b).

Read the table as a menu of the hit rates you’re allowed to have. At b = 24 you can be wrong 24 times out of 25 and still be flat. At b = 0.25 you can be right four out of five and still be flat

Applying to real life: when is Kelly the right tool?

Kelly needs a few inputs: a bankroll, a payoff you know, and a probability estimate.

Which of these is a Kelly problem?

  1. A prediction market contract trading at 30¢. You think it’s worth 45¢.
  2. You’re all-in-or-fold on the river with a read that you’re good 40% of the time, getting 3-to-1 from the pot.
  3. How much of your 401(k) to put in equities.
  4. Writing checks as an angel investor across 30 startups.
  5. Buying a weekly call on a biotech ahead of an FDA decision date.
  6. Whether to take the new job.
  7. Your buddy offers you 5-to-1 that it rains in Oakland tomorrow. The forecast says 30%.
  8. Buying homeowners insurance. The premium is clearly more than the expected loss — that’s how the insurer stays in business.
  9. Your neighbor doesn’t carry homeowners coverage. She banks the premium instead.
  10. Your auto policy offers a $500 deductible or a $2,500 deductible, for a $340/yr discount on the premium.
  11. You’ve got vested startup options. Exercising costs $40k out of pocket in strike, and you think there’s maybe a 15% chance the company gets somewhere that makes them worth $1M.
  12. A merger arb spread. Target’s at $46, deal price is $50, and it trades back to $38 if the deal breaks. You think it closes 90% of the time.
  13. Your agency spends 20 hours of unbilled time on a speculative pitch. You win about a quarter of them, and a win is worth 80 billable hours.

Solutions to Kelly Problems

  1. Yes. Cleanest case there is. Binary, known payoff, and the price provides b directly. Risk 30¢ to make 70¢, so b = 2.33. f* = 45% − 55%/2.33 = 21%.
  2. Yes. This is the canonical one. p = 40%, b = 3, f* = 40% − 60%/3 = 20% of your stack. The wrinkle is that in poker your stack isn’t really your bankroll. There’s a whole literature on pros using Kelly for bankroll management across sessions rather than for a single river decision.
  3. No. Not this version of it. Stock returns aren’t generally binary so there’s no p, q, or b. There’s a continuous analog called Merton’s Share, which is similarly rooted in reward vs variance. What Gamblers Can Teach the Buy-and-Hold Crowd can get you started.
  4. Sort of. The structure is right: repeated, roughly binary, long odds. The problem is that p is a guess and b is a bigger guess, and Kelly is violently sensitive to overestimating your edge. Garbage in, garbage out.
  5. Approximately. If you treat it as approve/reject it’s binary enough to size with. If the expiry aligns with the date such that you are betting strictly on the terminal intrinsic value, the option piece will inherit the binary modeling you imposed on the stock.
  6. No. The variables are too opaque.
  7. Yes. p = 30%, b = 5, f* = 30% − 70%/5 = 16%. Note, you’ll lose this bet more than twice as often as you win it, so your most likely scenario is losing 16%. You can shrink the Kelly fraction if this makes you uncomfortable.
  8. Wrong side of the equation. Run f* on this and you get a negative number, because you’re buying a negative-EV bet.
  9. Yes. It’s the same policy, so notice that the bet only exists on the insurer’s side of it! Every year your neighbor doesn’t buy, she collects a premium and writes a tail. Rebuild cost $500k, premium $3,000, call it a 1-in-500 chance of a total loss.

    p = 99.8%
    q = 0.2%
    b = 3,000 / 500,000 = 0.006 (risk $500k to win $3,000)

    f* = 99.8% − 0.2%/0.006 = 99.8% − 33.3% = 66.5%

    Positive, as expected since insurers price premiums well above fair value. In this case, her bet size is the $500k house. If the house is most of her net worth, she’s at 100% on a bet capped at 66%. Rather than overbet, she should buy the policy. If she’s worth $5M, she’s betting 10% when she’s allowed 66%, which puts her near quarter Kelly (~14%), and she could skip the insurance. If she’s worth $1mm, it’s a 50% bet, which is more than half Kelly. I’d say take the insurance but I’m a wimp. There are other considerations (would she have the liquidity to rebuild the home or maybe taking the insurance with a high deductible is a better fit), but just doing this exercise gives you a sense of how risky or conservative your choices are relative to the bet share that maximizes long-term wealth.

  10. Yes. Raising the deductible is like you writing a $2,000 policy and collecting $340 a year for it. You’re the insurer again, so work out what you need to believe. Take the high deductible and save $340. Have a claim, and you’re out $2,000 more, but you already banked the $340, so the loss is $1,660.

    b = 340 / 1,660 = 0.205

    f* = p − q/0.205 = p − 4.88q

    Set that to zero, and you get p = 4.88q, which, with p + q = 1, means q = 17%. Your breakeven is a claim every 5.9 years. Anything less frequent and you’re the one with the edge.

    Let’s say real-world collision frequency is more like 6%. So p = 94%:

    f* = 94% − 6%/0.205 = 94% − 29.3% = 64.7%

    Which says risk at most ~65% of your bankroll. The risk here is $1,660. That clears as long as you have about $2,600 in liquid savings, which is to say the sizing check is trivially satisfied for almost everyone so you should generally opt for the higher deductible. For quarter Kelly, we’d need savings of $1,660/(.647 * .25) = $10,262.

  11. Approximately. It’s not truly binary, but if you frame it in a way where you are comfortable with the no consolation prize of a medium outcome, you can see it as paying $40k for a 15% shot at $1M. b = 24, so f* = 15% − 85%/24 = 11.5% of your liquid net worth. This is quite sensitive to your estimate of p of course.
  12. Yes, a classic example of binary-type risk in markets. You risk $8 to make $4, so b = 0.5 — you’re laying odds, same as the moneyline. f* = 90% − 10%/0.5 = 70%. That number is only as good as the 90%. Revise p to 75% and f* is 25%.
  13. Yes, in a subtle way! Your bankroll is capacity, not cash. b = 80/20 = 4, so f* = 25% − 75%/4 = 6.25%. Twenty hours has to be 6% of what you’re working with, which means you can’t run this pitch out of a 100-hour month. If you’re the manager, you can put it in dollar terms by converting to wages.

Finally, I strongly recommend William Poundstone’s book Fortune’s Formula: The Untold Story of the Scientific Betting System That Beat the Casinos and Wall Street

Description:

In 1956, two Bell Labs scientists discovered the scientific formula for getting rich. One was mathematician Claude Shannon, neurotic father of our digital age, whose genius is ranked with Einstein’s. The other was John L. Kelly Jr., a Texas-born, gun-toting physicist. Together they applied the science of information theory—the basis of computers and the Internet—to the problem of making as much money as possible, as fast as possible.

Shannon and MIT mathematician Edward O. Thorp took the “Kelly formula” to Las Vegas. It worked. They realized that there was even more money to be made in the stock market. Thorp used the Kelly system with his phenomenally successful hedge fund, Princeton-Newport Partners. Shannon became a successful investor, too, topping even Warren Buffett’s rate of return. Fortune’s Formula traces how the Kelly formula sparked controversy even as it made fortunes at racetracks, casinos, and trading desks. It reveals the dark side of this alluring scheme, which is founded on exploiting an insider’s edge.

Shannon believed it was possible for a smart investor to beat the market—and William Poundstone’s Fortune’s Formula will convince you that he was right.

And this is from my notes, Insights From Fortune’s Formula:

a gripping narrative full of 20th century trivia that ties together the birth of information theory, some of the greatest scientific minds of the 1900s, the rise of quantitative finance, and the role of organized crime. These topics come alive in a fresh, memorable way when discovered through the lens of its colorful characters.

It chronicles the history of the efficient market hypothesis (MIT, U Chicago, Paul Samuelson). You can organize its conclusion around this excerpt:

There is much truth in the efficient market hypothesis. The controversy has always been over just how far the claim can be pressed. Asking whether markets are efficient is like asking whether the world is round. The best way to answer depends on the expectations and sophistication of the questioner. If someone is asking whether the world is round or flat, as fifteenth-century Europeans might have asked, then “round” is a better answer. If someone knows that and is asking whether the earth is a geometrically perfect sphere, the answer is no.

the “middle manager”

I ask for your indulgence, for what follows is an informal word-wall connecting conversations I’ve had with some W2 friends.

Pick your favorite gospel on disruption. Schumpeter’s creative destruction, Christensen’s innovator’s dilemma. Most companies, sometimes even industries, will eventually recognize themselves as a melting ice cube.

In a great talk back in 2009, media investor Peter Chernin warned:

One of the things I always used to say to the people who work for me is that you can’t protect your business, and your job is not to protect your business. Your job isn’t to protect. Your job is to maximize your business at any given time, but your real job is to grow new businesses faster than the old ones decay.

He uses record labels as the cautionary example. They set out to protect their business, but “to the degree you’re going to try and do it you’re going to get killed because technology is going to liberate audiences in such a way that they’re going to get what they want regardless.”

It feels like there are only bad choices for companies that are now run-off businesses. Re-investment feels too speculative, but panic will preclude any chance of a soft or at least dignified landing. Navigating these moments well is a sign of grace, but if grace were easy we’d call it something else.

Instead we get to witness corporate pathology.

The ice cube is melting. But it started as a very big ice cube. Big enough to bridge a middle-aged middle-manager’s soft landing into an early retirement if they can claim a larger share of a shrinking cube. Welcome to corporate hunger games on steroids.

The middle manager is a risk-averse incrementalist. That’s WHY they are middle managers. They will do no such thing as grow new businesses faster than the old ones decay. They need an accomplice to secure their spot on the life raft. This accomplice comes from an unexpected place. The partners.

You see, entrepreneurs have deranged brain chemistry. My friend Jason Buck and I were drinking coffee in my backyard on Wednesday when he described an entrepreneur as someone who works 100 hours a week for themselves to not have to work an hour for someone else.

The entrepreneur, the founder, the partner is a delusional optimist. A melting ice cube can be refrozen if you just find the right segment to sell to, the one that somehow eluded you when things were going well. The middle manager latches on to this hope.

He crafts a business plan to do things “radically” differently. Of course, radical in this context is a mere gesture in comparison to the extinction-level shift in the business environment. The partners, never ones to back down from a fight, support this can-do attitude from the manager who stepped up.

The manager, energized and enabled, moves to consolidate influence. By promoting their plan, any plan, doomed as it is, they portray unsupportive colleagues as quitters by virtue of their dissent.

“If you’re not on board with my [dumb] plan, you clearly hate the company and are not a team player.”

There’s no serious appraisal of the dissenting argument’s merit. And that’s probably because the dissenting argument is “Umm, we’re f’d, so we should focus on retention, not growth, because [insert analysis that actually makes sense]”. Any analysis that maximizes EV in a losing game will be hard to sell against a bad strategy employed with hope. We gamble to get back to even and we’re risk-averse when we’re ahead.

Our middle-manager doesn’t just knife out the dissenters. They fully larp optimism with expansion. They hire loyalists, veterans of this obsolete-but-not-yet-obviously-so strategy. The loyalists are relieved. They, themselves, were cooked seeing the same writing-on-the-wall, but they just got thrown a lifeline by the last firm running the old plays. The only plays they’re familiar with. The reunion with their middle-manager buddy is as predictable as you expect, complete with brown liquor and nostalgia for those nights during training when they were chasing tail in Murray Hill and scarfing khati roll together if they struck out that night. Ah, Indian food before bed. To be young again.

The loyalist cluster is pragmatic. Every year of health insurance and private school tuition extends the polyester harmony that is their home life. But for the middle manager, the loyalist cluster is strategic. He’s like an old piece of tile covering himself in linoleum. It’s all wrong but a bigger nuisance to remove. Become hard to kill by entangling yourself deeper in a hierarchy that you created and from which you are a convenient buffer between partners and minions whose names they never want to know.

The misalignment, politics, and the waste of human life force in the name of self-preservation of an artificial environment (that’s a big one to unpack, but y’all might have to come have coffee or maybe something a bit stiffer for that convo) are regrettable.

But you know what I find worse?

That our middle-manager protagonist was not as cunning as he seems. That his primary offense was stupidity. That he sincerely mistook the situation for something he could fix instead of a noble impulse towards calculated self-preservation.

Stupidity is uncivilized. It’s unpredictable. “Say what you want about the tenets of national socialism, at least it’s an ethos. These guys are nihilists.” That’s how I feel about our manager if he is stupid. That’s a barbarian. It might be unpleasant, but you can reason with the merely deceptive.

I’m not sure if our middle-manager character (whose likeness to any real-life individual is purely coincidental) is sincere and stupid or just trying to survive, but those accomplices on high, blinded by desperate optimism, channeled the old guy at the club instead of having the courage to reinvent or the grace to just go home.

Then again, being who they are is what got them to giant ice cube status in the first place.

futures premium cell

Back in my SIG days, every trader had a cell on their spreadsheet showing the SPUs (the name for SPX futures back in the day; it’s derived from the September symbol) premium or discount to the “cash”.

A little background

The cash is the SPX index price based on its components’ spot prices. You can compute the index value from the stock bids, and that would be the “SPX cash bid”. You can do this for the offer as well. The average of the bid and offer is SPX mid-market.

Futures contracts, based on no-arbitrage pricing, have a fair value based on what you gain or give up by owning the future instead of the cash basket. Since you save the interest on the cash it would cost to own the basket but forgo any dividends you would have received that offset the drop in the shares when they are paid the futures are valued as

SPX + (interest – dividends)

The interest and dividends are estimated from today until the futures’ expiry date.

Assume:

SPX cash mid market: 7,800

RFR = 3.5%

annual dividend yield =1.5%

t = 1 year

The 1-year future fair value is approximately 7800e(3.5% – 1.5%)*1 = 7957.57

The basis between the future and the cash index was known as the EFP (“exchange for physical”). In our example, it equals 157.57

If the future was trading 7977.57, we’d say “the futures are trading over”. In this case, it would be 20 points or about 25 bps rich as a percent of the cash index.

This is an arbitrageable difference as a trading firm can sell the future, buy all the components, and book a theoretical profit that will become a real profit if their carry assumptions hold true. As a market-maker on the floor, I was not involved in index arbitrage this directly (although most of my clerking experience was much closer to these strategies).

S&P Futures and Fair Value. | The Blue Collar Investor
You have probably seen this kind of graphic on CNBC

The futures are more liquid than the cash index so the basket price lags. If systematic bullish news hits the tape, you cross the tiny bid/ask on the futures. In fact, the ES or e-mini future actually leads the big SPUs, so ES was used for computing the premium or discount.

All of the market-makers had a cell on the spreadsheet on their handheld tablets that showed the futures’ premium/discount to fair value (FV = cash index + EFP). Those premiums or discounts were typically small, 10 bps or less, as the market bounced around. But huge news could catapult the futures 300 bps before much of the basket could blink. As you can imagine, index arbs get very uneven fills on the baskets they try to execute to close the gap. That’s because everyone making markets has not only pulled their stock offers, but lifting resting customer offers, often several levels through the NBBO that was posted a split second ago.

You can beta-weight the “amount over” or premium the futures are trading to estimate a new fair value for the stock. If the stock you trade has a 1.2 beta, then you might think its new fair value is 3.6% higher than the pre-news price if the futures are trading at a 3% premium. Of course, beta is just a statistical quantity so you have a confidence interval around the beta, which is pretty much life as a market maker. Futures are trading 300 bps over, what’s your 2-way on XYZ stock? Maybe I’m “2% over” bid, offered “4% over”. Then, you’d need to be quick to remember what bids and offers on the option chain are resting that you should race to lift (all the calls should increase by their delta * your beta-weighted change in fair value) and hit (all the puts will decrease by the same factor and this is not even considering the effect of gamma).

Today, all the quotes being streamed by traders will get pulled while they simultaneously blast all the resting orders in a flash, but this process happened a bit more at human speed 30 years ago (although you still weren’t gonna beat the floor…the competition for the resting orders would be market-maker on market-maker).

I’m mostly sharing this because it’s fun for the stragglers who read this far into the masochists section, but it’s worth mentioning the context that made me think of it.

This phenomenon sometimes leads to anti-data or an interpretation of data that is exactly the opposite of the typical inference.

Why?

If a contract or security trades on the offer, the assumption is generally that “paper” (trader language for customers) is buying and the market-maker is selling. But in the example above, the market-makers are the aggressors because they know the market is much higher than last sale. They are lifting calls and hitting puts, so your read on flow sentiment is exactly the opposite of a normal market condition.

Just another example of reality having a surprising amount of detail.

stock-bond correlation

Return Stacked’s RSSB gives you a dollar of global equities and a dollar of Treasuries on the same dollar of capital. They just listed options on it this week. It will be a useful option market to watch if it gathers liquidity because its implied vol allows you to back out an implied stock-bond correlation since we know the vol of the legs.

Risk-parity funds and strategies are diversified, which gives them cover to use leverage but also means they are structurally short stock-bond correlation. In theory, they should be natural buyers of RSSB vol. If they actually did that, I’d expect the implied vol to trade at a healthy risk premium since it’s a one-way risk. There’s no real natural seller of that correlation.

I asked the Moontower Agent to work out the details and generate a deck (which is also a heat check on its abilities.) It produced a deck I could download, walking through the variance identity, a IEF-for-GOVT proxy swap, a scenario table for implied ρ across RSSB IV levels, and the noise associated with realized correlation.

The deck is short and educational:

Reading Implied Correlation Through RSSB Options

Confidence Intervals on Correlation

This made me look up a time series of the correlation between stocks and bonds using VTI and IEF as respective representatives of the asset classes. The 30-day came back +0.43, the 1-year at +0.28.

The agent volunteered the following statistical insight:

N=22 daily returns is a small sample. The 95% confidence interval on a correlation of +0.43 with N=22 is roughly [+0.02, +0.72] — wide. You cannot confidently distinguish +0.43 from +0.20 or +0.60 at that sample size. The 1Y number (N=251) is statistically much tighter — CI roughly [+0.16, +0.39].

Normally, when you compute a confidence interval, you effectively make a market by scaling the standard deviation to your desired confidence (so 1.96 or “2 sigma” for ~95% confidence)

But take notice of the first interval: +0.02 to +0.72 isn’t centered on +0.43. The upper tail runs 0.29 above, the lower tail 0.41 below. Weird. You wouldn’t see lopsided error bars if the math was something like “estimate ± 1.96 × standard error”.

Agent teach me what you did and why.

The problem is that the error bar’s width depends on the answer

Correlation lives on [−1, +1]. Bounded. If the true ρ is 0.9, your sample estimates can’t overshoot by much (the ceiling is 1.0) but they can undershoot plenty.

That boundedness shows up in the standard error of r itself, which runs about (1 − r²)/√(N − 1). Look at what that does. At r = 0.43 with N = 22 the SE is 0.177. At r = 0.9 with the same 22 observations it’s 0.042, four times tighter. Not because you learned more, but because you got squeezed against the wall.

The agent used something it called Fisher’s fix to move to a coordinate system where the standard error stops depending on r, do the easy symmetric thing there, and come back.

The recipe

1. Transform the point estimate: z = arctanh(r) = ½ · ln[(1 + r) / (1 − r)]

2. Take the standard error in z-space: SE = 1/√(N − 3). Note what’s missing. No r. Sample size is the only input.

3. Build the interval symmetrically, the boring way you already know: z ± 1.96 · SE

4. Recover each endpoint separately: tanh(z_low) and tanh(z_high)

Article content

Broadly educational bits to notice

Step 1 does almost nothing to the estimate.

0.4332 becomes 0.4635. The transform is only there to get the r out of the standard error in step 2.

Step 4 is where the lopsidedness comes from.

Both 30-day endpoints sit exactly 0.4497 away from the center in z-space. Perfectly symmetric. But tanh squashes hard once you’re out past 0.5 but does almost nothing near the origin, so the top end at 0.9131 gets crushed down to +0.72 while the bottom end at 0.0139 is nearly untouched at +0.014. What matters is each endpoint’s distance from zero, not its distance from your estimate.

Step 2 accounts for sample size

Ten times the sample and the standard error only comes down by a factor of 3.6. At large N the confidence grows by √N scaling, but at small N, subtracting 3 makes the confidence grow slower than square root scaling.

2 lessons

  1. The object-level lesson: Correlations are volatile and because they are bounded from [-1,1] we need to recenter our point estimates before cuffing their ranges
  2. The meta-lesson: LLMs are great math tutors so when they volunteer info of the “I don’t know what I don’t know” variety they can hold your hand. I’ve shared how I use LLMs as a tutor before in Socrates 2026: how to use highlights

collar shopping

At the end of July, Dean Curnutt tweeted:

The thread should sound familiar. Weeks earlier, Dean tweeted about SNDK vols presenting attractive collar pricing for hedgers.

From high implied vol can work for or against you?:

The high vol is a gift to the natural holders. Millennial employees can lock in their unborn grandkids’ inheritance.

A non-technical way to appreciate how high vol creates this opportunity in upside call vs downside put differentials: Imagine a stock starts at $100. It gets to $125. From $125 to $150 is only 20%. But if the stock fell to $75 the distance from $75 to $50 is 33%. Both $50 and $150 are 50% from the initial price, but in a compounding sense 150 is much “closer” to the starting value than $50. The higher the vol the less “distance” a fixed dollar move represents. As implied vol increases OTM calls grow faster in value than OTM puts. This is the source of the attractive pricing you see in the risk reversal (ie option collar).

That pricing falls out of the risk-neutral world that rests on “no-arbitrage” assumptions. In fact, the lognormal return process, described in the quote without ever using the word “lognormal”, is only one of many assumptions. The forward price of the stock is also assumed to be a function of the risk-free rate, which pushes the price that the options are based on higher than the spot price. This, of course, pushes up calls and puts down by their respective deltas. These assumptions conspire to make calls look quite expensive relative to puts to someone comparing the risk-reward of options intuitively.

Intuition would likely lead to very different prices, but present an arbitrage in the process because you live in the real world not the risk-neutral world.

A few examples from that post:

  • Warren Buffett sells the long-dated puts because he believes the no-arbitrage assumption of the risk-free rate underestimates real-world forward prices. The option trader he faces is quite content to buy those puts since they are looking for an easy flip in the vol market.
  • FX carry. The speculator holds the future for the risky but generous profit the risk-neutral price creates. The derivatives trader is content to make a penny of arbitrage profit.

This is an admittedly mind-bending state of affairs, but it creates disagreement because of different horizons. This is the basis for trade!

The collars (also known as risk reversals or fences depending on the trading floor you grew up on) are another example of risk-neutral assumptions presenting wacky prices to those who don’t require arbitrage to trade.

The arbitrageur prices the collar like a contractor bidding a job. The daily hedges are materials and labor. The replication price is a manufacturing cost. Call it $500 per square foot. But the client trading the collar is looking at the finished house represented by the payoff of where the stock lands. They are happy to pay $500/ sq ft if the final product is worth $800/ft. The derivatives trader is not underwriting the final value but operating a cost plus business. They have no view on the value of the final product. It’s literally none of their business.

If you want to tangle with this deeper, you are welcome to revisit how to get arbed with perfect information (again) but take heart that you aren’t alone if you struggle with the idea of no-arbitrage replication. I always say it’s the “bridge of asses” for finance.

Collar Shopping

Dean’s tweets highlighted stocks where the calls you overwrite can finance a fairly high strike put, creating an attractive hedge and therefore risk-reward to a long position.

Consider an example where you buy a $100 stock and a 1-year 20% out-of-the-money put while selling a 1-year 40% out-of-the-money call at the same price as the put. In other words, a “zero-cost” collar. The most you can lose is 20% in the case where the stock tanks. Your upside is capped at 40% before your shares would be called away. If you think the stock has a symmetrical distribution where it’s going up or going down are the same probability this looks like a good bet because you are getting 2-1 odds.

If the collar costs $2 or 2% of the stock price, we can roll that into the basis (so $102). Now you are risking 22% to make 38%, a risk reward (R:R) of 1.73 to 1.

This is a snapshot from the Moontower Collars Workflow on 8/11/26 for options with about 3 months to expiry.

Every row on the screen prices the package of buying the .25 delta put, selling the .25 delta call, and buying the stock.

You can restrict your universe to certain sectors and re-sort. You can filter for stocks a certain threshold above their moving average or stocks that are strongly correlated to each other to hunt for the best bang for your buck for a given beta, or just any number of screens to narrow your candidates.

A thread for the masochists

The strikes here are delta-defined, so the vol gap between names is already baked in. You could imagine beta-adjusting the strikes to go a step further (although you’d need to use the API/MCP).

Beta is correlation times the vol ratio. So if correlation < 1.0, beta-adjusting pulls the call strike in toward at-the-money. You’re short that call in a collar, so a nearer strike prints fatter making the risk-reward look better. But it’s a trade-off. The result comes from truncating idiosyncratic upside that is not captured in beta.

Defensive-minded investors should take note (as Dean clearly has). The multi-year highs in deferred interest rates (pushes up forward prices) and the fact that high-flying stocks that have made concentrated investors quite rich are also relatively high-volatility presents those of us in the real (not risk-neutral) world an attractive menu of hedges.

A final puzzle for the masochists

Collared stock has the same hockey stock diagram as owning a call spread of the same strikes but the total cost will vary by some amount. How much is the amount and why? The hint lies in put-call parity. Solving this should feel like a puzzle and is a step on the “bridge of asses”.

the sound of inevitability

The market is 12-15.

12 bid. 15 offer.

The broker sizes up the offer.

“How many you got there? How about you? And you?”

A couple of the market makers get flakey once they see him counting.

“You know what, I’m 17 now.”

“Fine, but can you fill the size there?”

“This second, no more shopping.”

“Mine.”

A few minutes pass.

The broker comes back around. “How now?”

“Go fuck yourself. 20-40. Small.”

And that’s it.

It’s a liquidity-clearing trade. The market’s version of punctuated equilibrium.

Small lot sizes from here on out.

The price drifts higher over time. The same amount of volume moves the price by larger increments. The least-capitalized shorts who are also the smallest in the trade begrudgingly cover. Better to live another day. The larger ones lay off some of the headline greeks in related assets, but the basis leaks against them the whole way. At least it’s not a fresh flesh wound every day. Paper cuts aren’t mortal risks, but the job will be demoralizing for a while. We’re gonna do this again? Why don’t I learn? I’ve done well enough. Right? I don’t need this shit I say as I fire up loopnet on yet another chrome tab hoping to find a cap rate in the shape of an eternal palm tree.

Weeks pass. Maybe longer. Who’s counting?

You see the price. 63. Numb. Doesn’t mean anything anymore. You can’t taunt a ghost.

More time.

Wait, 56?

“Anybody doing anything in this?”

“Nah, maybe just recent sympathy with hawkish Fed chatter.”

“You think this thing being up 300% has anything to do with basis points. C’mon.”

“Yea, I don’t know. I’m trying to book Odyssey tickets on IMAX for 3am, can we do this later.”

“Neverm—”

[Ringing. The hoot flashes.]

[Groans and picks up.]

“Sal, I thought I told you to cut the line. The fuck you want?”

“How is it today?”

“I don’t know, Saaaal, why don’t you tell me how it is?”

“At 45”.

“We’re 5 minutes from the close, I can’t show.”

“At 33. They’re gonna trade”.

“I’ll round you out just to be social.”

Next day.

How?

“15-25. Your move, Sal.”

Leo

Matt Levine:

[On the modern AI thesis] Aschenbrenner was early and smart and articulate about it. This allowed him to raise money, and the fact that it has been basically correct allowed him to return 270% through May.

Leopold Aschenbrenner’s fund, Situational Awareness LP (SALP) started in late 2024 when he was 23. He raised $225mm and, through the use of leverage and being right as hell, rode a legendary heater with assets at peak over $25B about a month ago.

His portfolio concentrated heavily on hardware and chip shares of CoreWeave, Nebius, Bloom Energy, Iris Energy, Micron and the Korean company SK Hynix, as well as a sizable private stake in Anthropic. Meanwhile, he was shorting traditional software companies. You only need to pull up the charts of his longs and shorts to explain how his returns had been so stellar.

Many of his longs peaked in late June. On July 24th, he sent a memo to investors stating that the fund “has not been immune” to recent market volatility. He invited existing investors to commit fresh capital effective Aug 1, citing the most attractive opportunity set since early 2025.

Within the week, SALP would proceed to lose 2/3 of its assets and liquidate its public portfolio to Citadel.

Back to the pit

Trading is about pricing liquidity. Handicapping the price to move a chunk of risk in a particular period of time.

Leopold was very right on his security selection. So right that even net of the liquidation, the fund is still up 80% on the year! That’s got to be unprecedented. But being liquidated in the first place was inevitable.

Let’s go back to the stylized story from the opening.

When an asset rips higher as quickly as these longs did, it exhausts the supply of offers that maintain a sensible relationship to a concept of fair value. For the sake of legibility, we’ll call those sellers the natural investors. The ones whose bids and offers are tied to some semblance of a fundamental model.

If a stock you sold when it 3x’d continues on its way to a 10x in a short window of time, like on the order of months, there is a paradigm shift in its liquidity. In a squeeze, there is a shortage of supply, but those episodes are faster and less mysterious. In the opening example, it’s not a supply squeeze but reluctance.

With the “naturals” long taken out of the stock, the marginal liquidity provider on the offer is an atheist. In other words, a trader. They have no religion about the short. It’s an HFT, a market maker trading some low-capacity intraday basis they discovered in a linear regression, or some passive mechanism with a rebalance toggle.

What do these sellers have in common?

They’re hyper-tuned to the risk. They have a volatility number somewhere in their trade lifecycle.

Fine, who’s buying when the stock is now twice the price that itself was twice the price of anything sensible?

For starters, the original fanatic, especially if they are receiving inflows based on the marks they are reinforcing. Who else is buying? The momos and fomos. Weak hands. The momos are executing a simple bandwagon python script. This is a weak hand by design. The fomo weak hands come in 2 forms. The ones who haven’t had an original thought in their life and the ones chasing a benchmark because they devoted their life to mimicry with just enough leeway to preserve the illusion that their creativity matters. There’s a price for everything, so no shade, but calling a spade a spade, this is the weakest hand.

The stock is in a liminal zone. Levitating on the echo flows from the original disturbance. The stock market’s microstructure always presents a wash of back-and-forth trading. The withdrawal of real liquidity is less visible than the example in the opening sequence, which caricatures a market in a derivative or obscure contract that trades on Clearport by appointment but lacks a DOM. But make no mistake, the liquidity of the super stock is broken just like the fake derivative example.

The liminal zone, to quote Kindergarten Cop, “lacks discipline”. The sellers are atheists, the buyers are momos, fomos, and the original mover gathering flows, literally “high on his own supply”, with a mandate to trade on a thesis he publicly telegraphed in an adversarial game. The marketing effect of this strategy was powerful but not free. Why not? Because it rings the dinner bell which reverberates with the sound of inevitability.

Mordecai

When I graduated college in 2000, my mother took my sis and me to visit our family in Sydney and travel Australia for 3 weeks. In Darwin, we took one of those boat tours on the Adelaide River where they hang massive slabs of meat over the sides so us tourists can watch the crocs coil below and then leap high for their meals. Once you get on the river, the swarm of eyes comes out of the weeds as the sound of the motor signifies meal time. I’ll never forget the sheer size and thus the name of the croc they told us was the river’s alpha — Mordecai.

Leo’s wild success summoned Mordecai.

Crocs don’t need to chase. They don’t waste energy. The strike happens in a muddy thrash, and shortly after, the ripples of water slow as the trees and surrounding fauna relax in the wake of violent awe.

And even if the alpha crocs turns over, replaced by a new alpha, this species lives forever.

If I can prove how apt this analogy is, you will believe, like I do, that this liquidation was inevitable.

If you’ve been following this saga, you will notice I haven’t yet introduced the true culprit — Leverage + Concentration. The crocs aren’t the villains. In the words of Jack White, “if you’re headed to the grave you don’t blame the hearse”.

The moment Leo chose 4x leverage on a concentrated book, he splashed loudly into the river. From there, crocs just do what they do.

Byrne Hobart, in the Diff:

When there’s an economic actor whose day job is to identify forces that will lead to short-term flows in and out of particular stocks, and whose long-term model is to periodically pounce on distressed companies, these models will tend to converge into a model where they trade in advance of the blowup, and then exit and reverse that trade in the rescue.

The economic actor Byrne is referring to specifically is Ken Griffin’s Citadel, but generally, it is the dealer. The primary function of a dealer in any market, whether it’s securities, art, cars, or even being a link in a supply chain, is to price liquidity and manage inventory. SALP’s performance was a confession of Leverage + Concentration. Crank the virtuous loop of momentum and flows into thin liquidity, and those dead eyes surface for a look. If understanding liquidity was easy, then market-making would be less profitable. It simply would not be as valuable a service. So we can forgive Leopold for not realizing his gross market value was in shallower waters than he thought.

Market-making is a psychological grind. Again, go to the story from the open. You get paid $10 to flip million-dollar coins, and every now and then you find out you’re on the wrong side of a rigged coin. But even rarer than getting picked off is the chance to feast on fat prey. Now, to be fat, they must have been doing something right, but the weight makes it harder to maneuver than it used to be, and in Leo’s case, “used to be” was quite recent. It only takes a moment of indiscretion to show your belly. Markets are unforgiving because you’re only as sturdy as your worst mistake.

The crocs are always there. Griffin was also there to buy Amaranth out of their positions. Citadel has been in nat gas since the Centaurus era, and with John Arnold retired, Citadel has been an alpha croc in gas trading for well over a decade. If you search my writing, you’ll see a recurring theme of “what equity traders can learn from commodity futures markets”. Futures are zero-sum, so not all the lessons apply, but the ruthlessness will let you borrow a healthy amount of paranoia.

This is @LepoulpePoulpo:

So this was the view I always had wrt equities before
Vs commodities where a lot of shady stuff happens all the time but everyone knows about these games

On the other hand, I’m really less sure now. “It wasn’t certain how close SALP were to a margin call. Wasn’t certain they would have to liquidate in a block” -> I think if you had an idea of their leverage, and saw the price action on all his names, significantly worse than other semis, you could anticipate he would be close to force unwind/liquidate and try to squeeze him.

I admit I don’t know any equities trading team where people would do this kind of thing, but it happens often in commodities.

The word liquidity is a reminder that this is a biological system. Leo priced his trade as if the distribution were exogenous. He would never admit that, but his actions suggest his understanding was purely academic.

Byrne Hobart again:

AI people obsess about existential risk in theory and Leopold has publicly spoken about being aware of it, from a financial standpoint, in practice. But if you want someone who really feels existential risk in their bones, you’re better off talking to a hedge fund manager in Miami.

A mental model for the commodity market I’ve at times lamented and at times celebrated is that for the most part it’s a boring business of blocking and tackling. But now and then a well-capitalized outsider hops Chesterton’s Fence to see if he can force the market to cry uncle. If they’re especially crafty, it can work for a while. You can always beat a dealer on the way in. But you need liquidity to get out and now you don’t have the element of surprise on your side. Eventually, the old illuminati of the business lock arms to go on a hunting expedition. We used to call this “running them in”. If you are forced to cover, there’s little risk to me to bid ahead of you. (That’s not to say the “clean up” isn’t competitive. This is a good thread.)

When the Hunt Brothers cornered silver, the exchange eventually disallowed opening buy orders and raised margin requirements, depleting all the fuel. The exchange used to be owned by traders. They were literally called “members”. You can imagine their position at the top.

The sound of inevitability

In Jurassic Park, Michael Crichton folded an introduction to the field of complexity into the story. The park’s creators’ overconfidence in linear scientific thinking led to disaster. Complexity focuses on chaotic systems like weather (the proverbial butterfly flaps its wings and causes a hurricane across the globe) where models resist equations. There’s a greater emphasis on simulation and higher-order effects. A popular analogy from complexity science is the sandpile. Eventually the sandpile collapses but nobody would say the nth grain of sand causes the avalanche. It simply reveals that the pile angle had become unstable.

When observers consider the timeline (SALP’s peak was likely in late June) they are trying to label the nth grain. Fully embracing the butterfly, here’s my list:

  • the SpaceX IPO
  • the World Cup
  • the uptick in long-term yields
  • box spread rates reflecting funding costs rising as demand for leverage increased, with those costs passed straight through to levered ETFs
  • a slowing trend increasing chop, which increases drag in levered ETFs, which wears down the momos’ patience faster
  • the Knicks winning
  • Kris visits Rome for the first time

Inevitability means none of these matter.

So why was a liquidation inevitable? Why was at least one croc guaranteed a meal?

I already said it. It’s for the same reason LTCM, Hwang, Alameda, and Brian Hunter remain cautionary tales:

Leverage + Concentration.

But what’s so lethal about this combination? Why must it lead to liquidation?

Stated as plainly as possible:

As soon as you assert leverage, you are saying not only am I right, I’m right on timing AND path.

It’s a continuous time parlay. Even if he is right on the destination within the time frame of his choosing he can’t tolerate a large drawdown in the interim.

There’s just no give in the math.

Here’s quant Richard Craib:

But the outcome was never about being right or wrong on AI. At ~150% vol, variance drag alone is ~113%/yr, and risk of ruin is roughly a coin flip over the fund’s life. A child can do the math on a napkin (Claude did it for me: “ruin wasn’t unlikely, it was roughly even money”).

Volatility that high pierces every other fact about a portfolio: the thesis, the timing, the talent. The initial success and the margin call are draws from the same distribution.

And this isn’t really conditioning on the reality that a market’s price discovery function means they will push to a clearing price on a faster schedule than your lender would like. If you didn’t have a lender, this is not a concern!

Martin Shkreli had great coverage on the SALP story on TBPN. But thrown in at the end of the interview is this terrific section:

Kelly famously came up with what is now called the Kelly Criterion. It started as a gambling concept before becoming a finance concept, and it mathematically proves the optimal bet size. The formula is your edge minus the reciprocal of the odds. So, if you have a 55% edge, your optimal bet size is about 10%.

Even that is quite volatile for most people, which is why many investors use half-Kelly or quarter-Kelly sizing. The reality is that most traders don’t actually have an edge, yet they trade as if they have a four- or five-times Kelly edge.

That might sound like they’re simply taking a lot of risk, but if you run the simulation, you’ll go to zero almost every time. The simulator is a really powerful tool because it shows that even if you had a 60/40 edge on every trade—which nobody has in the stock market—you’ll still go bust if you overbet.

That’s a real eye-opener. Position sizing matters just as much as having an edge. It’s something I had to learn the hard way over many years: I was almost always overbetting. I think most hedge funds do it to some extent, and certainly most retail investors do. Very few people actually simulate their portfolios to understand what the appropriate position sizing should be.

After I left the Tiger Cub fund where I worked, I spent a short time in the office of a former SAC Capital (now Point72) portfolio manager. He was one of the best managers I’d ever seen—a quiet guy that almost nobody has heard of, now retired. I had the chance to watch him for a few months before launching my own hedge fund, where I proceeded to do the exact opposite and massively overbet everything.

I group LTCM in with other victims of the Leverage + Concentration poison. As a quant fund, their business was actually to lever diversified edges. But once the correlations of their positions converged, their cocktail was spiked with mathematical Concentration. On the surface, you might say what does a position in corn have to do with Treasury basis, but when a single commingled fund cross-collateralizes its leverage, then their size in the market imports a temporary synchronization of price returns.

Elm Wealth’s Victor Haghani has done a public good by commuting his pain as LTCM partner to teaching the necessity of sound bet sizing. His famous coin-flipping studies show how econ and finance professionals manage to continuously blow up 60/40 advantages by betting far more than what Kelly prescribes.

It gets better. In Fortune’s Formula, I learned that someone betting the prescribed Kelly fraction of their bankroll (which btw implies they have an edge in the first place) has a 50% chance of experiencing a 50% drawdown and a 1/3 chance of experiencing a 50% drawdown before doubling up. In other words, the Kelly fraction is not even conservative. Many traders and gamblers I know will max their betting at half-Kelly.

[The book points out that halving your Kelly fraction will cut your drawdown risk in half but your return only by a quarter so even though your long-term wealth compounds more slowly, the risk-reward is better. That fact alone tells you that the scaling law is extremely punitive if you overbet at all nevermind overbet at the rate Leo was.]

Why Leo, why?

I’m not attacking Leo. I mean he’s very rich, and a bona fide genius. But if the goal is to learn from what we see, I can’t shy from documenting the mistakes despite the optics of seemingly picking on someone half my age.

[His defenders are quick to point out that he’s still up 80% for the year, but all this does is highlight the thin line between zero and hero. We overfit narratives with a comfort that is comically out of tune with what is warranted by circumstance. If Leo loses an extra 25%, an utter blip given his vol, at 4x leverage his investors are zeroed. He made a great call to liquidate, but the presence of liquidity to do so is never a given. In the final hours of a deeply fragile situation, every routine event, hell a Trump tweet, has butterfly potential. Putting yourself in a situation where there is no margin for error is itself a mistake. Leo’s future will revise and buff down the pointy edges of the story’s path dependence, but honesty demands acknowledging that no matter what he becomes, today he is neither lion nor lamb but liquidation was inevitable. He is a man alternating as we do between grace and folly, with neither ever being our full legacy.]

Let’s proceed.

Structure Mistakes

Matt Levine:

If you are all-in on this thesis, you might be more than all-in on this thesis. You won’t put 100% of your money (and your investors’ money) into the AI boom. You’ll put, like, 300% of your money into the AI boom. You’ll borrow money to lever up your bets on the AI boom. As your AI stocks go up, you’ll borrow more to buy more. Getting a 200% return on your money by buying SK Hynix stock is great, but getting a 1,000% return on your money requires borrowing more money to buy more stock. This is a naturally long-term trade. Aschenbrenner’s famous June 2024 essay series is titled “Situational Awareness: The Decade Ahead.” The point is not, like, “SK Hynix will beat earnings expectations next quarter”; the point is stuff like “by the end of the decade, we are headed to $1T+ individual training clusters, requiring power equivalent to >20% of US electricity production.”

You have a vision of the future and want to make a fortune; you need to match your funding to the duration it will take to see the thesis play out. The use of recourse leverage, subject to daily revaluation, is wholly inconsistent with the horizon.

He must know this.

That’s why companies issue equity. They don’t want to worry about the next loan payment. The duration of the financing and vision are aligned. A business that cannot fund ops from cash flows is depleting capital and will need to issue more equity. In that sense, its leverage is not reevaluated every day its beholden to investor appetites at discrete points when it refinances. Meanwhile, a self-sustaining profit machine is more like permanent capital. If it doesn’t like investor bids, it can create its own liquidity by buying itself back.

There was a failure in appreciating structure. The vehicle you use to express a vision is no less important than the vision itself. Founders Fund is Peter Thiel’s GOAT-level investing vehicle, which uses locked-up capital to invest in private companies. Meanwhile, Clarium, his hedge fund that invested in public markets, lost 90% and closed in the wake of the GFC. That a hyper-opinionated genius could succeed and fail so loudly in seemingly similar tasks should alert you to the nature of edge, its prerequisites, and limitations with respect to how you express it.

Hubris?

The most famous Leopold I knew of before Aschenbrenner was another genius.

Nathan Leopold.

In bullet form:

  • First words at four months and three weeks
  • Studied fifteen languages, claimed five fluently.
  • Graduated in his teens Phi Beta Kappa at Chicago, headed for Harvard Law.
  • A nationally recognized ornithologist at nineteen

Enamored with Nietzsche’s Übermensch (“supermen”) as transcendent individuals with superior intellect, Leopold wrote to his a precocious friend Loeb, that such a man is “exempted from the ordinary laws which govern men.”

They conspired to get away with the perfect murder as proof and tribute to their superiority. They spent 7 months planning the abduction, disposal, and even a ransom demand purely as misdirection.

All this only to be caught by eyeglasses dropped near the body. While the glasses had an ordinary prescription and an ordinary frame, they featured an unusual hinge sold to three customers in Chicago. One was Leopold.

Kelly math would have been trivial to Aschenbrenner by the time he was 10. He probably would have used the word “trivial”.

Ed Thorp, another genius, was able to connect the dots from John Kelly’s equation to its use in investing, effectively inventing the world’s first quant fund (which incidentally seeded Ken Griffin when Thorp shut down and gave Ken all his documents since he saw Ken knew what to do with it all). But I’m increasingly of the mind that Thorp’s genius also included suppressing his own ego enough to take Kelly seriously. It’s an intersection of classical genius and wisdom which itself needn’t be so rare. It’s neither here nor there, but I think the public recognizes that the type of genius we are getting out of Silicon Valley is far narrower than the Thorpian variety.

[Related: The connection and friendship between Thorp and Buffett, whose approach to investing was vastly different, was one of the audience’s favorite parts of the talk I gave at Arbor].

Back to Shkreli referring to a trader he once worked with:

What amazed me was that he managed roughly $300–400 million of his own capital but almost never used it. Eighty to ninety percent of the portfolio was simply cash. He would make these tiny trades—little nibbles—and over more than 20 years, I don’t think he ever had a down quarter. He generated 20–30% annual returns while barely putting capital at risk.

It was an incredible lesson. Then, of course, the moment I got the opportunity to manage capital myself, I was running eight times leverage. Looking back, it was one of the dumbest things I could have done. You live and you learn.

Apparently, you can’t learn risk management. You can only live it. Allocators take note.

The allocator’s mistake

Speaking of allocators…why do they keep falling for this grand thesis routine?

I mean, what makes us human, right?

We love stories. We need stories. We want to see athletes fly. We want the impossible dream. We love that truth is stranger than fiction, after all, fiction is restrained by its need to make sense.

The optimism required to back the impossible is the same optimism required to attempt the impossible. Leopold’s investors were not teachers’ pensions and bean counters. It was his singularity-pilled entrepreneur brethren.

Even if he zeroed the damage would be contained to those who can afford it, so my view is no harm, no foul.

Just to share a personal thought.

I’m deeply uninterested in whiz kids that are too cool for the mundane. I strayed from this once and was burned by giving money to some hotshot who I have no doubt is a genius. It wasn’t a fraud or blowup. Hell, it’s still operating as a respectable, institutionally-approved fund. It’s just expensive mediocrity. If I wanted that, I could find some value fogies quoting Cicero.

I went against my better judgement.

Where is the repurposed crusty trader who’s been turned upside down a few times but never had a losing year, even if sometimes it’s a T-bill? I don’t need him originating the ideas, but I need him to call you an idiot and ask hard questions because he’s just as dubious of smart kids as he is of the government. When you use all your brilliance to dress a pump-and-dump in new tech and memes, he reminds you the tail you’re selling is tied to a statute of limitations, and anything like that is a non-starter.

Gimme the corny words. A warden of capital. A trustee. A fiduciary. At least they have a chance of not being a grift. I still need to figure out if they’re made of what I want at the point of sale and for monitoring the books. Steady hands. Paranoid. Path-aware. Fat Tony. If it leads with sexy, get me outta here.

I want the C in CAGR because I want to minimize drag.

I want the C in curmudgeon because you’ve heard the line:

There are old pilots and bold pilots, but no old, bold pilots.

We have been in a regime where many of the people who can raise money have built returns and stories in a boom environment, but those environments paper over bad habits, which increases the allocator’s adverse selection risk.

If you’re giving someone money for the long run, you want a battle-tested framework honed by the drudgery of risk monitoring, outtrades, system outages, and the paranoia from the memory of a hardcoded number in a spreadsheet getting you picked off.

Nobody is bigger than the market

A lesson we learn again and again, is that nobody is bigger than the market. Not even the crocs. They survive because they respect it. They’ve seen so much in the course of both providing liquidity and also occasionally taking it to manage risk, that they can have no other relationship to it other than respect.

That means keeping concentration away from leverage. There’s no price worth giving up control of your fate.

There seems to be something deep about risk management that eludes genius alone.

  • Leopold couldn’t have been ignorant of betting math.
  • Leopold should have been able to understand structure.
  • Leopold knew growth would require getting thesis, path, and timing right.

I’m left to conclude that the ancient root of most major errors is at hand. Hubris.

I’m even more confident in this because history tells us that being smart doesn’t inoculate you from hubris and, as the murdering Leopold story suggests, can actively fuel it. But since I have never spent a day being a genius, any more than I jump like Jordan or sing like Sinatra, my sense of limitation leaves me unable to empathize with Leopold’s blind spot.

This is a point of encouragement to everyone.

When you are limited, you seek approaches in light of your limitations, which explains the reality we seek all around us. That the quality of our decisions which determines flourishing has no relationship with excessive intelligence.

Since investing and managing money is a decision overlay that sits on top of research and analysis functions, the manager’s efficacy is rate-limited not by brains but by wisdom.

We should be a bit more obsessed with where wisdom comes from than continue to fall for dazzling minds.

If there’s a bit of poetry in this episode, it’s that Leopold’s only out was a singularity bigger than the market’s boring old constructs like liquidity and collateral. Maybe he could have gotten there. And he still has a chance. His mistake was turning it into a race with the crocs.

If you don’t mix Leverage + Concentration, you never have to get in the water.