Taylor Expansion โ€” Cheatsheet


The idea in one sentence

Stand at one spot on a curve, measure everything you can there, and use those measurements to guess the height somewhere else.

You’d do this when the curve is hard to compute everywhere but easy to measure at one point, or when you want to see what drives a change. A bond’s price is a familiar case: you know its yield and duration today and want to know what happens if rates move.

The guess is built in layers. Each layer uses one more thing you measured at the anchor. The first layer is a straight line. The second bends it. The third bends the bend. You stop when the layers stop mattering, or when you run out of measurements.


Worked example: guess xยณ at x = 1.2 using only what you know at x = 1

The curve is y = xยณ. You are standing at x = 1. The true answer is 1.2ยณ = 1.728, but pretend you can’t compute it.

What you know at x = 1

ThingHow you get itValue at x = 1
Heightxยณ1
Slopederivative, 3xยฒ3
How fast the slope changesderivative of that, 6x6
How fast that changesderivative of that6
Anything furtherderivative of a constant0

The walk. Destination minus start: 1.2 โˆ’ 1 = 0.2. Call it h.

Layer 1: pretend the slope stays 3. Guess = height + slope ร— walk = 1 + 3 ร— 0.2 = 1.6. Off by 0.128.

Layer 2: the slope drifts. It goes up at 6 per unit, so over the walk it rises from 3 to 3 + 6 ร— 0.2 = 4.2. Use the average slope, 3.6, instead of 3. Guess = 1 + 3.6 ร— 0.2 = 1.72. Off by 0.008.

Written as a separate correction: the new piece is ยฝ ร— 6 ร— 0.2ยฒ = 0.12, added to the 1.6.

Layer 3: the drift rate drifts. Same move one level deeper. The correction is 6 ร— 0.2ยณ รท 6 = 0.008. Guess = 1.728. Off by exactly 0.

Layer 4 and beyond: the measurement is 0, so every further correction is 0.

The guess is now perfect at x = 1.2, and it’s perfect at every other x too. xยณ only has three pieces of information in it. Use all three and you have rebuilt the function.

Try it at x = 3 (walk h = 2): 1 + 3(2) + 3(2)ยฒ + (2)ยณ = 1 + 6 + 12 + 8 = 27 = 3ยณ. Still exact, even two units from the anchor.

Three layers rebuild xยณ exactly, everywhereHeight y; the dashed vertical is the anchor x = 1-50510152025-10123xxยณ = layer 3 27layer 1 7layer 2 19

Layer 1 is the tangent line. Layer 2 bends it into a parabola that hugs the curve near x = 1 but misses on both sides. Layer 3 (dashed) lies on top of the black xยณ curve at every x, which is the whole point: for a polynomial, enough layers is exactly the function.


Why you divide by 2, then 6, then 24

Each layer’s measurement gets divided before it’s used: layer 1 by 1, layer 2 by 2, layer 3 by 6, layer 4 by 24. Those are 1!, 2!, 3!, 4!. Two ways to see why.

The averaging picture. Layer 2 used the average slope over the walk. The slope was a ramp going from 3 to 4.2, and the average of a ramp is halfway: that’s the รท2. Layer 3 needs the average of something that grows like a parabola, and a parabola from zero spends most of the walk being small, so its average is only a third of its end value: another รท3. Stack them: 2 ร— 3 = 6. One more level and the average of a cubic is a quarter: 2 ร— 3 ร— 4 = 24.

Check the parabola claim with numbers. Sample tยฒ at t = 0.1, 0.2, โ€ฆ, 1.0: you get 0.01, 0.04, 0.09, 0.16, 0.25, 0.36, 0.49, 0.64, 0.81, 1.00. They add to 3.85, average 0.385, and with finer sampling it settles to โ…“.

The power-raising picture. Differentiating hโฟ gives n ร— hโฟโปยน: lowering a power by one multiplies by that power. So raising a power by one divides by it. To turn a constant measurement into a term in hยณ you raise the power three times, paying รท1, รท2, รท3 along the way. That product is 3! = 6.

LayerMeasurement (xยณ at 1)Divide byTerm
1313h
2623hยฒ
366hยณ
40240

Worked example: ln x, where the corrections stop helping

Same game, anchor at x = 1. Height is ln 1 = 0. Slope is 1/x, so 1 at the anchor.

What you know at x = 1

LayerDerivativeValue at 1Divide byTerm
11/x11h
2โˆ’1/xยฒโˆ’12โˆ’hยฒ/2
32/xยณ26hยณ/3
4โˆ’6/xโดโˆ’624โˆ’hโด/4
524/xโต24120hโต/5
6โˆ’120/xโถโˆ’120720โˆ’hโถ/6

The measurements never hit zero. They alternate sign and grow. So there is always another correction, forever. Whether the corrections help depends on how far you walk.

Short walk: x = 1.5, so h = 0.5. True value ln 1.5 = 0.4055.

Layers usedGuessGap
10.50.0945
20.3750.0305
30.41670.0112
40.40100.0044
50.40730.0018
60.40470.0008

Each layer roughly halves the gap. Keep going and it heads to zero.

Long walk: x = 2.5, so h = 1.5. True value ln 2.5 = 0.9163.

Layers usedGuessGap
11.50.5837
20.3750.5413
31.50.5837
40.23440.6819
51.75310.8368
6โˆ’0.14531.0616

The gap gets worse. Each term is ยฑhแต/k, and with h = 1.5 the hแต grows faster than the k can shrink it: 1.5, 1.125, 1.125, 1.27, 1.52, 1.90, โ€ฆ The guess swings wider and wider around the truth.

Past x = 2, more layers make the ln x guess worseHeight y; anchor at x = 1; shaded band = radius of convergence, 0 to 2layers helplayers hurt-3-2-101230.511.522.533.54xlayer 1layer 2layer 3layer 4layer 5layer 6ln x

Inside the band the colored curves pile onto the black one, each layer tighter than the last. Right of x = 2 they fan out: layer 3 shoots up, layer 4 dives, layer 5 shoots higher, layer 6 dives harder. Every extra layer swings further from ln x instead of closer.

The rule. For ln x about 1, the corrections help when |h| < 1 and hurt when |h| > 1. That distance, 1, is the radius of convergence. It’s set by where the function itself breaks: ln x blows up at x = 0, exactly one unit left of the anchor, and the series can’t reach further right than it can reach left.

xยณ had no such limit because its corrections ran out before they could misbehave. That is the difference between a polynomial and everything else.


The formula, decoded

Everything above, written the standard way:

Pโ‚™(x) = ฮฃโ‚–โ‚Œโ‚€โฟ fโฝแตโพ(xโ‚€) / k! ยท (x โˆ’ xโ‚€)แต

Spelled out for the first few terms:

Pโ‚™(x) = f(xโ‚€) + fโ€ฒ(xโ‚€)(x โˆ’ xโ‚€) + fโ€ณ(xโ‚€)/2 ยท (x โˆ’ xโ‚€)ยฒ + fโ€ด(xโ‚€)/6 ยท (x โˆ’ xโ‚€)ยณ + โ‹ฏ

Every symbol, in the order it appears:

SymbolRead it asWhat it isIn the xยณ example
f“the function”the curve you’re guessing; f(x) is its height at xxยณ
x“x”the destination, any point on the horizontal axis1.2
xโ‚€“x-nought”the anchor, where you stood and took measurements1
x โˆ’ xโ‚€“the walk”destination minus start, also written h0.2
P“the polynomial”the guess; a polynomial because it’s a sum of powers of the walk1 + 3h + 3hยฒ + hยณ
n“n”how many layers you used; the highest power in the guess3
Pโ‚™(x)“P-n of x”the guess using n layers, evaluated at xPโ‚ƒ(1.2) = 1.728
k“k”the counter: which layer you’re on, running 0, 1, 2, โ€ฆ up to n0, 1, 2, 3
ฮฃโ‚–โ‚Œโ‚€โฟ“sum from k = 0 to n”add up the term for every k from 0 through nfour terms
fโ€ฒ, fโ€ณ, fโ€ด“f-prime, double-prime, triple-prime”first, second, third derivative: slope, rate of slope, rate of that3xยฒ, 6x, 6
fโฝแตโพ“f-k”the k-th derivative; fโฝโฐโพ is f itself, fโฝยนโพ is fโ€ฒ, and so onfโฝยฒโพ = 6x
fโฝแตโพ(xโ‚€)“f-k at x-nought”the k-th derivative evaluated at the anchor, a plain number1, 3, 6, 6
k!“k factorial”1 ร— 2 ร— โ€ฆ ร— k, the divide-by column; 0! = 11, 1, 2, 6
(x โˆ’ xโ‚€)แต“the walk to the k”the walk raised to the layer number1, 0.2, 0.04, 0.008
f(x) โˆ’ Pโ‚™(x)“the gap”true height minus guess0
ฮพ“xi” (Greek letter)some unknown point between xโ‚€ and x; used only in the error formulasomewhere in [1, 1.2]

So the k = 2 term of the sum is: take the second derivative (6x), evaluate at the anchor (6), divide by 2! (3), multiply by the walk squared (0.04). That’s 0.12, the layer-2 correction from the worked example.

How big is the gap? It’s controlled by the next measurement you didn’t use, taken somewhere along the walk:

f(x) โˆ’ Pโ‚™(x) = fโฝโฟโบยนโพ(ฮพ) / (n+1)! ยท (x โˆ’ xโ‚€)โฟโบยน for some ฮพ between xโ‚€ and x

Read it as: the gap is small when the walk is short (the hโฟโบยน is tiny), when the next derivative is tame, or when n is large enough that (n+1)! dominates. The gap is large when the walk is long and the higher derivatives are big, which is exactly what happened to ln x at x = 2.5.


Real-world uses

Four places you’ve met this without the name. Each is worked with numbers.

Bond prices: duration and convexity. A 10-year zero at a 4% yield is priced 100/1.04ยนโฐ = 67.56. The anchor is 4%. The two measurements are duration (layer 1, slope) = 10/1.04 = 9.62 and convexity (layer 2) = 10 ร— 11/1.04ยฒ = 101.7.

Yield rises 1%, so the walk is 0.01:

LayersGuessTrue priceGap
1 (duration only)67.56 ร— (1 โˆ’ 9.62 ร— 0.01) = 61.0661.390.33
2 (add convexity)61.06 + 67.56 ร— ยฝ ร— 101.7 ร— 0.01ยฒ = 61.4061.390.01

Yield rises 3%, walk 0.03: duration alone says 48.07, convexity pulls it to 51.16, true is 50.83. The gap is 30 times larger than for the 1% move. That’s the long walk. Traders quote duration and convexity for exactly the reason your tool quotes delta and gamma.

How a calculator computes sin. There’s no sin key inside the chip; it sums the series about 0: x โˆ’ xยณ/6 + xโต/120 โˆ’ xโท/5040 + โ€ฆ.

sin(0.5): 0.5 โˆ’ 0.0208 + 0.0003 = 0.4794. True value 0.4794. Three terms.

sin(3): 3 โˆ’ 4.5 + 2.025 โˆ’ 0.434 + 0.050 โˆ’ 0.004 = 0.137. True value 0.141. Six terms and still off in the third decimal. The series converges everywhere, unlike ln x, but a long walk needs many more layers. Calculators dodge this by folding the input back to a small angle first, which is the “walk less” strategy.

Pendulum clocks. The textbook period T = 2ฯ€โˆš(L/g) comes from replacing sin ฮธ with ฮธ, which is layer 1 of the sine series. The next layer says the true period is longer by a factor of about 1 + ฮธโ‚€ยฒ/16, where ฮธโ‚€ is the swing in radians.

Swingฮธโ‚€ in radiansCorrectionPeriod error if you ignore it
5ยฐ0.0871.00050.05%
20ยฐ0.3491.00760.8%
60ยฐ1.0471.0697%

A clock built on layer 1 keeps time at small swings and drifts at big ones. Same story: the anchor is ฮธ = 0 and the walk is the amplitude.

Compound growth. (1 + r)โฟ about r = 0 is 1 + nr + n(nโˆ’1)rยฒ/2 + โ€ฆ. For 5% over 10 years: layer 1 says 1.50, layer 2 says 1.50 + 45 ร— 0.0025 = 1.61, true is 1.63. Layer 1 alone is the “simple interest” mental shortcut, and the layer-2 term is exactly how much compounding beats it.

All four break the same way ln x did. Past some size of move the layers you kept stop describing the function, and the fix is either more layers, a shorter walk, or computing the real thing.


Quick reference

Recipe for any function about any anchor

  1. Pick the anchor xโ‚€. Compute the walk h = x โˆ’ xโ‚€.
  2. Take derivatives of f until you have as many as you want, and evaluate each at xโ‚€.
  3. Divide the k-th one by k!, multiply by hแต.
  4. Add them up. That’s the guess. The leftover is the gap.
  5. If the derivatives hit zero, the guess becomes exact. If they don’t, check whether the terms are shrinking; if not, you walked too far.

The two examples side by side

xยณ about 1ln x about 1
Derivatives at anchor1, 3, 6, 6, 0, 0, โ€ฆ0, 1, โˆ’1, 2, โˆ’6, 24, โ€ฆ
k-th term3h, 3hยฒ, hยณ, then 0(โˆ’1)แตโบยน hแต / k
Exact after3 layersnever
Works forevery x0 < x < 2 only
Whypolynomial: information runs outln x breaks at 0, one unit from the anchor

Factorials

kk!Average of tแต over [0, h]
11h/2
22hยฒ/3
36hยณ/4
424hโด/5
5120hโต/6

Common series about 0, for reference

FunctionSeriesConverges for
eหฃ1 + x + xยฒ/2 + xยณ/6 + โ€ฆall x
sin xx โˆ’ xยณ/6 + xโต/120 โˆ’ โ€ฆall x
cos x1 โˆ’ xยฒ/2 + xโด/24 โˆ’ โ€ฆall x
1/(1โˆ’x)1 + x + xยฒ + xยณ + โ€ฆ|x| < 1
ln(1+x)x โˆ’ xยฒ/2 + xยณ/3 โˆ’ โ€ฆโˆ’1 < x โ‰ค 1

The last two have a radius because the function breaks at x = 1 or x = โˆ’1. The first three never break, so the series works everywhere even though it never terminates.


Key Insights

WhatWhy it matters
Anchor and walkEverything is measured at one point; the guess only ever knows about that point
Layers = derivativesEach derivative at the anchor buys one more correction; that is all the information you have
Divide by k!Raising a power costs a division each time; averaging a ramp, parabola, cubic costs รท2, รท3, รท4
Polynomials terminateDerivatives hit zero, so finitely many layers rebuild the function exactly, everywhere
Radius of convergenceFor everything else, past the distance to the nearest breakdown the layers make the guess worse
The gap formulaThe error is the first term you dropped, evaluated somewhere on the walk

soulcraft

My second favorite thing after a piece of art, writing, music, movie, and sport that I love is watching someone else gush over things they love. Itโ€™s why the Lost in Vegas guys are so endearing, especially when the channel first started. You were watching hip-hop heads not only discover but unpeel the layers in Toolโ€™s music.

One of my dad friends recently bought a 1979 Jeep to restore with his son and asked if my son Zak (13) would like to help. Zakโ€™s โ€œhell yeaโ€ beat the sound of my last syllable when I asked him. He started watching YouTube, and we landed on a channel where these obsessed details comb the U.S. looking for โ€œbarn findsโ€ to clean for the owners. For free! Itโ€™s free because the channel has almost 2 million subscribers clamoring to watch an impossible cleaning.

I watched a few videos. I get it. Renewal of beauty, just like the objects of renewal never goes out of style. The process is as timelessly rewarding as the result.

Caitlinโ€™s tweet is an appeal to our spirit of obsession and craft.

I like โ€˜em thick (Adam Mastroianni)

This essay not only feels good but is an important lens on the fast-approaching experiment of whether infinite electric monkeys will pound out Shakespeare.

Normally I might apologize for extensive excerpting but they are the point. I can do no better than this.

It opens:

I owe an apology to every English teacher I ever had. I always assumed that so-called โ€œgreatโ€ literature was a hoax, a punishment inflicted upon adolescents for the crime of being young. These books did not have anything special about them, and covering up that fact was simply a make-work exercise for former English majors, a sort of โ€œjobs for snobsโ€ program.

I was wrong about this. There is such thing as greatness. More specifically, there is such thing as thickness. Great works of fictionโ€”for that matter, great works of any artโ€”unfurl in response to your attention. The more time you spend with them, the more you get out of them. That kind of responsiveness is so addicting that it can lead people to do crazy things, like try to teach literature to high schoolers.

But thickness is tricky, because rewarding the careful reader often means repelling the casual one. And this is where I would like an apology in return from my English teachers, because while this might have been obvious to them, they never made it obvious to me.

I was presented with art and literature as if it was self-explanatory, and that everything wonderful about it was plainly visible from the outside. But those works were much more like dark, winding caves with treasure stashed inside of them. My teachers were like, โ€œRight, well, into the cave you go!โ€ and I was like โ€œBut thereโ€™s nothing in thereโ€ and they were like โ€œEntering the cave is 30% of your gradeโ€ and so I took a few steps into the darkness and I was like โ€œJust as I suspected: an empty caveโ€ and then I came trudging back out and pretended that I saw something.

It goes on to describe spectacular examples of โ€œthicknessโ€. You wonโ€™t want to miss the The Garden of Earthly Delights painting and its unintended invitation to hear โ€œbutt musicโ€.

Adam presents 4 qualities that make something thick with examples from a childrenโ€™s book, the โ€œearsโ€ in Hamlet, and Penn (of Penn and Teller fame) eating fire.

Thereโ€™s an amazing takedown of what passes for popular non-fiction.:

Reading a book like this feels like wandering through a Potemkin village. Touch any of the ideas, and they tip over.

He contrasts gilded examples of non-fiction with solid gold:

Thickness comes from surfacing a few facts well, and in such a way that you realize the existence of entire universes of additional facts that could be knownโ€ฆ.[Jane Jacobs book] is pointing out a fact that millions of people observe every day, but almost none of them notice.

He addresses an easily anticipated objection which if youโ€™ll be familiar with if you have read any of deBoer on โ€œpoptimismโ€:

If you allow for the existence of secret treasures that can only be accessed with effort and analysis, then you empower the elitists and the snobs. โ€œBuddy, donโ€™t even talk to me until youโ€™ve been in the cave!โ€

But look around. The snobs are in full retreat. We have swung the pendulum so far toward poptimism, toward the blinkered idea that all art is equal because all humans are equal, toward the ethos that guilty pleasures are simply pleasures, that Iโ€™m not sure if we can ever swing it back.

And finally, Adam addresses the elephant:

Erasing the line between the thick and the thin has left us defenseless against slop at the exact moment of its onslaught. Everyone can sense thereโ€™s something amiss with the prose that comes out of the machines, but we lack the language to talk about it, and so weโ€™ve converged on the idea that slop simply means using too many em dashes, bullet points, and line breaks.

What separates substance from slop is thickness. Slop holds no secrets; it signifies nothing. Under scrutiny, it evaporates. All it can offer is bottomlessnessโ€”sure, thereโ€™s nothing on, but at least there are infinite channels!

Thatโ€™s why Iโ€™m neither surprised nor dismayed when studies find that people prefer AI art to human art. Of course they do! In the short term, thinness prevails. When people are making snap judgments, they want pretty flowers, poems that rhyme, pleasing pablum, the simulacrum of thought. But none of these lastโ€ฆ

About once a week, I get a pitch from some AI startup that wants to automate some part of my writing. The most recent one says itโ€™s โ€œbuilt for credible thinkers who have a bookโ€™s worth of ideas but not the time it typically takes to write oneโ€.

Iโ€™m sorry, but if youโ€™re building or using a tool like this, then youโ€™ve got slop for brains. There is no such thing as having a โ€œbookโ€™s worth of ideasโ€ that are all ready to go except for the small matter of choosing the right words and putting them in the right order.

I know exactly the feeling that these slop-trepeneurs are preying on, because I feel it all the time: Iโ€™ve got these thoughts in my head, and boy oh boy theyโ€™re good ones, all-timers, really, and itโ€™s so annoying that I have to spend all this time making the words sound good, when the ideas behind the words are already so good!

But this is an illusion. The ideas are not already good. They need to be thickened. I understand why itโ€™s tempting to force a machine do the hard part for you, but it canโ€™t, and the hard part is the only part worth doing anyway.

Moontower #329

In this issue:

  • soulcraft
  • grab a coffee and derive this with me

Friends,

My second favorite thing after a piece of art, writing, music, movie, and sport that I love is watching someone else gush over things they love. Itโ€™s why the Lost in Vegas guys are so endearing, especially when the channel first started. You were watching hip-hop heads not only discover but unpeel the layers in Toolโ€™s music.

One of my dad friends recently bought a 1979 Jeep to restore with his son and asked if my son Zak (13) would like to help. Zakโ€™s โ€œhell yeaโ€ beat the sound of my last syllable when I asked him. He started watching YouTube, and we landed on a channel where these obsessed details comb the U.S. looking for โ€œbarn findsโ€ to clean for the owners. For free! Itโ€™s free because the channel has almost 2 million subscribers clamoring to watch an impossible cleaning.

I watched a few videos. I get it. Renewal of beauty, just like the objects of renewal never goes out of style. The process is as timelessly rewarding as the result.

Caitlinโ€™s tweet is an appeal to our spirit of obsession and craft.

Which brings me to a luxurious reading experience:

I like โ€˜em thick (Adam Mastroianni)

This essay not only feels good but is an important lens on the fast-approaching experiment of whether infinite electric monkeys will pound out Shakespeare.

Normally I might apologize for extensive excerpting but they are the point. I can do no better than this.

It opens:

I owe an apology to every English teacher I ever had. I always assumed that so-called โ€œgreatโ€ literature was a hoax, a punishment inflicted upon adolescents for the crime of being young. These books did not have anything special about them, and covering up that fact was simply a make-work exercise for former English majors, a sort of โ€œjobs for snobsโ€ program.

I was wrong about this. There is such thing as greatness. More specifically, there is such thing as thickness. Great works of fictionโ€”for that matter, great works of any artโ€”unfurl in response to your attention. The more time you spend with them, the more you get out of them. That kind of responsiveness is so addicting that it can lead people to do crazy things, like try to teach literature to high schoolers.

But thickness is tricky, because rewarding the careful reader often means repelling the casual one. And this is where I would like an apology in return from my English teachers, because while this might have been obvious to them, they never made it obvious to me.

I was presented with art and literature as if it was self-explanatory, and that everything wonderful about it was plainly visible from the outside. But those works were much more like dark, winding caves with treasure stashed inside of them. My teachers were like, โ€œRight, well, into the cave you go!โ€ and I was like โ€œBut thereโ€™s nothing in thereโ€ and they were like โ€œEntering the cave is 30% of your gradeโ€ and so I took a few steps into the darkness and I was like โ€œJust as I suspected: an empty caveโ€ and then I came trudging back out and pretended that I saw something.

It goes on to describe spectacular examples of โ€œthicknessโ€. You wonโ€™t want to miss the The Garden of Earthly Delights painting and its unintended invitation to hear โ€œbutt musicโ€.

Adam presents 4 qualities that make something thick with examples from a childrenโ€™s book, the โ€œearsโ€ in Hamlet, and Penn (of Penn and Teller fame) eating fire.

Thereโ€™s an amazing takedown of what passes for popular non-fiction.:

Reading a book like this feels like wandering through a Potemkin village. Touch any of the ideas, and they tip over.

He contrasts gilded examples of non-fiction with solid gold:

Thickness comes from surfacing a few facts well, and in such a way that you realize the existence of entire universes of additional facts that could be knownโ€ฆ.[Jane Jacobs book] is pointing out a fact that millions of people observe every day, but almost none of them notice.

He addresses an easily anticipated objection, which youโ€™ll be familiar with if you have read any of deBoer on โ€œpoptimismโ€:

If you allow for the existence of secret treasures that can only be accessed with effort and analysis, then you empower the elitists and the snobs. โ€œBuddy, donโ€™t even talk to me until youโ€™ve been in the cave!โ€

But look around. The snobs are in full retreat. We have swung the pendulum so far toward poptimism, toward the blinkered idea that all art is equal because all humans are equal, toward the ethos that guilty pleasures are simply pleasures, that Iโ€™m not sure if we can ever swing it back.

And finally, Adam addresses the elephant:

Erasing the line between the thick and the thin has left us defenseless against slop at the exact moment of its onslaught. Everyone can sense thereโ€™s something amiss with the prose that comes out of the machines, but we lack the language to talk about it, and so weโ€™ve converged on the idea that slop simply means using too many em dashes, bullet points, and line breaks.

What separates substance from slop is thickness. Slop holds no secrets; it signifies nothing. Under scrutiny, it evaporates. All it can offer is bottomlessnessโ€”sure, thereโ€™s nothing on, but at least there are infinite channels!

Thatโ€™s why Iโ€™m neither surprised nor dismayed when studies find that people prefer AI art to human art. Of course they do! In the short term, thinness prevails. When people are making snap judgments, they want pretty flowers, poems that rhyme, pleasing pablum, the simulacrum of thought. But none of these lastโ€ฆ

About once a week, I get a pitch from some AI startup that wants to automate some part of my writing. The most recent one says itโ€™s โ€œbuilt for credible thinkers who have a bookโ€™s worth of ideas but not the time it typically takes to write oneโ€.

Iโ€™m sorry, but if youโ€™re building or using a tool like this, then youโ€™ve got slop for brains. There is no such thing as having a โ€œbookโ€™s worth of ideasโ€ that are all ready to go except for the small matter of choosing the right words and putting them in the right order.

I know exactly the feeling that these slop-trepeneurs are preying on, because I feel it all the time: Iโ€™ve got these thoughts in my head, and boy oh boy theyโ€™re good ones, all-timers, really, and itโ€™s so annoying that I have to spend all this time making the words sound good, when the ideas behind the words are already so good!

But this is an illusion. The ideas are not already good. They need to be thickened. I understand why itโ€™s tempting to force a machine do the hard part for you, but it canโ€™t, and the hard part is the only part worth doing anyway.


Money Angle

People seemed to like the post I wrote 2 weeks ago teaching readers how to compute Kelly optimal bet sizes in their heads. A lot of people reached out saying that despite learning Kelly in the past, this treatment not only made it clearer but also helped them appreciate how its approach informs risk-taking in wider contexts.

Panoptica graciously asked me to republish it under their own banner:

After this post you will be sizing bets in your head (Panoptica)

Just to put a bow on it, Iโ€™ll restate what I think are the most crucial lessons without dwelling on the formula:

  1. Even educated people are terrible at sizing bets. Itโ€™s not because itโ€™s so complex, but I guess itโ€™s like squatting. It seems like you should just know how to do it, but itโ€™s actually something you need to learn the mechanics of.
  2. Overbetting is incinerating money. This is something thatโ€™s hard to appreciate until you see the math. The reason you size smaller is because of the asymmetry of being wrong on your edge. If you underbet, you slow your growth rate but slow risk even faster. At least you’re exchanging lower returns for a better risk/reward. But if you overbet, you lower your growth rate AND increase your risk even faster than you reduce your reward. Both the numerator and denominator of your risk/reward move in the wrong directions!

If interested, Matt does this neat meta series “Notes on Notes”. It’s a short chat about how and why a particular article comes together in the first place. You can watch it here:

via Matt Zeigler@CultishCreative

A short yet broad-ranging talk with @KrisAbdelmessih on @Panoptica_ai for his @EpsilonTheory: Unplugged essay, “After this post you will be sizing bets in your head” Kelly bets, using AI to learn math, creativity… NEW Notes on Notes!

Article content

To round out this Kelly sprint, I have 2 more bits that are again overtures to curious learners who may find math intimidating.

1) Slides

The first is a condensed slide version, which I hope makes this accessible. It was born out of teaching Kelly to my 8th grader at breakfast this past Tuesday. โ€œZak, I wanna see if I can teach you something neat in 5 minutes.โ€ He rolled his eyes, but at least he humored me while scarfing down his cereal.

๐Ÿ“ŠMoontower Kelly Slides

2) An empowering derivation

The derivation of the Kelly formula is so fun because as it rolls downhill, a number of concepts we talk about in this letter stick to it, so when we get to the end it feels like something grand, but itโ€™s so damn compact.

Another teaching experiment. Letโ€™s see if I can narrate the derivation in a way so that as you follow along it never feels โ€œhardโ€. I want to prove that this is fun to do and while I donโ€™t expect to convert everyone, I do think thereโ€™s a bunch of you whoโ€™d like to be able to learn this but feel blocked because you have gaps in your foundations or canโ€™t remember HS math, or just lack some confidence.

Screw all that, Iโ€™ll lay my jacket over the puddles so you can see that itโ€™s not so bad out here. Just come along.

Money Angle For Masochists

Deriving Kelly From Scratch

Stating the question

You have a bet. You win with probability p and lose with probability q = 1 โˆ’ p. If you win, you get paid B times what you risked. If you lose, you lose what you risked. B is also just a return. So if you double your money on a bet, B=1=100% return.

Youโ€™re going to bet the same fraction f of your bankroll every time and let it ride.

Whatโ€™s the f that yields the highest compounded return?

Step 1: What one bet does to your wealth

If your wealth is W and you bet the fraction f:

Win: W โ†’ W(1 + Bf) Lose: W โ†’ W(1 โˆ’ f)

Example:

Your starting wealth is $100 and you bet 50% of it on a coin flip. Remember B =1 because when you win you make 100%. When you lose you always lose f which is your bet size.

Win: 100 โ†’ 100(1 + 1*.50) = $150 Lose: 100 โ†’ 100(1 โˆ’ .50) = $50

You keep the part you didnโ€™t bet no matter what. A win adds B times your stake. A loss takes your stake away.

Step 2: What many bets do to your wealth

Play n bets. You win h of them and lose the other n โˆ’ h. Each bet multiplies whatever the last one left you, so the multipliers stack:

Wโ‚™ = Wโ‚€(1 + Bf)สฐ(1 โˆ’ f)โฟโปสฐ

The order of wins and losses doesnโ€™t matter. Only the count does.

Example:

I bet 50% of my wealth each turn on the coin. I play 4 times and I win 3 of them.

100(1+1*.5)3(1-.5)1= 168.75

Order doesnโ€™t matter. If I lost 50% on the first flip, then won 3 in a row:

Start: 100 Lose (ร—0.5): 50 Win (ร—1.5): 75 Win (ร—1.5): 112.50 Win (ร—1.5): 168.75

Step 3: Turn total growth into a growth rate

Wโ‚™ is where you end up. We want a per-bet growth rate.

You already know how to do this. If an investment grew by a factor of (1 + 8%)ยนโฐ over 10 years, you take the 10th root to get the CAGR back. Same move here: take the nth root of this equation: Wโ‚™ = Wโ‚€(1 + Bf)สฐ(1 โˆ’ f)โฟโปสฐ

Note that Wโ‚€ can be divided out, just as if your starting identity was 150 = 100(1+8%)ยนโฐ and you turned that into 1.5 = 1.08ยนโฐ before taking the 10th root to get to annual growth rate.

This equation took our simple total growth equation and turned it into a growth rate equation:

Article content

Step 4: Let probabilities take over

We can clean up our new growth equation with several handy notation substitutions.

Ps and Qs

h/n is the share of bets won

(n-h)/n is the share of bets lost

Over a long run, the fraction of bets you win settles down to your win probability:

h/n โ†’ p and (n โˆ’ h)/n โ†’ q

G

Wโ‚™/Wโ‚€ is just a wealth multiple. If you made 200% on your investment portfolio over a decade, your wealth multiple is 3 because you started with $1 and ended up with $3.

Weโ€™ll call that wealth multiple G (for โ€œgross multipleโ€)

So the per-bet growth factor becomes:

G = (1 + Bf)แต–(1 โˆ’ f)แ‘ซ

G = 1.05 means your bankroll grows 5% per bet on a compounded basis. In traditional investing, weโ€™d substitute the word โ€œyearโ€ for โ€œbetโ€. Each flip is like a year.

Our job is to find the f that makes G (the gross multiple of our wealth) as big as possible.

Step 5: Take the log to make the math easy

Maximizing G means taking its derivative with respect to f and setting it to zero.

Article content
when someone says โ€œderivativeโ€

Itโ€™s 2026, you donโ€™t need to know how to actually differentiate an equation. You just have to awaken that part of your brain that knows:

a) a derivative is the slope of a function at a given point on a curve

b) when the slope of a curve is 0 this is a maximum or minimum

c) intuitively, we can reason that this is a growth curve is a hill with no bumps. Its slope starts positive and only ever gets smaller as you bet more, so it can hit zero exactly once, and when it does, youโ€™re at the top. (the jargon version: this growth curve is concave: it bends downward everywhere from the max, with no inflection points)

Article content

We are still here:

G = (1 + Bf)แต–(1 โˆ’ f)แ‘ซ

But G is a product of powers, which is nasty to differentiate.

But remember from simple pleasures, logs fix this. They are the inverse of exponentiation, allowing them to turn multiplication into addition!

Push the cobwebs away to recall that this works:

log(100) = log(10 ร— 10) = log(10) + log(10) = 1 + 1 = 2

Thatโ€™s not all logs do.

They pull exponents down in front: log(xแต–) = p ยท log(x).

We can prep that equation for differentiation by taking the log of both sides and using both of those sexy log features:

log(G) = p ยท log(1 + Bf) + q ยท log(1 โˆ’ f)

Maximizing log(G) gives the same f as maximizing G, because log always rises when its input rises.

Article content
rapidtables.com

Thereโ€™s another little bonus.

Log(G) is the log return per bet, the same quantity as ln(Sโ‚/Sโ‚€).

It re-expresses a compounded, multiplicative growth factor as a continuously compounded rate, and rates add cleanly.

Step 6: Differentiate and set to zero

Once you know this is a derivative problem because you are trying to maximize then we can rely on crutches (Claude) for thing thatโ€™s hard to remember.

Namely that:

  • the derivative of log(x) is 1/x.
  • If thereโ€™s something inside the log, also multiply by the derivative of that inside piece (the chain rule)

So we have our equation:

log(G) = p ยท log(1 + Bf) + q ยท log(1 โˆ’ f)

then we differentiate our terms to be added with respect to f:

  • p ยท log(1 + Bf) โ†’ pB/(1 + Bf) (the inside, 1 + Bf, has derivative B)
  • q ยท log(1 โˆ’ f) โ†’ โˆ’q/(1 โˆ’ f) (the inside, 1 โˆ’ f, has derivative โˆ’1)

Set the sum equal to zero, which is where the growth curve is flat at its peak:

pB/(1 + Bf) โˆ’ q/(1 โˆ’ f) = 0

Step 7: Solve for f

Move the second term across and cross-multiply:

pB(1 โˆ’ f) = q(1 + Bf)

pB โˆ’ pBf = q + qBf

pB โˆ’ q = pBf + qBf

pB โˆ’ q = Bf(p + q)

But donโ€™t turn your pattern recognition noggin off!

Since p + q = 1, this collapses to:

f = (pB โˆ’ q) / B

f = p โˆ’ q / B

Can you reproduce this right now on a blank sheet of paper?* Probably not. But go through it once the way you used to trace when you learned the motions for drawing comic book characters or flowers or in my case TMNT. After a single reproduction by hand, I assure you some neural pathways will re-open that have had construction signs in front of them for years.

And to tie this back to the beginning, that is the joy of thickness.

[Get your mind back over here, this is a family letter.]

*You will be able to reproduce it on your own after 2 or 3 attempts. I donโ€™t know why the motion of the pencil is a 10x better instructor than reading it a bunch of times (generation effect maybe?). But this is one of those learning principles I take seriously and impress on my kids.

The world will seduce you with ease where youโ€™d be better served by friction. It is a 21st-century skill to have a point of view on the difference.

Kelly Criterion โ€” Cheatsheet Derivation

Kelly Criterion โ€” Cheatsheet Derivation


Step 1 One Period Expectancy

E = pB โˆ’ q

where p = win probability, q = 1โˆ’p = loss probability, B = net odds (win B per unit staked, lose 1).

Why: Weighted average of outcomes. Win B with probability p, lose 1 with probability q.


Step 2 Per-Flip Wealth Multipliers

Bet fraction f of current wealth W:

  • Win: Wโ‚ = W(1+Bf)
  • Lose: Wโ‚ = W(1โˆ’f)

Why: You keep the unbet portion W(1โˆ’f) regardless. On a win you collect B times your stake Wf on top. On a loss your stake Wf is gone.


Step 3 Wealth After n Flips

After h wins and (nโˆ’h) losses:

Wโ‚™ = Wโ‚€(1+Bf)h(1โˆ’f)nโˆ’h

Why: The flips compound โ€” each one rescales whatever the previous left. That means multiply, not add. Order doesn’t matter, only h and nโˆ’h.


Step 4 Per-Flip Growth Rate G

Take the nth root of total growth to extract the per-period rate:

G = (Wโ‚™/Wโ‚€)1/n = (1+Bf)h/n ยท (1โˆ’f)(nโˆ’h)/n

As n โ†’ โˆž, law of large numbers: h/n โ†’ p, (nโˆ’h)/n โ†’ q.

G = (1+Bf)p(1โˆ’f)q

Why nth root: Same logic as extracting r from (1+r)n = total growth. Geometric mean, not arithmetic, because the process is multiplicative.


Step 5 Take ln Before Differentiating

Define g = ln(G). Since ln is monotonically increasing, maximizing g gives the same f* as maximizing G.

Apply two log rules:

  • ln(AB) = ln(A) + ln(B)  โ†’  product becomes sum
  • ln(Ap) = pยทln(A)  โ†’  exponent drops to coefficient
g = pยทln(1+Bf) + qยทln(1โˆ’f)

Why: Differentiating a product of powers is a mess. A sum of logs is trivial. Valid because ln is monotone โ€” same maximum, easier math.


Step 6 Differentiate and Set to Zero

Rule: ddx[ln(x)] = 1x. Chain rule: multiply by derivative of the inside.

  • ddf[pยทln(1+Bf)] = pB1+Bf  โ† chain rule gives B from inside (1+Bf)
  • ddf[qยทln(1โˆ’f)] = โˆ’q1โˆ’f  โ† chain rule gives โˆ’1 from inside (1โˆ’f)

Set dg/df = 0:

pB1+Bf โˆ’ q1โˆ’f = 0

Step 7 Solve for f*

Cross-multiply:

pB(1โˆ’f) = q(1+Bf)

Expand:

pB โˆ’ pBf = q + qBf

Collect f terms:

pB โˆ’ q = pBf + qBf = Bf(p+q)

Since p+q = 1:

f* = pBโˆ’qB = p โˆ’ qB

The Answer

f* = pB โˆ’ qB

Read as: edge / odds

  • Numerator pBโˆ’q is your expected profit per unit bet
  • Denominator B scales it by the odds

Special case B=1 (even money): f* = pโˆ’q

Your optimal bet equals your raw edge.


Key Insights

What Why it matters
Multiplicative wealth function One bad bet can’t be offset by other bets โ€” sizing matters
Geometric mean not arithmetic Compounding processes need per-period rates, not averages
ln transform Turns product into sum without moving the maximum
Chain rule on ln(1โˆ’f) The โˆ’1 derivative is what creates a finite optimum
p+q=1 The simplification that closes the algebra cleanly

“end behavior”

Alex is an options trader you should follow in case he ever tweets a lot. Because he doesnโ€™t, when he posted the question below a year ago, it got few responses. I took the liberty of posting it myself this week.

Article content

This was fun because it led to a lot of discussion on the timeline and DMs. I was told it sparked a bunch of quant debate on one traderโ€™s desk.

The most popular answer, which was still less than 1/3 of the responses, was the correct answer.

Why?

The maximum value of a put is the strike. The maximum value of a call is the stock price.

Straddle is C + P so $100+$100 = $200

Notice how this means all call spreads go to zero since the calls are worth the same โ€” the stock price. All put spreads go to their max valueโ€” the distance between strikes because the puts themselves are worth the strikes.

Logic for delta:

Delta is the change in option price per change in stock.

But the putโ€™s strike is fixed, so the value of the put doesnโ€™t depend on the stock price. The put has zero delta. Itโ€™s always worth $100. Which means it has no gamma either ๐Ÿ™‚

The call is $100 because the max value of the call is the stock price. The call value moves 1-to-1 with the stock, so it has a delta of 1 or 100%

The max value of a straddle is therefore the stock price plus the strike price.

If you sell the straddle or either option at max value and hedge on its delta one time (this is known as a static hedge in contrast to dynamic hedging where you would rebalance as your hedge ratio changes), you cannot lose. It is that simple fact of arbitrage that makes it the upper bound.

To address the second most popular response in the poll, those who said the straddle is $100 (wrong) and has a 1.00 delta (correct), we will demonstrate why this is incorrect.

Whatโ€™s your p/l if you sell 1 straddle at $100 and buy 100 shares against, if the stock goes to $300?

The straddle will be worth $400, so you lose $300 but make $200 on your long share.

Hmm, maybe Iโ€™m just underhedged. Fine, what if I hedge on a 200 delta?

In that case, you actually make money; you win $400 on your 2 shares more than offsetting the $300 straddle loss. But what if the stock went to zero?

Your straddle p/l is unchanged, but you lost $200 on the long stock position. Arbitrage max value means you cannot lose if you sell at that price. Since we found a losing scenario, the price is not the maximum arbitrage bound. If you sell the straddle at $200 and buy a single share of stock, thereโ€™s no scenario where you lose. It is the lowest straddle value for which this no-lose scenario is true, thus itโ€™s the arbitrage bound.

Of course, this is but a toy problem where the call and put go to their maximum values because itโ€™s a degenerate case of infinite time or vol. But learning how a function (an option price is just a function) behaves by observing its boundaries is good for intuition. You did this in 9th grade. Khan Academy can jog your memory:

Article content

In the real world, you can fleetingly find options that trade beyond their arbitrage values:

Article content

Earlier in the week, @DeepDishEnjoyer aka p4 wrote a thread about a dividend mispricing.

It led to some back and forth with passersbys who use options but appear to have large gaps in the fundamentals.

Between the maximum value poll and p4โ€™s dividend lesson, itโ€™s worth saying it:

In a proper option education, you spend a lot of time on arbitrage relationships, cost of carry, and synthetics before you ever hear the word โ€œvolatilityโ€.

I didnโ€™t study formal math but I imagine thereโ€™s a lot in common with the process of proofs. Arbitrages rest heavily on assumptions. So to understand the relationships, you are forced into an intimate familiarity with the assumptions. And in the extremes of everything, itโ€™s the failure to examine assumptions that leads to being blindsided. But also, when things get extreme, to go on the attack means asking yourself, โ€œWhoโ€™s on autopilot? Is this price resting on a stale assumption?โ€ The arbitrage relationships give you the highest conceptual ROI that derivatives offer, you never learn the most useful thing derivatives can teachโ€ฆpassage over the โ€œbridge of assesโ€.

If you want to see more examples of why option basics are so key to understanding assumptions and opportunities when things get weird:

blindsided

My cousin Nicole just had her first child in the past year and is now compiling resources for new parents in light of where her attention has obviously been.

In our family chat, she asked:

When you became a new parent, what blindsided you the most?

Iโ€™ll share my answer, which I qualify with both the awareness that having a child is a gamble on many levels and that conception itself is a miracle and should never be taken for granted. What was I blindsided by?

[9:16 AM, 9/16/2026] Kris Abdelmessih: that they were gonna be so awesome

[9:17 AM, 9/16/2026] Kris Abdelmessih: that last one is important when we live in times where people have less kids and talk about it as though itโ€™s a chore (it is) but not the amazing upside

The most transcendent single moment of my life thus far was to hear my sonโ€™s voice the day he came into the world. I donโ€™t know if thatโ€™s every parentโ€™s experience, but immediately I felt the joy and clarity of purpose. Before that cry, I donโ€™t think I would have said I had no purpose, but the moment revealed that I didnโ€™t believe I did. The speed and intensity of this rush of belief was a novel feeling. Thus, irreparably blindsided.

Share your own answers in the comments and Iโ€™ll share them with Nicole. Thank you!


I offered a couple of less serious answers to her question as well.

Article content

On that last one, I wrote about that 2 years ago in our minds love to betray us:

Article content

When I was at the Sphere with my family over Spring Break, I wouldnโ€™t ride the long, exposed escalators. I took the elevator where I found the rest of the scaredy-cats.

If thereโ€™s anything good about having a phobia, itโ€™s empathy for the range of what can go on in peopleโ€™s minds and bodies.

Iโ€™m watching this poor guy thinking, donโ€™t do it man, itโ€™s not worth and all he wants is a glimpse:

Man crawls through his phobia of heights to get a view of the Atlantic Ocean from the edge of a cliff. ๐Ÿ˜…

Article content

1:44 PM ยท Sep 12, 2026 ยท 36.9M Views


2.99K Replies ยท 5.3K Reposts ยท 131K Likes

The comment section understands, and based on the number of โ€œlikesโ€, many others do too.

Article content

Anyway, I blame my kids for my embarrassment when I go to the Sphere to see Metallica with a group of guys next month and have to explain that Iโ€™ll meet them at the seats.

Moontower #328

In this issue:

  • blindsided
  • โ€œend behaviorโ€

Friends,

Blindsided

My cousin Nicole just had her first child in the past year and is now compiling resources for new parents in light of where her attention has obviously been.

In our family chat, she asked:

When you became a new parent, what blindsided you the most?

Iโ€™ll share my answer, which I qualify with both the awareness that having a child is a gamble on many levels and that conception itself is a miracle and should never be taken for granted. What was I blindsided by?

[9:16 AM, 9/16/2026] Kris Abdelmessih: that they were gonna be so awesome

[9:17 AM, 9/16/2026] Kris Abdelmessih: that last one is important when we live in times where people have less kids and talk about it as though itโ€™s a chore (it is) but not the amazing upside

The most transcendent single moment of my life thus far was to hear my sonโ€™s voice the day he came into the world. I donโ€™t know if thatโ€™s every parentโ€™s experience, but immediately I felt the joy and clarity of purpose. Before that cry, I donโ€™t think I would have said I had no purpose, but the moment revealed that I didnโ€™t believe I did. The speed and intensity of this rush of belief was a novel feeling. Thus, irreparably blindsided.

Share your own answers in the comments and Iโ€™ll share them with Nicole. Thank you!


I offered a couple of less serious answers to her question as well.

Article content

On that last one, I wrote about that 2 years ago in our minds love to betray us:

Article content

When I was at the Sphere with my family over Spring Break, I wouldnโ€™t ride the long, exposed escalators. I took the elevator where I found the rest of the scaredy-cats.

If thereโ€™s anything good about having a phobia, itโ€™s empathy for the range of what can go on in peopleโ€™s minds and bodies.

Iโ€™m watching this poor guy thinking, donโ€™t do it man, itโ€™s not worth and all he wants is a glimpse:

Man crawls through his phobia of heights to get a view of the Atlantic Ocean from the edge of a cliff. ๐Ÿ˜…

Article content

1:44 PM ยท Sep 12, 2026 ยท 36.9M Views


2.99K Replies ยท 5.3K Reposts ยท 131K Likes

The comment section understands, and based on the number of โ€œlikesโ€, many others do too.

Article content

Anyway, I blame my kids for my embarrassment when I go to the Sphere to see Metallica with a group of guys next month and have to explain that Iโ€™ll meet them at the seats.


Money Angle

A couple of Option Trench episodes to share:

๐Ÿ“บTerminal vs Path-Dependent Value Explained Using Collars | 39 min

๐Ÿ“บThe (Not So) Efficient Market Hypothesis? | 59 min

The first one will be useful for anyone wanting to learn more about option collars, which Iโ€™ve been writing a lot about. A video might be a gentler format so check that out.

The second one applies to investing broadly. I also use the Paradox of Provable Alpha at the end to answer a good question Erik asks.


Money Angle For Masochists

Alex is an options trader you should follow in case he ever tweets a lot. Because he doesnโ€™t, when he posted the question below a year ago, it got few responses. I took the liberty of posting it myself this week.

Article content

This was fun because it led to a lot of discussion on the timeline and DMs. I was told it sparked a bunch of quant debate on one traderโ€™s desk.

The most popular answer, which was still less than 1/3 of the responses, was the correct answer.

Why?

The maximum value of a put is the strike. The maximum value of a call is the stock price.

Straddle is C + P so $100+$100 = $200

Notice how this means all call spreads go to zero since the calls are worth the same โ€” the stock price. All put spreads go to their max valueโ€” the distance between strikes because the puts themselves are worth the strikes.

Logic for delta:

Delta is the change in option price per change in stock.

But the putโ€™s strike is fixed, so the value of the put doesnโ€™t depend on the stock price. The put has zero delta. Itโ€™s always worth $100. Which means it has no gamma either ๐Ÿ™‚

The call is $100 because the max value of the call is the stock price. The call value moves 1-to-1 with the stock, so it has a delta of 1 or 100%

The max value of a straddle is therefore the stock price plus the strike price.

If you sell the straddle or either option at max value and hedge on its delta one time (this is known as a static hedge in contrast to dynamic hedging where you would rebalance as your hedge ratio changes), you cannot lose. It is that simple fact of arbitrage that makes it the upper bound.

To address the second most popular response in the poll, those who said the straddle is $100 (wrong) and has a 1.00 delta (correct), we will demonstrate why this is incorrect.

Whatโ€™s your p/l if you sell 1 straddle at $100 and buy 100 shares against, if the stock goes to $300?

The straddle will be worth $400, so you lose $300 but make $200 on your long share.

Hmm, maybe Iโ€™m just underhedged. Fine, what if I hedge on a 200 delta?

In that case, you actually make money; you win $400 on your 2 shares more than offsetting the $300 straddle loss. But what if the stock went to zero?

Your straddle p/l is unchanged, but you lost $200 on the long stock position. Arbitrage max value means you cannot lose if you sell at that price. Since we found a losing scenario, the price is not the maximum arbitrage bound. If you sell the straddle at $200 and buy a single share of stock, thereโ€™s no scenario where you lose. It is the lowest straddle value for which this no-lose scenario is true, thus itโ€™s the arbitrage bound.

Of course, this is but a toy problem where the call and put go to their maximum values because itโ€™s a degenerate case of infinite time or vol. But learning how a function (an option price is just a function) behaves by observing its boundaries is good for intuition. You did this in 9th grade. Khan Academy can jog your memory:

Article content

In the real world, you can fleetingly find options that trade beyond their arbitrage values:

Article content

Earlier in the week, @DeepDishEnjoyer aka p4 wrote a thread about a dividend mispricing.

It led to some back and forth with passersbys who use options but appear to have large gaps in the fundamentals.

Between the maximum value poll and p4โ€™s dividend lesson, itโ€™s worth saying it:

In a proper option education, you spend a lot of time on arbitrage relationships, cost of carry, and synthetics before you ever hear the word “volatility”.

I didnโ€™t study formal math but I imagine thereโ€™s a lot in common with the process of proofs. Arbitrages rest heavily on assumptions. So to understand the relationships, you are forced into an intimate familiarity with the assumptions. And in the extremes of everything, itโ€™s the failure to examine assumptions that leads to being blindsided. But also, when things get extreme, to go on the attack means asking yourself, โ€œWhoโ€™s on autopilot? Is this price resting on a stale assumption?โ€ The arbitrage relationships give you the highest conceptual ROI that derivatives offer, you never learn the most useful thing derivatives can teachโ€ฆpassage over the “bridge of asses”.

If you want to see more examples of why option basics are so key to understanding assumptions and opportunities when things get weird:

From My Actual Life

I leave you with another pic from our family chat where my wife posted a photo of where she was walking.

Article content

I donโ€™t know how many Egyptian Arabic speakers we got in the crowd but โ€œshib-shibโ€ is like a slipper. Momโ€™s weapon of choice. Apparently this is a broader thing:

Mothers and shoes

Stay groovy

โ˜ฎ๏ธ


Moontower Weekly Recap

Variance & Covariance Cheat Sheet

Variance & Covariance Cheat Sheet

Starting points (where every derivation begins)

Everything below is derived from these definitions. They’re the raw material โ€” average squared deviation for variance, average product of deviations for covariance. When a derivation feels stuck, come back here and plug in.

Variance โ€” average squared deviation from the mean
Var(X) = E[(X โˆ’ ฮผ)2]    ฮผ = E[X]
Covariance โ€” average product of deviations
Cov(X, Y) = E[(X โˆ’ ฮผX)(Y โˆ’ ฮผY)]
Sum of squared deviations (the un-averaged version)
SS = ฮฃ (xi โˆ’ ฮผ)2    Var = SSn

Variance is just SS divided by n (or nโˆ’1 for a sample). Same object, before you average.

The move in every derivation: plug into one of these, expand the square or product (pure algebra), apply E using linearity, then recognize the Var/Cov patterns that fall out. The computational forms below (E[X2] โˆ’ (E[X])2, E[XY] โˆ’ E[X]E[Y]) are results of doing this, not starting points.


The identities

Variance from the definition
Var(X) = E[(X โˆ’ ฮผ)2] = E[X2] โˆ’ (E[X])2

Average of the squares minus the square of the average. Worth showing where that second form comes from, since every later grind reuses this exact collapse. Start from the deviation definition and expand the square:

1n ฮฃ(xi โˆ’ x)2 = 1n ฮฃ(xi2 โˆ’ 2xix + x2)

Average term by term. The key is that x is a constant (already computed), so it pulls out of the sums:

  • First term: (1/n)ฮฃxi2 = E[X2]
  • Middle term: (1/n)ฮฃ(โˆ’2xix) = โˆ’2x ยท (1/n)ฮฃxi = โˆ’2x ยท x = โˆ’2(E[X])2
  • Last term: (1/n)ฮฃx2 = x2 = (E[X])2 (averaging a constant returns the constant)

Put them together โ€” and notice the last term carries a coefficient of 1, not 2:

E[X2] โˆ’ 2(E[X])2 + (E[X])2 = E[X2] โˆ’ (E[X])2

The โˆ’2 and +1 combine to โˆ’1. That collapse โ€” middle and last terms both becoming (E[X])2 and partially cancelling โ€” is the same move behind every Var/Cov identity on this sheet.

Covariance from the definition
Cov(X, Y) = E[(X โˆ’ ฮผX)(Y โˆ’ ฮผY)] = E[XY] โˆ’ E[X]E[Y]

Average of the products minus the product of the averages.

Variance is covariance with itself
Cov(X, X) = Var(X)
Scaling rule for variance
Var(aX) = a2 ยท Var(X)

Constants pull out as their square.

Scaling rule for covariance
Cov(aX, bY) = ab ยท Cov(X, Y)
Shifting rule for covariance
Cov(X + c, Y) = Cov(X, Y)

Adding a constant doesn’t change covariance.

Covariance with a constant is zero
Cov(X, c) = 0

Constants don’t co-vary.

Variance of a sum
Var(X + Y) = Var(X) + Var(Y) + 2 ยท Cov(X, Y)
Variance of a weighted sum (the workhorse)
Var(aX + bY) = a2 Var(X) + b2 Var(Y) + 2ab ยท Cov(X, Y)

Here a and b are the amounts held of each asset. They’re portfolio weights when they sum to 1. Var(X+Y) above is just this formula with a = b = 1 โ€” one unit of each, no weighting lever. The weights are what turn a raw sum into a portfolio.

Worked example. Two assets: ฯƒX = 20%, ฯƒY = 10%, ฯ = 0.3. Equal weights a = b = 0.5.
  • Var(X) = 0.04,   Var(Y) = 0.01
  • Cov(X, Y) = ฯ ยท ฯƒX ยท ฯƒY = 0.3 ยท 0.2 ยท 0.1 = 0.006
  • Var(P) = 0.25ยท0.04 + 0.25ยท0.01 + 2ยท0.5ยท0.5ยท0.006 = 0.01 + 0.0025 + 0.003 = 0.0155
  • ฯƒP = โˆš0.0155 โ‰ˆ 12.4%
Compare to the naive weighted-average vol, 0.5ยท20% + 0.5ยท10% = 15%. The cross term (with ฯ < 1) is what pulls portfolio vol below the average of the two vols. That gap is the diversification benefit.
Variance of a difference (spread variance / pair-trading formula)
Var(X โˆ’ Y) = Var(X) + Var(Y) โˆ’ 2 ยท Cov(X, Y)

Same as Var(X+Y) but cross term flips sign. When X and Y are highly correlated, spread variance is small โ€” the math behind why pair trades work.

Bilinearity of covariance
Cov(X, A + B) = Cov(X, A) + Cov(X, B)

Same in the first slot by symmetry.

Where it’s used. This is the move that lets you compute an asset’s covariance with a whole portfolio without re-deriving anything. Say a portfolio P = 0.5A + 0.5B and you want how asset A co-moves with the portfolio it sits in:
Cov(A, P) = Cov(A, 0.5A + 0.5B) = 0.5 Var(A) + 0.5 Cov(A, B)
Distribute across the sum, pull the weights out. That number โ€” an asset’s covariance with its own portfolio โ€” is its marginal contribution to portfolio risk, and it’s exactly what you FOIL out when you expand Var(w1X1 + โ€ฆ + wnXn) into the full covariance matrix. Bilinearity is the engine under every portfolio-variance calculation.
Variance of a binomial
Var(H) = np(1โˆ’p)   where H = ฮฃ Xi

Derived in two steps, both from scratch.

Step 1 โ€” variance of a single flip. One flip X is 1 with probability p, 0 with probability (1โˆ’p). Mean is E[X] = p. Plug into the squared-deviation definition โ€” deviations are (1โˆ’p) for heads and (0โˆ’p) = โˆ’p for tails, each weighted by its probability:

Var(X) = p(1โˆ’p)2 + (1โˆ’p)p2

Factor out p(1โˆ’p): the bracket is (1โˆ’p) + p = 1, so

Var(X) = p(1โˆ’p)

(Peaks at p = 0.5, value 0.25 โ€” the fair coin is the most uncertain, most variance per flip.)

Step 2 โ€” n flips. Write H as a sum of n independent single flips, H = X1 + โ€ฆ + Xn. Variance of a sum adds the pairwise Cov terms, but independent flips have Cov(Xi, Xj) = 0, so every cross term drops. The n identical variances just add:

Var(H) = ฮฃ Var(Xi) = n ยท p(1โˆ’p)

The np(1โˆ’p) isn’t handed to you โ€” it falls out of one Bernoulli’s p(1โˆ’p) times n, because independence kills the covariances.

Standard deviation scaling
StDev(aX) = |a| ยท StDev(X)
Correlation definition
ฯ = Cov(X, Y)ฯƒX ยท ฯƒY

When ฯ = 1: Cov(X, Y)2 = Var(X) ยท Var(Y).


The derivation recipe

Every identity in this neighborhood comes out of the same five moves. When you see a Var or Cov of something built from linear combinations of random variables, this is the procedure.

  1. Plug into the definition. Use Var(Z) = E[Z2] โˆ’ (E[Z])2 or Cov(X, Y) = E[XY] โˆ’ E[X]E[Y] depending on what you’re computing.
  2. Expand squares and products. Pure algebra on the random variables. FOIL out any binomials. No expectations yet.
  3. Apply E using linearity. Distribute E across sums, pull constants out of expectations. This is the step that does the most work. Always handle linearity first when you have the chance โ€” squaring is not linear, so you simplify E first and let the square wrap what’s left.
  4. Group matching terms. Line up the things that share factors (a2 terms together, ab terms together, b2 terms together, etc.).
  5. Factor and recognize. Pull out shared factors and spot the patterns: (E[X2] โˆ’ (E[X])2) is Var(X), and (E[XY] โˆ’ E[X]E[Y]) is Cov(X, Y).

The reason this recipe always closes: variances and covariances are quadratic in the underlying random variables, so expanding any square or product of linear combinations only generates more variances and covariances. Step 5 is recognition, not computation. There’s nowhere else for the algebra to land.


Two applications

Interview problem: E[H ยท T] for n coin flips

Flip a fair coin n = 100 times. H = heads, T = tails. Find E[H ยท T]. Worked slowly, because the one-line answer hides about six moves.

Step 1 โ€” first reach, and why it fails. The instinct is E[H ยท T] = E[H] ยท E[T] = 50 ยท 50 = 2,500. But splitting a product of expectations like that is only legal when the two variables are independent. Check the precondition: H + T = 100, so knowing H pins down T exactly. Not independent. The naive split is off by a correction.

Step 2 โ€” name the correction. That correction is what covariance is. Rearranging Cov(X, Y) = E[XY] โˆ’ E[X]E[Y]:

E[H ยท T] = E[H] ยท E[T] + Cov(H, T)
Independent โ†’ Cov = 0 โ†’ naive split exact. Locked โ†’ Cov โ‰  0 โ†’ you need the term.

Step 3 โ€” get the sign first. H + T = 100, so when H is above its mean, T is forced below. They move opposite, always. So Cov(H, T) is negative, and the true answer lands below 2,500.

Step 4 โ€” compute Cov(H, T) via substitution. Since T = 100 โˆ’ H, write Cov(H, T) = Cov(H, 100 โˆ’ H) and split with bilinearity:

Cov(H, 100 โˆ’ H) = Cov(H, 100) + Cov(H, โˆ’H)
First term is covariance with a constant โ†’ 0. Second term, pull out the โˆ’1 (scaling rule, ab = 1 ยท (โˆ’1) = โˆ’1):
= 0 โˆ’ Cov(H, H) = โˆ’Var(H)
So Cov(H, T) = โˆ’Var(H). Now it’s earned, not asserted.

Step 5 โ€” Var(H) is the binomial variance. H is the count of heads in n flips, so Var(H) = np(1โˆ’p) = 100 ยท 0.5 ยท 0.5 = 25.

Step 6 โ€” land it.

E[H ยท T] = 2,500 โˆ’ 25 = 2,475

Where n(nโˆ’1) comes from. Keep everything in symbols instead of plugging in. E[H] = np and E[T] = n(1โˆ’p), so E[H]ยทE[T] = n2p(1โˆ’p), and Var(H) = np(1โˆ’p). Then:

E[H ยท T] = n2p(1โˆ’p) โˆ’ np(1โˆ’p) = p(1โˆ’p)[n2 โˆ’ n] = n(nโˆ’1)p(1โˆ’p)
The n2 is the naive product, the โˆ’n is the variance shortfall, and factoring out p(1โˆ’p) leaves n(nโˆ’1). Sanity check: 100 ยท 99 ยท 0.25 = 2,475. โœ“

The through-line: the answer falls short of E[H]ยทE[T] by exactly Var(H), because Cov(H, T) = โˆ’Var(H) whenever H and T sum to a constant.

Two-stock equal-weight portfolio with equal variances ฯƒ2 and correlation ฯ

Start from the workhorse with a = b = 0.5:

Var(P) = 0.25 Var(X) + 0.25 Var(Y) + 2(0.5)(0.5) Cov(X, Y)

Impose equal variances Var(X) = Var(Y) = ฯƒ2, and write the cross term with correlation, Cov(X, Y) = ฯฯƒ2 (since Cov = ฯ ยท ฯƒX ยท ฯƒY and both vols are ฯƒ):

Var(P) = 0.25ฯƒ2 + 0.25ฯƒ2 + 0.5ฯฯƒ2 = 0.5ฯƒ2 + 0.5ฯฯƒ2

Factor out 0.5ฯƒ2:

Var(P) =  ฯƒ2(1 + ฯ)2
Why this is the instructive form. Weights are fixed (50/50) and both vols are fixed (ฯƒ). The only thing left moving is ฯ. So the entire diversification effect is carried by the single factor (1 + ฯ)/2 โ€” a clean dial from 0 to 1 that multiplies the single-name variance. Sweep ฯ and read what correlation actually does:
ฯ Var(P) ฯƒP (vol) vs. one stock
+1ฯƒ2ฯƒno benefit โ€” identical names
+0.50.75ฯƒ20.87ฯƒ13% vol cut
00.5ฯƒ20.71ฯƒ29% vol cut (the โˆšยฝ case)
โˆ’0.50.25ฯƒ20.5ฯƒhalf the vol
โˆ’100risk fully cancels
The variance scales linearly in ฯ, but the thing you feel โ€” vol, ฯƒP = ฯƒโˆš((1+ฯ)/2) โ€” scales as the square root, so the first chunk of decorrelation buys more than the last. Going from ฯ = 1 to ฯ = 0.5 already takes 13% off your vol. You do not need negative correlation to diversify; anything below +1 helps. Negative correlation is just the strong form, and ฯ = โˆ’1 is the perfect hedge where the two positions cancel outright.

The whole two-name diversification story lives in that (1 + ฯ)/2 factor. Same vols, same weights, and correlation alone moves you from “no benefit” to “risk gone.”

Two-stock unequal-weight portfolio with equal variances ฯƒ2 and correlation ฯ
Var(P) = ฯƒ2 [1 โˆ’ 2w(1โˆ’w)(1โˆ’ฯ)]

Diversification benefit is the product of a weight piece (2w(1โˆ’w), maxed at w = 0.5) and a correlation piece (1โˆ’ฯ). Need both to get benefit. With equal variances, equal weighting is optimal โ€” any tilt from 50/50 sacrifices diversification.

Two-stock unequal-variance portfolio โ†’ inverse-variance weighting

Now drop the equal-variance assumption. Keep ฯƒX2 and ฯƒY2 separate. Weights w and (1โˆ’w):

Var(P) = w2ฯƒX2 + (1โˆ’w)2ฯƒY2 + 2w(1โˆ’w)ฯฯƒXฯƒY

Minimize over w. Var(P) is an upward parabola in w (positive coefficient on w2), so the critical point is the min. Differentiate term by term and set to zero:

2wฯƒX2 โˆ’ 2(1โˆ’w)ฯƒY2 + 2(1โˆ’2w)ฯฯƒXฯƒY = 0

Divide by 2, expand, collect the w terms on the left and constants on the right, factor w out:

w(ฯƒX2 + ฯƒY2 โˆ’ 2ฯฯƒXฯƒY) = ฯƒY2 โˆ’ ฯฯƒXฯƒY
w* =  ฯƒY2 โˆ’ ฯฯƒXฯƒYฯƒX2 + ฯƒY2 โˆ’ 2ฯฯƒXฯƒY

The denominator is Var(X โˆ’ Y) โ€” the spread variance from earlier. The numerator is ฯƒY2 โˆ’ Cov(X, Y).

The payoff โ€” set ฯ = 0 (independent names):
w* = ฯƒY2ฯƒX2 + ฯƒY2 = 1/ฯƒX21/ฯƒX2 + 1/ฯƒY2
That’s inverse-variance weighting: each asset’s weight is its inverse variance over the sum of inverse variances. The quieter asset gets more money. This is the result behind Kalman filters, weighted least squares, and meta-analysis โ€” anywhere you optimally combine noisy estimates, you weight by precision (1/variance). It’s also the “optimal” cousin of the inverse-vol risk-parity heuristic, which ignores correlations.

Intuition โ€” hold ฯƒX fixed at 20% (ฯƒX2 = 0.04), turn the ฯƒY knob:

ฯƒY ฯƒY2 w* on X
000
10%0.010.20
20%0.040.50
40%0.160.80
โˆžโˆžโ†’ 1

X’s weight is driven by Y’s variance, not its own. The noisier the alternative, the more you pile into X. Three anchors: ฯƒY2 = 0 โ†’ w* = 0 (Y is riskless, hold only Y); ฯƒY2 = ฯƒX2 โ†’ w* = 0.5 (equal variances recover equal weighting); ฯƒY2 โ†’ โˆž โ†’ w* โ†’ 1 (Y is pure noise, flee into X). The cleanest limit: if ฯƒX2 = 0, then w* = 1 โ€” a riskless X takes the whole book. Precision is just quietness, and you trust the quiet estimate more.

Three-variable portfolio variance โ†’ why the matrix shows up

Same Form A grind, one more variable. Expand (aX + bY + cZ)2, apply E, subtract the squared-mean term. Every squared term becomes a variance, every cross term a covariance:

Var(aX + bY + cZ) = a2Var(X) + b2Var(Y) + c2Var(Z) + 2ab Cov(X,Y) + 2ac Cov(X,Z) + 2bc Cov(Y,Z)

Counting the terms. For n assets you always get:

  • n variance terms โ€” one per asset (the ai2 Var pieces)
  • nC2 = n(nโˆ’1)/2 covariance pairs โ€” one per distinct pair

Total = n + nC2. For n = 3: 3 + 3 = 6. For n = 100: 100 variances + 4,950 covariance pairs. The cross terms grow as n2, which is exactly why nobody writes portfolio variance longhand past n = 3 โ€” you switch to the matrix form.

The double-sum / matrix form. Organize every term into a grid indexed by asset pairs. With weights wi and returns ri:
Var(P) = ฮฃi ฮฃj wi wj Cov(ri, rj) = wโŠคฮฃw
Reading the double sum: the outer ฮฃ over i and the inner ฮฃ over j together form every ordered pair (i, j). For each pair you drop in one term, wiwjCov(ri, rj), and add them all up. For n = 3 that’s 3 ร— 3 = 9 cells:
X (j=1) Y (j=2) Z (j=3)
X (i=1)a2Var(X)ab Cov(X,Y)ac Cov(X,Z)
Y (i=2)ab Cov(X,Y)b2Var(Y)bc Cov(Y,Z)
Z (i=3)ac Cov(X,Z)bc Cov(Y,Z)c2Var(Z)
Sum all nine cells and you get the six-term formula above. Two things to see:
  • Diagonal (i = j, shaded): Cov(ri, ri) = Var(ri), so the diagonal is the n variance terms.
  • Off-diagonal (i โ‰  j): each unordered pair appears twice โ€” cell (X,Y) and cell (Y,X) are identical โ€” and those two copies are exactly where the factor of 2 on each covariance comes from. You never write the 2 by hand; the grid double-counts it for you.
So ฮฃ is the covariance matrix: variances down the diagonal, covariances off it. wโŠคฮฃw just says “sweep every cell of the grid, weight it, sum it.” The n + nC2 count is the matrix โ€” diagonal plus (doubled) upper triangle. You’ve already discovered why the matrix form is inevitable; it’s just bookkeeping for the term explosion.
Figure โ€” the double sum is the matrix: one sweep, three views
Var(P) = ฮฃi ฮฃj wiwj Cov(ri, rj) an instruction for sweeping a grid: for every row i, for every column j, add that cell inner ฮฃ over j → picks the column X (j=1) Y (j=2) Z (j=3) outer ฮฃ over i → picks the row X (i=1) Y (i=2) Z (i=3) w1² Var(X) w1w2 Cov(X,Y) w1w3 Cov(X,Z) w2w1 Cov(X,Y) w2² Var(Y) w2w3 Cov(Y,Z) w3w1 Cov(X,Z) w3w2 Cov(Y,Z) w3² Var(Z) Diagonal (i = j): Cov(ri, ri) = Var(ri) — the 3 variance terms. Twin cells (i,j) & (j,i): identical — every covariance is visited twice. That double-count is the ×2 you wrote by hand. The grid supplies it for free. add all 9 cells = wT ฮฃ w ฮฃ is the covariance matrix: variances on the diagonal, covariances off it. n assets = the same sweep on an n×n grid. Nothing new happens; the grid just grows.
Figure โ€” three forms of the same variance, and where the 2s go
โ‘  ALGEBRAIC — you write the 2s by hand a²Var(X) + b²Var(Y) + c²Var(Z) 2ab Cov(X,Y) + 2ac Cov(X,Z) + 2bc Cov(Y,Z) โ‘ก MATRIX — the 2s vanish into the symmetry Var(P) = wT ฮฃ w,   where ฮฃ = Var(X) Cov(X,Y) Cov(X,Z) Cov(X,Y) Var(Y) Cov(Y,Z) Cov(X,Z) Cov(Y,Z) Var(Z) Each covariance sits in two matched-color cells (mirrored across the diagonal). Summed, the two cells are the 2 from Stage 1. Variances (diagonal, purple) have no mirror — that’s why they’re never doubled. โ‘ข DOUBLE SUM — the matrix written in math Var(P) = ฮฃ n i=1 ฮฃ n j=1 wiwj Cov(ri, rj) Cov(ri, ri) = Var(ri) — covariance with itself is just variance, so the diagonal needs no special case. Both sums run 1 to n, so the sweep is n² terms. 10 assets → 100 terms, not 20. That quadratic blow-up is exactly why the compact matrix form earns its keep.
Notation you’ll see in practice: wโŠคฮฃw. You’ll run into this constantly in risk models, optimizers, and quant papers. It’s the same portfolio variance, packaged as a matrix operation. Reading it piece by piece:
  • w โ€” the weight vector, weights stacked in a column.
  • wโŠค โ€” “w transpose,” the same weights laid flat as a row. Transpose just tips a column over into a row.
  • ฮฃ โ€” the covariance matrix (the grid above). Watch out: this capital-sigma is the matrix, not a summation sign. Variances on the diagonal, covariances off it.
So wโŠคฮฃw is row-of-weights ร— matrix ร— column-of-weights, which multiplies out to a single number โ€” the portfolio variance. It’s identical to the double sum, and in a spreadsheet it’s literally =MMULT(MMULT(TRANSPOSE(w), ฮฃ), w).

Why bother, when the double sum already shows everything? Three practical reasons, none of them “it’s more correct.” It doesn’t grow โ€” three symbols whether n is 2 or 2,000, where the double sum for 500 names is 250,000 terms. It’s how software actually computes portfolio variance (one fast matrix op). And optimization only speaks matrix: the minimum-variance weights come out as ฮฃโˆ’11 normalized, and the inverse ฮฃโˆ’1 has no double-sum spelling. The two-asset inverse-variance weighting derived above is ฮฃโˆ’1 for n = 2 โ€” the matrix form is how that generalizes. For understanding, the double sum is enough; this is the notation for doing things with it.


The ladder

Each rung is built from the one before it โ€” the definition first, then the algebra of scaling and adding, then portfolios, then the jump to the matrix. Nothing is assumed that wasn’t derived earlier.

1.  Variance as average squared deviation
2.  Var(X) = E[X2] โˆ’ (E[X])2 โ€” from the definition
3.  Cov(X, Y) = E[XY] โˆ’ E[X]E[Y] โ€” from the definition
4.  Var(aX) = a2ยทVar(X)
5.  Cov(aX, bY) = abยทCov(X, Y)
6.  Cov(X + c, Y) = Cov(X, Y) and Cov(X, c) = 0
7.  Var(X + Y) = Var(X) + Var(Y) + 2ยทCov(X, Y)
8.  Var(aX + bY) โ€” the full weighted-sum workhorse
9.  Two-stock equal-weight portfolio variance โ†’ the (1+ฯ)/2 diversification factor
10. Two-stock unequal-weight portfolio variance (equal variances)
11. Var(X โˆ’ Y) โ€” spread variance / pair-trading formula
12. Variance of a binomial = np(1โˆ’p)
13. Two-stock unequal-variance portfolio โ†’ inverse-variance weighting
14. Var(aX + bY + cZ) โ€” three variables, and the n + nC2 term count
15. General n-asset portfolio variance โ€” the wโŠคฮฃw matrix form

Is this one lesson in a math course?

No. This would be roughly half a semester of a first probability course, or a full chapter and a half of a more applied stats book.

Rough mapping to a standard curriculum:

  • Variance from the definition, E[X2] identity: one lecture, plus a problem set
  • Covariance and the product identity: one lecture
  • Scaling rules, bilinearity, variance of a sum: one to two lectures
  • Portfolio variance, weighted sums, correlation: one lecture in the probability course, or the opening week of a finance/portfolio-theory course
  • Binomial variance, applications: another lecture or two

So this is the equivalent of maybe four to six lectures of material, plus the problem sets that go with them. The reason it feels like a lot is that this sheet does the whole pipeline โ€” derivation, intuition, numerical examples, applications โ€” for each piece, instead of showing a formula and moving on.

The trade-off is real: this is slower but produces durable understanding. A typical math course shows you Var(aX + bY) on day one, leaves the “why it’s that and not something else” fuzzy, and you pattern-match for the rest of the semester. Done this way, when ฯƒ2 ยท (1+ฯ)/2 turns up in a textbook two years later, you see the bilinearity FOIL behind it instead of recognizing a memorized formula.

That’s the trade. Slower, but it sticks.

why i’m not more bearish equities compared to 6 months ago

Programming Note: I normally publish Munchies on Wednesday and the paid post on Thursday, but this week I will post both a day early, as they are market-related and I want to get them out ahead of the Fed meeting.


Friends,

Iโ€™ll start with updating some broad market observations I laid out in March and then square that with what option surfaces are telling us in the context of the Fed meeting and beyond. The Fed meeting is a highly skewed event with a 25 bp hike more than about 90% priced in.

The probability of Fed target rate between 3.75% and 4% has shot from 35% to 90% in under 3 weeks

Recent context:

The Jackson Hole keynote was on 8/28/26. The probability of a rate hike shot from 35% to 57% in one day.

The 10-year yield has rallied from 4.67% on 8/26 to 4.96% as I write on 9/14.

In that same window, the SPX is down a mere 50 bps and the NDX is basically unched.

Letโ€™s rewind to my March post trading is like a sudoku puzzle with prices as the given numbers. I compared earnings yields to bond yields to get an equity risk premium. This comparison in the modern era of massive budget deficits, where a large โ€œGโ€ in the Kalecki-Levy world tells us that money ends up as private sector nominal income by identity*, never flatters equities. We can reconcile skinny equity risk premiums the same we reconcile low earnings yield any stockโ€ฆthereโ€™s an expectation of growth.

*While this is reductionist in the sense that the flow doesnโ€™t definitionally have to end up there, it also happens to be where it has ended up.

If we are desensitized to the low absolute level of equity risk premiums, it is because it has been easy to presume nominal growth. We have relatively low unemployment, technology companies growing quickly despite massive scale, and, I almost forgot, A GOVERNMENT THAT IS EXISTENTIALLY POT-COMMITTED TO DEFICIT SPENDING.

โ€œBut, we have a Republican president.โ€

Ok, define Republican. Because if you look closely at our fiscalโ€ฆoh never mind. Nobody cares. Weโ€™re well past fiscal discipline as a political delimiter.

The more money there is, the more surface area there is for theft/grift/graft.

[While this is now quite obvious and exploited by both right and left, Donโ€™s brazenness feels like itโ€™s giving everyone permission. He is a populist wrestling everyoneโ€™s birthright to cheat away from the cloaked political elites who once monopolized it. That sentence is the dressโ€ฆwhether you see blue or gold is up to you.]

Alrighty then, where were we? Ahh, yes, nominal growth as a given, because our society is a passenger in a global economic trolley problem. Fantastic. Instead of fretting over the level of the equity risk premium, we can look at the past 6 months to consider the change or, in a sudoku-esque way, solve for what needs to happen for the risk premium to not get worse.

Thus far, higher energy prices and higher bond yields havenโ€™t produced the equity decline in the scenario I outlined. My hunch is the market got more expensive on a relative basis, but letโ€™s investigate.

First, the updated numbers.

Market quotes below are September 14 intraday observations, taken at different times. The equity calculations use the same snapshot as the charts.

WTI oil is up 18.5% since the March post (March price from 3/27/26)

RBOB gasoline is up ~28%

IEF on a div-adjusted basis is down 1.9%, about half what the duration (which is a snapshot like delta) expects because you picked up over 200 bps of carried interest (ie those divs) in the meantime.

Equity valuation as the missing Sudoku number

In late March, I used an equity earnings yield of roughly 5% vs 4.4% ten-year Treasury yield. This represented a 60 bps nominal equity risk premium and ~300 bps in real terms. My downside scenario considers what might be expected if the Treasury yield reached 5.5% and equities needed to offer another 50 bps above that. The target nominal earnings yield would rise to 6%, requiring roughly 17% equity decline if earnings stayed flat.

On partial probabilities

My thinking was on a subset of causes for yields to rise. I was thinking about the inflation channel via energy price pass-through in the event that futures prices rolled up to war-bolstered petroleum spot prices. This is supply-shock inflation, but there is also demand-pull inflation. Yields can also increase because of concerns about sovereign creditworthiness. Asset prices are complex because they embed expectations about many variables, some of which reinforce each other and some offset. Arrows in every direction. Complex portfolios can be constructed to isolate bets on conditional or partial probabilities.

An example from sports betting:

Say you bet $900 to win $100 on the Seahawks not winning the Super Bowl, and $100 to win $600 on them winning the NFC. If they donโ€™t win the NFC, the bets cancel. If they reach the Super Bowl, you make $700 if they lose and lose $300 if they win. Youโ€™ve constructed a conditional bet: Seattle to lose the Super Bowl, provided they get there.

The prices imply a 10% chance of winning the Super Bowl and a 14.3% chance of getting there. Divide those and you get a 70% chance of winning conditional on getting there. Your combined position bets against that 70%.

The way Iโ€™m reasoning in a macro way is implicitly partial. All relative value bets have this property, but be aware that if you bet on cross-asset, you are not truly isolating bets as cleanly as the Seahawks example. You are hand-waving all the other ways a bond yield, oil price, or stock earnings can change relative to one another. Thereโ€™s no equivalent to โ€œIf they donโ€™t win the NFC, the bets cancelโ€ because the relationships in assets are not deterministic in the same way that winning the Super Bowl encompasses โ€œwinning the NFCโ€.

This undermines all the logic of my trade ideas to the extent that it sets up lots of ways for the market to creatively โ€œmiddleโ€ you, just like the bettor who lays off a sports bet and gets careless about that half a point. But since all relative value trading deserves the error bars Iโ€™m dancing between, I feel a bit better having disclaimed them. Confidence sells, but when it comes to markets itโ€™s craven. Something to keep in mind when youโ€™re listening to your next podcast.

Of course, if earnings growth increased fast enough, equities donโ€™t need to decline at all to maintain the same equity risk premium to bond yields.

That leaves us with two questions:

  • What spread do todayโ€™s earnings estimates offer?
  • How much more earnings would restore the 50 bp equity risk premium?

At SPX 7,637.79, the approximately $397 of earnings expected over the next twelve months gives us a 5.20% earnings yield. Against a 4.96% Treasury yield, thatโ€™s only a 24 bp spread.

To get 50 bps over Treasuries, we need a 5.46% earnings yield:

7,637.79 ร— 5.46% = $417 of annual EPS.

How close are we to earning that much?

The first-half figures total approximately $181 per share, up 39% from the same quarters in 2025. Second-half estimates total approximately $183. The forecast calls for roughly maintaining the first-half earnings level through the rest of 2026, which represents 26% growth over the second half of 2025. Historically high, but a slower pace vs the first halfโ€™s increase over the same period a year earlier.

For the full calendar years, consensus is $362 in 2026 and $417 in 2027. Thatโ€™s another 15% growth after this yearโ€™s expected 32% increase. Together, those forecasts take annual earnings from $275 in 2025 to $417 in 2027โ€”52% growth in two years. These figures come from the same FactSet earnings series.

Back to the valuation. The calendar-2027 estimate already gets us almost exactly to our $417 target and assumes EPS growth of 15%, much more reasonable compared to historical changes and far slower growth than 2026 experienced.

Despite the rise in yields and inflation, consensus equity pricing offers an earnings path that supports approximately the original 50 bp spread benchmark. It requires delivering the rest of 2026 plus โ€œonlyโ€ 15% subsequent growth next year.

The equity risk premium has been skinny, but EPS growth has delivered such that those risk premiums arenโ€™t actually shrinking despite the continued outperformance of equities vs bonds. I wouldnโ€™t call that bullish exactly, but itโ€™s not bearish versus where we were 6 months ago. When I first started this investigation with the high energy prices and rising yields, I expected the โ€œmissing Sudoku numberโ€ of equity risk premium to look even more paltry, but sparkling earnings have bailed the multiples out.

Stay groovy

โ˜ฎ๏ธ


End Notes

Gasoline futures and CPI

Persistently high fuel prices burden consumers and businesses. But keeping the price high doesnโ€™t mean inflation stays high. If gasoline rises from $2 to $3 and stays there, year-over-year inflation is positive until the comparison catches up. Comparing $3 with $3 gives zero inflation.

The spread benchmark

โ€œEquity risk premiumโ€ here is shorthand for earnings yield minus the nominal ten-year Treasury yield. It is a valuation comparison, not a complete estimate of equitiesโ€™ expected excess return. Earnings are not contractual interest payments and are not necessarily distributed to shareholders.

Comparing that earnings yield with the roughly 2% TIPS yield gives a 300 bp difference. That is a comparison with a real bond yield, not a separately calculated real equity risk premium. My mental shorthand for real equity returns is that they typically realize between 300 and 600 bps over the risk-free rate. Equity valuation is on the high side (ie low real earnings yield over CPI), but it has been for a while, as it has been priced for growth which has been repeatedly confirmed.

Market snapshot and forward calculation

The calculations hold SPX at 7,637.79 and the ten-year Treasury yield at 4.96%, using September 14 intraday observations from MarketWatch and Trading Economics. These are not closing prices.

Tracking 2026

FactSet lists Q1 2026 EPS of $80.99 and Q2 EPS of $100.28. Q1 is shown as actual; Q2 remains marked as estimated. The $181.27 first-half total therefore should not be described as entirely finalized.

Index composition

The S&P 500โ€™s membership changes. Depending on how a historical series is constructed, year-over-year earnings changes can reflect additions, deletions and changes in index representation, as well as growth within businesses.

Comparing each periodโ€™s membership differs from comparing a fixed set of companies across both periods. The chart therefore describes the published index earnings series, not a constant-company measure of organic growth. No adjustment for composition has been made.

after this post you will be sizing bets in your head

One of the most important concepts in risk-taking is bet sizing. Which is unfortunate because people are quite bad at it, while the effort to be way above average is quite low.

A jarring and famous demonstration of this is the Haghani-Dewey Coin Flipping study, which showed how even college grads with business, economic, and technical backgrounds incinerated their capital or massively underperformed the expected profits presented to them by a game they knew was rigged in their favor.

You can read my synopsis in Bet Sizing Is Not Intuitive.

For a binary wager (ie win or lose), if you know the payoffs and the probability of winning, both of which were known to the participants, the solution is to use the Kelly Criterion.

The tragedy is that it is incredibly simple to compute and applies to many conventional gambles and decisions (the examples in the quiz will span various life situations!).

If something is both easy and widely relevant, it should be common knowledge. So letโ€™s fix that today. Iโ€™ll show you how easy it is to use, and youโ€™ll forever be able to do it in your head.

First, a succinct definition:

Kelly is the bet size, as a fraction of bankroll, that maximizes the long-run compounded growth rate of your wealth. Itโ€™s a mathematical solution to bet size that doesnโ€™t seek to maximize expected profit per trial, but the size that optimally balances compounding rate and survival.

If you want to go deep on this, see Moontowerquantโ€™s Kelly Criterion Resources, but todayโ€™s focus is on getting straight to usability.

We will use this formulation of Kelly because itโ€™s general:

f* = p โˆ’ q/b

where:

p = probability of winning

q = 1โˆ’p or probability of losing

b = the odds youโ€™re getting โ†’ what you win divided by what you risk

The easiest way to learn it is just jump right in with a few worked examples:

Fair coin wager (even odds style bet)

p =50%

q= 50%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 50% – 50% / 1 = 0 โ†’ bet nothing, zero edge

Coin biased in your favor (even odds style bet)

p =60%

q= 40%

b =1 (ie even money, for a $1 bet you either lose a $1 or make a $1 profit)

f* = 60% – 40% / 1 = .20 โ†’ bet 20% of your bankroll

Roll a 6 on a die (underdog bet where you get odds)

p =1/6

q= 5/6

b =8 (for a $1 bet, you either lose a $1 or make an $8 profit)

f* = 1/6 – (5/6) / 8

f* =8/48 – 5/48 = 3/48 โ†’ bet 6.25% of your bankroll

If f* is 0 or negative, you have no edge, so not betting is prescribed

Sports moneyline (betting as a favorite where you lay odds)

A โˆ’200 favorite. You risk $2 to win $1, and the line implies 2/3, but you think itโ€™s closer to 3 in 4.

p = 75%
q = 25%
b = 0.5 (getting 50% return on the amount you risk)

f* = 75% โˆ’ 25% / 0.5 = 75% โˆ’ 50% = 25% โ†’ bet 25% of your bankroll

Wait a minute, these are large bets?!!

If these bet sizes seem surprisingly large for the given advantages, then your senses are well-tuned. For most people, โ€œfullโ€ Kelly is too big!

Kelly maximizes long-run growth on the assumption your probability is correct. Well, it probably isnโ€™t because the world is messy. We can inject some humility by using a fraction of Kelly:

  • โ€œHalf Kellyโ€ gives up about a quarter of the growth rate and roughly halves the drawdowns. If you invert that, you see that the Kelly scaling law means as you bet bigger, you get diminishing returns per unit of risk. Extrapolating that logic, betting more than โ€œfull Kellyโ€ is incinerating compounded wealth even if the individual bet has positive EV.
  • โ€œQuarter Kellyโ€ or less is far more common in practice.

Please donโ€™t let the equation scare you, itโ€™s intuitive and easy to remember

Look at the equation again:

f* = p โˆ’ q/b

Itโ€™s just โ€œhow often you winโ€ minus โ€œhow often you lose.โ€ Itโ€™s just that the second term incorporates the payoff. The loss term gets divided by b, which represents the return you collect when youโ€™re right.

  • When b = 1, youโ€™re getting even money. A 100% return. Dividing by 1 leaves q alone, and the equation collapses to pure hit rate: p โˆ’ q. Thatโ€™s the coin case where you bet $1 to make $1.
  • When b > 1, youโ€™re getting long odds. The division shrinks the loss term. This is why the die works. You lose 5 out of 6 rolls. Straight subtraction says youโ€™re down 66 cents on the dollar and should never play, but youโ€™re paid 8-to-1, so that 5/6 becomes 5/48, and suddenly the 1/6 win percentage is the bigger number. Long odds forgive a bad hit rate.
  • When b < 1, youโ€™re laying odds. Now division stretches the loss term. The moneyline: you only lose a quarter of the time, but at โˆ’200 each loss costs you two units to earn back one, so that 25% loss percentage behaves like 50%. Being right three times out of four barely clears the bar. Lay enough odds and even a very good record is a losing proposition.

The graphic shows how you can think of the odds (the denominator) as shrinking or inflating q, as you collapse your thinking to a comparison of p vs q.

A word on b

b trips people up because โ€œoddsโ€ is loaded gambler jargon. A wider interpretation of b is that itโ€™s a percent return.

Itโ€™s what you make divided by what you risk. Even money is b = 1: risk a dollar, make a dollar. Thatโ€™s a 100% return on the amount at stake. 3-to-1 is b = 3, a 300% return. Laying โˆ’200 is b = 0.5 because if you risk two to make one, itโ€™s a 50% return.

[Return is a profit, while multiples donโ€™t subtract your initial risk. Itโ€™s the difference between โ€œI 2xโ€™d my moneyโ€ vs โ€œI made 100%โ€ or โ€œI 10xโ€™d my moneyโ€ vs โ€œI made 900%โ€. The percent return is the multiple minus one because we subtract our initial risk.]

The reason itโ€™s a return and not just โ€œthe oddsโ€ is that Kelly assumes a loss wipes out the whole stake. The denominator is always the same number: everything you put up. b is comparable across a coin, a die, and a moneyline because itโ€™s the return on risk, always measured against a total loss.

b = 1 is a natural reference point. At 100% return, a win exactly cancels a loss, so you need to win more than half the time. The breakeven hit rate changes with b.

Set f* = 0 and you get p = 1/(1+b).

Read the table as a menu of the hit rates youโ€™re allowed to have. At b = 24 you can be wrong 24 times out of 25 and still be flat. At b = 0.25 you can be right four out of five and still be flat

Applying to real life: when is Kelly the right tool?

Kelly needs a few inputs: a bankroll, a payoff you know, and a probability estimate.

Which of these is a Kelly problem?

  1. A prediction market contract trading at 30ยข. You think itโ€™s worth 45ยข.
  2. Youโ€™re all-in-or-fold on the river with a read that youโ€™re good 40% of the time, getting 3-to-1 from the pot.
  3. How much of your 401(k) to put in equities.
  4. Writing checks as an angel investor across 30 startups.
  5. Buying a weekly call on a biotech ahead of an FDA decision date.
  6. Whether to take the new job.
  7. Your buddy offers you 5-to-1 that it rains in Oakland tomorrow. The forecast says 30%.
  8. Buying homeowners insurance. The premium is clearly more than the expected loss โ€” thatโ€™s how the insurer stays in business.
  9. Your neighbor doesn’t carry homeowners coverage. She banks the premium instead.
  10. Your auto policy offers a $500 deductible or a $2,500 deductible, for a $340/yr discount on the premium.
  11. Youโ€™ve got vested startup options. Exercising costs $40k out of pocket in strike, and you think thereโ€™s maybe a 15% chance the company gets somewhere that makes them worth $1M.
  12. A merger arb spread. Targetโ€™s at $46, deal price is $50, and it trades back to $38 if the deal breaks. You think it closes 90% of the time.
  13. Your agency spends 20 hours of unbilled time on a speculative pitch. You win about a quarter of them, and a win is worth 80 billable hours.

Solutions to Kelly Problems

  1. Yes. Cleanest case there is. Binary, known payoff, and the price provides b directly. Risk 30ยข to make 70ยข, so b = 2.33. f* = 45% โˆ’ 55%/2.33 = 21%.
  2. Yes. This is the canonical one. p = 40%, b = 3, f* = 40% โˆ’ 60%/3 = 20% of your stack. The wrinkle is that in poker your stack isnโ€™t really your bankroll. Thereโ€™s a whole literature on pros using Kelly for bankroll management across sessions rather than for a single river decision.
  3. No. Not this version of it. Stock returns arenโ€™t generally binary so thereโ€™s no p, q, or b. Thereโ€™s a continuous analog called Mertonโ€™s Share, which is similarly rooted in reward vs variance. What Gamblers Can Teach the Buy-and-Hold Crowd can get you started.
  4. Sort of. The structure is right: repeated, roughly binary, long odds. The problem is that p is a guess and b is a bigger guess, and Kelly is violently sensitive to overestimating your edge. Garbage in, garbage out.
  5. Approximately. If you treat it as approve/reject itโ€™s binary enough to size with. If the expiry aligns with the date such that you are betting strictly on the terminal intrinsic value, the option piece will inherit the binary modeling you imposed on the stock.
  6. No. The variables are too opaque.
  7. Yes. p = 30%, b = 5, f* = 30% โˆ’ 70%/5 = 16%. Note, youโ€™ll lose this bet more than twice as often as you win it, so your most likely scenario is losing 16%. You can shrink the Kelly fraction if this makes you uncomfortable.
  8. Wrong side of the equation. Run f* on this and you get a negative number, because youโ€™re buying a negative-EV bet.
  9. Yes. Itโ€™s the same policy, so notice that the bet only exists on the insurerโ€™s side of it! Every year your neighbor doesnโ€™t buy, she collects a premium and writes a tail. Rebuild cost $500k, premium $3,000, call it a 1-in-500 chance of a total loss.

    p = 99.8%
    q = 0.2%
    b = 3,000 / 500,000 = 0.006 (risk $500k to win $3,000)

    f* = 99.8% โˆ’ 0.2%/0.006 = 99.8% โˆ’ 33.3% = 66.5%

    Positive, as expected since insurers price premiums well above fair value. In this case, her bet size is the $500k house. If the house is most of her net worth, sheโ€™s at 100% on a bet capped at 66%. Rather than overbet, she should buy the policy. If sheโ€™s worth $5M, sheโ€™s betting 10% when sheโ€™s allowed 66%, which puts her near quarter Kelly (~14%), and she could skip the insurance. If sheโ€™s worth $1mm, itโ€™s a 50% bet, which is more than half Kelly. Iโ€™d say take the insurance but Iโ€™m a wimp. There are other considerations (would she have the liquidity to rebuild the home or maybe taking the insurance with a high deductible is a better fit), but just doing this exercise gives you a sense of how risky or conservative your choices are relative to the bet share that maximizes long-term wealth.

  10. Yes. Raising the deductible is like you writing a $2,000 policy and collecting $340 a year for it. Youโ€™re the insurer again, so work out what you need to believe. Take the high deductible and save $340. Have a claim, and youโ€™re out $2,000 more, but you already banked the $340, so the loss is $1,660.

    b = 340 / 1,660 = 0.205

    f* = p โˆ’ q/0.205 = p โˆ’ 4.88q

    Set that to zero, and you get p = 4.88q, which, with p + q = 1, means q = 17%. Your breakeven is a claim every 5.9 years. Anything less frequent and youโ€™re the one with the edge.

    Letโ€™s say real-world collision frequency is more like 6%. So p = 94%:

    f* = 94% โˆ’ 6%/0.205 = 94% โˆ’ 29.3% = 64.7%

    Which says risk at most ~65% of your bankroll. The risk here is $1,660. That clears as long as you have about $2,600 in liquid savings, which is to say the sizing check is trivially satisfied for almost everyone so you should generally opt for the higher deductible. For quarter Kelly, weโ€™d need savings of $1,660/(.647 * .25) = $10,262.

  11. Approximately. Itโ€™s not truly binary, but if you frame it in a way where you are comfortable with the no consolation prize of a medium outcome, you can see it as paying $40k for a 15% shot at $1M. b = 24, so f* = 15% โˆ’ 85%/24 = 11.5% of your liquid net worth. This is quite sensitive to your estimate of p of course.
  12. Yes, a classic example of binary-type risk in markets. You risk $8 to make $4, so b = 0.5 โ€” youโ€™re laying odds, same as the moneyline. f* = 90% โˆ’ 10%/0.5 = 70%. That number is only as good as the 90%. Revise p to 75% and f* is 25%.
  13. Yes, in a subtle way! Your bankroll is capacity, not cash. b = 80/20 = 4, so f* = 25% โˆ’ 75%/4 = 6.25%. Twenty hours has to be 6% of what youโ€™re working with, which means you canโ€™t run this pitch out of a 100-hour month. If youโ€™re the manager, you can put it in dollar terms by converting to wages.

Finally, I strongly recommend William Poundstoneโ€™s book Fortuneโ€™s Formula: The Untold Story of the Scientific Betting System That Beat the Casinos and Wall Street

Description:

In 1956, two Bell Labs scientists discovered the scientific formula for getting rich. One was mathematician Claude Shannon, neurotic father of our digital age, whose genius is ranked with Einstein’s. The other was John L. Kelly Jr., a Texas-born, gun-toting physicist. Together they applied the science of information theoryโ€”the basis of computers and the Internetโ€”to the problem of making as much money as possible, as fast as possible.

Shannon and MIT mathematician Edward O. Thorp took the “Kelly formula” to Las Vegas. It worked. They realized that there was even more money to be made in the stock market. Thorp used the Kelly system with his phenomenally successful hedge fund, Princeton-Newport Partners. Shannon became a successful investor, too, topping even Warren Buffett’s rate of return. Fortune’s Formula traces how the Kelly formula sparked controversy even as it made fortunes at racetracks, casinos, and trading desks. It reveals the dark side of this alluring scheme, which is founded on exploiting an insider’s edge.

Shannon believed it was possible for a smart investor to beat the marketโ€”and William Poundstone’s Fortune’s Formula will convince you that he was right.

And this is from my notes, Insights From Fortuneโ€™s Formula:

a gripping narrative full of 20th century trivia that ties together the birth of information theory, some of the greatest scientific minds of the 1900s, the rise of quantitative finance, and the role of organized crime. These topics come alive in a fresh, memorable way when discovered through the lens of its colorful characters.

It chronicles the history of the efficient market hypothesis (MIT, U Chicago, Paul Samuelson). You can organize its conclusion around this excerpt:

There is much truth in the efficient market hypothesis. The controversy has always been over just how far the claim can be pressed. Asking whether markets are efficient is like asking whether the world is round. The best way to answer depends on the expectations and sophistication of the questioner. If someone is asking whether the world is round or flat, as fifteenth-century Europeans might have asked, then โ€œroundโ€ is a better answer. If someone knows that and is asking whether the earth is a geometrically perfect sphere, the answer is no.