Workflows are Moontower’s screeners organized around what you’re trying to do: sell premium (Income), buy protection (Defensive), finance that protection with your upside (Collars).
We just released a new one to help directional traders scan across the market for the best payoff for a given move: Verticals
About the Verticals Workflow
It starts from a simple need:
How do I compare payoffs across tickers in volatility-adjusted terms?
You don’t want to screen across names for the best bang-for-your-buck on a 10% rally because 10% means something very different in SPY vs MU or TSLA.
Instead we use standard deviations computed from the name’s own surface. Pick a move, say +1 SD by the November expiry, and ask the same question of every name: what’s the cheapest vertical that pays in full if the stock gets there?
The Verticals workflow returns the answer in a grid based on the watchlist you care about.
Let’s see how it works (and learn some option math in the process).
Two knobs
You set two things:
An expiry. The picker lists every expiry any name in your list carries, with a count of how many names list it. Every row in the grid is the same maturity, so you’re comparing like with like.
A signed move. ±0.5, ±1 or ±1.5 SD. Positive means call spreads, negative means put spreads.
For each name, the grid shows the tightest vertical that is fully in the money at that move. Tightest means adjacent listed strikes. Fully in the money means if the stock lands exactly on the target at expiry, you collect the whole strike width.
Finding the breakpoint
The target strike for a z-SD move is:
Then pick strikes:
Calls: the short strike is the closest listed strike at or below K_z. The long strike is the next one down.
Puts: the short strike is the closest listed strike at or above K_z. The long strike is the next one up.
Both rules keep the whole spread inside the target, so a stock that gets to the breakpoint maximizes the max spread value (ie it pays the width of the strikes.)
Marking the spread
The obvious price is the spread’s mid, but it’s far too noisy if each leg is 40 cents wide and the spread is worth 30 cents.
Collecting an accurate mark for a spread is a bit of an art, combining curve fitting and option pricing. In the app, the Price you see for the spread is not mid but a Moontower fair value which lives within the bid-ask.
moontower.ai
Reading the grid
The headline number is Payout:
A 1-point spread priced at $0.20 pays 4 to 1 if the stock gets to the target.
Payout is the first column your eye goes to, but we include several columns to judge how much to trust a mark.
Filtering for confident markets
Spread bid / ask: the spread’s own market.
Mkt width: the legged market, both legs’ bid/ask widths added together. It’s what you’d pay to cross both legs.
Mkt / strike width: Mkt width as a share of the strike width. A 20-cent-wide market on a $1 vertical might be hard to execute nearfair value.
Leg width / vol: the wider leg’s bid/ask expressed in vol points, as a share of its IV. It puts a 20-cent market on a $400 stock and a 20-cent market on a $30 stock on the same footing.
Mid: Compare with Price. When these two sit close together, the market and the model agree. When they’re far apart, the Moontower fair value is probably more reliable.
Liq: the name’s liquidity tier, which we provide in all our filters. It’sa function of how wide the markets as a percent of vol allowing us to compare across names.
The grid also drops names whose markets are too wide to mark at all. They’re counted in the “N names excluded” line under the grid, so you can see what got left out.
A few screens to start with
Tight markets only. Mkt / strike width under 20%, Liq High, sorted by Payout. This is the honest version of the leaderboard.
Where the market and the model agree. Add the Mid column and keep rows where Price is within a few cents of it.
Everything filters, groups and sorts like the other workflows, and your choices persist between visits.
This tool will let you look at the market and answer where the cheapest put spread or call spread to bet on a .5 or 1 standard deviation move by some expiry date.
You can also just reverse sort if you’re looking for spreads to sell. Knowing how much people love to sell options, I should probably rebrand this as the iron condor destroyer (an iron condor is a package of 2 OTM vertical spreads typically marketed as a less risky way to sell a strangle).
Exponential Wealth: Centuries of Stock and Bond Returnsfree PDF · 380 pp.
In 1976, Roger Ibbotson and Rex Sinquefield published “Stocks, Bonds, Bills, and Inflation” (SBBI), the first long-run total-return history of the major asset classes. Economists had talked about the equity risk premium for years but as Ibbotson puts it, “SBBI gave us a measure of it.” The data became the standard reference for historical return assumptions, from asset allocation to cost-of-capital work.
This is the 50th-anniversary update, rebuilt from CRSP after Morningstar discontinued the original indices, with the dataset now spanning a full century, 1926–2025.
It recaps what stocks, bonds, bills, and inflation have delivered, then widens out to global markets since 1900, the US before 1926, commodity futures since 1871, bubbles and crashes since 1792, and a forecast to 2050.
It’s a free, encyclopedic reference I recommend adding to your favorite LLM’s investing project folder. It’s like the Farmer’s Almanac for money.
It’s also a great reference for methodologies since the studies entail constructing price series and indices, measuring their statistical features, and addressing pitfalls such as survivorship bias, float vs full-cap weighting, shifting size definitions, overlapping observations, and effective sample sizes.
A sprinkling of fun facts:
$1 in US large caps in 1926 → $14,751 by end of 2025; $814 after inflation. Long Treasuries: $117 nominal, under $8 real. T-bills: $25 nominal, under $2 real.
Real compound returns: stocks 6.9%, long Treasuries 1.9%, bills 0.3%.
About a third of the century was spent below a prior real high: 1929–WWII, 1966–82 (16 years of zero real growth), 2000–2012.
A 50% three-year run-up was followed by another 50% run-up ~3x as often as by a full reversal (144 vs 50 of 394 episodes, 1792–2024). Crashes were followed by full recovery even more reliably.
Top 25 companies are ~50% of US market cap, a level last seen in the 1930s. Tech is ~40% of cap.
This is a non-obvious logical bit: Index total return is attainable by any investor but not all investors, for example dividend reinvestment can’t scale to everyone. Reminds me of these fallacy of composition peculiarities like “paradox of thrift” or treating government finances as if they are a household.
Investors had the least cash when expected returns are highest.
Micro cap: 16.5% arithmetic, 11.1% geometric, σ = 37%. Mid cap matches micro on compound return with σ = 24%. Post-1976, micro had the lowest compound return of any size bucket.
October 1987 was a 1-in-10¹²⁸ event under lognormal i.i.d. The chapter links fat tails to the persistent apparent overpricing of OTM S&P puts.
The book closes with a forecast for 2026–2050, using the same approach Ibbotson and Sinquefield used in 1976. It has 3 steps:
They anchor interest rates and inflation to today’s Treasury yield curve.
Resample historical risk premiums.
Simulate thousands of possible paths.
How did this method fare from 1976 to today?
Pretty well, actually. Although the composition of the return wasn’t quite what they expected.
From 1976 to 2025, the forecast called for a median of ~13.3% a year, but stocks only returned ~11.9%. However, it assumed ~6.7% inflation (it was the 70s after all) when actual inflation was only a bit more than half at 3.6%. Which means nominal returns fell short of the forecast while real returns beat it (~8.1% vs ~6.2%).
At this time, inflation and rates are closer to their 100-year averages, so the yield curve isn’t pushing the forecast much in either direction. What does it say about 2026–2050?
On raw US history: 9.5% nominal, 7.0% real per year.
The authors think that bakes in an unusually fortunate US century, so they pull the mean toward the global historical experience.
Preferred forecast: 8.0% nominal, 5.6% real, implying an equity risk premium of about 4.7% over cash. In their words, a premium “very close to what it always has been.”
Seems rosy. We’re not allowed to be this optimistic. Let’s get a second opinion. This one comes from the market itself as translated by Elm Wealth. Their capital markets assumptions, out this week, are implied from current bond yields and valuations (a cyclically adjusted earnings yield) instead of historical premiums.
[Whether CAPE is a more or less reliable forecasting tool is up for debate, but any attempt to forecast returns feels about as ill-fitting as using a vacuum to pull out a splinter. I’m in the camp that it’s a pointless exercise, while vols are more predictable and therefore a sounder basis for “know nothing” sizing, which I need because I know nothing.]
Elm reports as of September 30:
US stocks: 2.92% real, 5.30% nominal
Non-US stocks: 5.78% real, 8.15% nominal
10-year TIPS: 2.91% real
US equity risk premium over TIPS: 0.01%
Where do Elm and the authors agree?
Both take the rate and inflation pieces from market yields.
Both lean away from projecting the US past forward. Ibbotson adjusts toward global history.
Where do they depart from each other?
The equity premium. History says ~4.7% over cash. Today’s valuations say roughly zero over TIPS. Measured against the same TIPS yield, Ibbotson’s forecast is still ~2.7 points above it.
Ibbotson does address valuation. The chapter flags P/E expansion as inflating the historical record and tests stripping it out, but that actually lowered the forecast less than the global adjustment did. In other words, using global historical valuation was more conservative.
There’s some timing mismatch. Ibbotson uses the end-2025 curve while Elm’s publishing 9 months later when we know 10-year real yields are up over a full 100 bps.
Caveats both sides attach
Elm: the risk premium is a poor predictor of near-term direction; momentum is positive and its risk indicator reads low.
Ibbotson: a forecast is the center of a wide distribution, “not a guarantee.”
A lazy man’s framing
Ibbotson forecasts 5.7% real equity returns for a couple of decades. The CAPE method is close to 0. Split the difference, and you are close to the current 2.90% 10-year TIPS yield you can lock in now (for reference, Treasuries returned 1.9% real over the last 100 years.)
The market in general doesn’t appear to be especially frothy, but the risk-reward in bonds has gotten far more attractive with the latest surge in yields. If AI turns out to be deflationary and hurts employment at the same time that housing rolls over because the cost to own and finance has skyrocketed, then bonds will have looked like insurance policies with ex-ante positive carry. In English, a pretty sweet deal.
But I get it, it’s no spaceship outta the underclass.
It’s funny, when you learn about investing, you associate greed with bullishness. Turns out the textbooks have it exactly backwards in the context of our modern economic pathology psychology.
I’m trying something new that has become much easier with Meta’s personal AI assistant Muse: turning Moontower posts into Instagram Stories and Reels.
Math and personal finance ideas, one frame at a time. Muse handles the heavy lifting of reading the posts and synthesizing the arc. I then edit the frames and add music.
Just a few fun things up here since the bulk of today’s letter is heavy on investing stuff.
I Was the Wedding Planner for the Guns N’ Roses “November Rain” Ceremony and Reception |4 min read
Oh McSweeneys. I knew your writers must be jaded hipster millennials.
“Uh, no”
And here’s a clip I’m very, very late to, but making up for lost time by replaying it 20x in a row.
I just love how the announcer who gets caught off guard is so quick to come up with the most charitable interpretation of the question. Also that voice?! He was engineered in a lab to do that job.
Money Angle
Exponential Wealth: Centuries of Stock and Bond Returnsfree PDF · 380 pp.
In 1976, Roger Ibbotson and Rex Sinquefield published “Stocks, Bonds, Bills, and Inflation” (SBBI), the first long-run total-return history of the major asset classes. Economists had talked about the equity risk premium for years but as Ibbotson puts it, “SBBI gave us a measure of it.” The data became the standard reference for historical return assumptions, from asset allocation to cost-of-capital work.
This is the 50th-anniversary update, rebuilt from CRSP after Morningstar discontinued the original indices, with the dataset now spanning a full century, 1926–2025.
It recaps what stocks, bonds, bills, and inflation have delivered, then widens out to global markets since 1900, the US before 1926, commodity futures since 1871, bubbles and crashes since 1792, and a forecast to 2050.
It’s a free, encyclopedic reference I recommend adding to your favorite LLM’s investing project folder. It’s like the Farmer’s Almanac for money.
It’s also a great reference for methodologies since the studies entail constructing price series and indices, measuring their statistical features, and addressing pitfalls such as survivorship bias, float vs full-cap weighting, shifting size definitions, overlapping observations, and effective sample sizes.
A sprinkling of fun facts:
$1 in US large caps in 1926 → $14,751 by end of 2025; $814 after inflation. Long Treasuries: $117 nominal, under $8 real. T-bills: $25 nominal, under $2 real.
Real compound returns: stocks 6.9%, long Treasuries 1.9%, bills 0.3%.
About a third of the century was spent below a prior real high: 1929–WWII, 1966–82 (16 years of zero real growth), 2000–2012.
A 50% three-year run-up was followed by another 50% run-up ~3x as often as by a full reversal (144 vs 50 of 394 episodes, 1792–2024). Crashes were followed by full recovery even more reliably.
Top 25 companies are ~50% of US market cap, a level last seen in the 1930s. Tech is ~40% of cap.
This is a non-obvious logical bit: Index total return is attainable by any investor but not all investors, for example dividend reinvestment can’t scale to everyone. Reminds me of these fallacy of composition peculiarities like “paradox of thrift” or treating government finances as if they are a household.
Investors had the least cash when expected returns are highest.
Micro cap: 16.5% arithmetic, 11.1% geometric, σ = 37%. Mid cap matches micro on compound return with σ = 24%. Post-1976, micro had the lowest compound return of any size bucket.
October 1987 was a 1-in-10¹²⁸ event under lognormal i.i.d. The chapter links fat tails to the persistent apparent overpricing of OTM S&P puts.
The book closes with a forecast for 2026–2050, using the same approach Ibbotson and Sinquefield used in 1976. It has 3 steps:
They anchor interest rates and inflation to today’s Treasury yield curve.
Resample historical risk premiums.
Simulate thousands of possible paths.
How did this method fare from 1976 to today?
Pretty well, actually. Although the composition of the return wasn’t quite what they expected.
From 1976 to 2025, the forecast called for a median of ~13.3% a year, but stocks only returned ~11.9%. However, it assumed ~6.7% inflation (it was the 70s after all) when actual inflation was only a bit more than half at 3.6%. Which means nominal returns fell short of the forecast while real returns beat it (~8.1% vs ~6.2%).
At this time, inflation and rates are closer to their 100-year averages, so the yield curve isn’t pushing the forecast much in either direction. What does it say about 2026–2050?
On raw US history: 9.5% nominal, 7.0% real per year.
The authors think that bakes in an unusually fortunate US century, so they pull the mean toward the global historical experience.
Preferred forecast: 8.0% nominal, 5.6% real, implying an equity risk premium of about 4.7% over cash. In their words, a premium “very close to what it always has been.”
Seems rosy. We’re not allowed to be this optimistic. Let’s get a second opinion. This one comes from the market itself as translated by Elm Wealth. Their capital markets assumptions, out this week, are implied from current bond yields and valuations (a cyclically adjusted earnings yield) instead of historical premiums.
[Whether CAPE is a more or less reliable forecasting tool is up for debate, but any attempt to forecast returns feels about as ill-fitting as using a vacuum to pull out a splinter. I’m in the camp that it’s a pointless exercise, while vols are more predictable and therefore a sounder basis for “know nothing” sizing, which I need because I know nothing.]
Elm reports as of September 30:
US stocks: 2.92% real, 5.30% nominal
Non-US stocks: 5.78% real, 8.15% nominal
10-year TIPS: 2.91% real
US equity risk premium over TIPS: 0.01%
Where do Elm and the authors agree?
Both take the rate and inflation pieces from market yields.
Both lean away from projecting the US past forward. Ibbotson adjusts toward global history.
Where do they depart from each other?
The equity premium. History says ~4.7% over cash. Today’s valuations say roughly zero over TIPS. Measured against the same TIPS yield, Ibbotson’s forecast is still ~2.7 points above it.
Ibbotson does address valuation. The chapter flags P/E expansion as inflating the historical record and tests stripping it out, but that actually lowered the forecast less than the global adjustment did. In other words, using global historical valuation was more conservative.
There’s some timing mismatch. Ibbotson uses the end-2025 curve while Elm’s publishing 9 months later when we know 10-year real yields are up over a full 100 bps.
Caveats both sides attach
Elm: the risk premium is a poor predictor of near-term direction; momentum is positive and its risk indicator reads low.
Ibbotson: a forecast is the center of a wide distribution, “not a guarantee.”
A lazy man’s framing
Ibbotson forecasts 5.7% real equity returns for a couple of decades. The CAPE method is close to 0. Split the difference, and you are close to the current 2.90% 10-year TIPS yield you can lock in now (for reference, Treasuries returned 1.9% real over the last 100 years.)
The market in general doesn’t appear to be especially frothy, but the risk-reward in bonds has gotten far more attractive with the latest surge in yields. If AI turns out to be deflationary and hurts employment at the same time that housing rolls over because the cost to own and finance has skyrocketed, then bonds will have looked like insurance policies with ex-ante positive carry. In English, a pretty sweet deal.
But I get it, it’s no spaceship outta the underclass.
It’s funny, when you learn about investing, you associate greed with bullishness. Turns out the textbooks have it exactly backwards in the context of our modern economic pathology psychology.
Money Angle For Masochists
Workflows are Moontower’s screeners organized around what you’re trying to do: sell premium (Income), buy protection (Defensive), finance that protection with your upside (Collars).
We just released a new one to help directional traders scan across the market for the best payoff for a given move: Verticals
About the Verticals Workflow
It starts from a simple need:
How do I compare payoffs across tickers in volatility-adjusted terms?
You don’t want to screen across names for the best bang-for-your-buck on a 10% rally because 10% means something very different in SPY vs MU or TSLA.
Instead we use standard deviations computed from the name’s own surface. Pick a move, say +1 SD by the November expiry, and ask the same question of every name: what’s the cheapest vertical that pays in full if the stock gets there?
The Verticals workflow returns the answer in a grid based on the watchlist you care about.
Let’s see how it works (and learn some option math in the process).
Two knobs
You set two things:
An expiry. The picker lists every expiry any name in your list carries, with a count of how many names list it. Every row in the grid is the same maturity, so you’re comparing like with like.
A signed move. ±0.5, ±1 or ±1.5 SD. Positive means call spreads, negative means put spreads.
For each name, the grid shows the tightest vertical that is fully in the money at that move. Tightest means adjacent listed strikes. Fully in the money means if the stock lands exactly on the target at expiry, you collect the whole strike width.
Finding the breakpoint
The target strike for a z-SD move is:
Then pick strikes:
Calls: the short strike is the closest listed strike at or below K_z. The long strike is the next one down.
Puts: the short strike is the closest listed strike at or above K_z. The long strike is the next one up.
Both rules keep the whole spread inside the target, so a stock that gets to the breakpoint maximizes the max spread value (ie it pays the width of the strikes.)
Marking the spread
The obvious price is the spread’s mid, but it’s far too noisy if each leg is 40 cents wide and the spread is worth 30 cents.
Collecting an accurate mark for a spread is a bit of an art, combining curve fitting and option pricing. In the app, the Price you see for the spread is not mid but a Moontower fair value which lives within the bid-ask.
moontower.ai
Reading the grid
The headline number is Payout:
A 1-point spread priced at $0.20 pays 4 to 1 if the stock gets to the target.
Payout is the first column your eye goes to, but we include several columns to judge how much to trust a mark.
Filtering for confident markets
Spread bid / ask: the spread’s own market.
Mkt width: the legged market, both legs’ bid/ask widths added together. It’s what you’d pay to cross both legs.
Mkt / strike width: Mkt width as a share of the strike width. A 20-cent-wide market on a $1 vertical might be hard to execute nearfair value.
Leg width / vol: the wider leg’s bid/ask expressed in vol points, as a share of its IV. It puts a 20-cent market on a $400 stock and a 20-cent market on a $30 stock on the same footing.
Mid: Compare with Price. When these two sit close together, the market and the model agree. When they’re far apart, the Moontower fair value is probably more reliable.
Liq: the name’s liquidity tier, which we provide in all our filters. It’sa function of how wide the markets as a percent of vol allowing us to compare across names.
The grid also drops names whose markets are too wide to mark at all. They’re counted in the “N names excluded” line under the grid, so you can see what got left out.
A few screens to start with
Tight markets only. Mkt / strike width under 20%, Liq High, sorted by Payout. This is the honest version of the leaderboard.
Where the market and the model agree. Add the Mid column and keep rows where Price is within a few cents of it.
Everything filters, groups and sorts like the other workflows, and your choices persist between visits.
This tool will let you look at the market and answer where the cheapest put spread or call spread to bet on a .5 or 1 standard deviation move by some expiry date.
You can also just reverse sort if you’re looking for spreads to sell. Knowing how much people love to sell options, I should probably rebrand this as the iron condor destroyer (an iron condor is a package of 2 OTM vertical spreads typically marketed as a less risky way to sell a strangle).
At the Robinhood Summit last week, Mat Cashman and I hosted a “Trading Lab” emphasizing that options are surgical tools. You only want to reach for them when you have a specific view. While their payoff is highly levered to nailing a particular timing, they are unforgiving to vague forecasts like “I generally want to be long this stock”. If you buy the stock instead, you know your exposure. Your delta is 1.00 and doesn’t change.
If you buy an option, your exposure changes due to time passing alone. So that initial decision thrusts you into a series of future decisions, each one challenging you to tighten your thesis and face yet another bid/ask spread.
To demonstrate an appropriate use of options we stepped through a highly specific example during the lab and even executed the trade live!
I’m going to walk you through it here today.
Setup
The session was on Wednesday, September 30th. Nike (NKE) reports earnings after the close Thursday, October 1st.
Some context:
✔️NKE is down ~75% since its lifetime peak which occurred 5 years ago. While falling short of the peak, it did stage a large rally from late 2022 into 2023, before resuming its downward March which accelerated this year with the stock down >40% YTD at the time of the session.
✔️NKE typically pays about $1.40/yr in dividends or greater than 4% yield at recent stock prices. However, it can no longer fund the dividend from current cash flows. They could fund it from balance sheet assets (ie dip into their savings) to be able to pay it for several more years. I learned this from Brett Caughran’s tweet:
✔️I used moontower.ai’s MCP (which is now included with individual plans 😉) to ask Claude if the options market was anticipating a cut. It turns out the implied dividend about 1 year out is 50% of what NKE has been paying. NKE’s woes are out in the open, the stock and options markets see them.
With such a backdrop, we expect this to be an important earnings call. After all, if a plane is nosediving, we expect the captain to say something. (We know the content of the words will be reassuring, but we’ll listen to the tone of his voice to infer how scared or optimistic we should be.)
Pulling up the option chain
With the above context in mind, we load the October 2nd option chain. This is clean as a market gets around an earnings move since the options expire 24 hours after the earnings call.
During the session, I reason aloud through what I’m looking at.
The stock is ~$36. The 36 straddle, which can be interpreted as the average expected move size, is ~$3 or 8% of the stock price. Earnings are a big deal.
I gracefully elide the question of whether you should be bullish or bearish. If I was privy to such information, I’d be a Market Wizard. But many people have views and wanna bet. I can help you reason through your choices given you have a view. So let’s get specific.
What are you paying for?
We arbitrarily side with the bulls, but the analysis that follows works for bears as well if they substitute different contracts.
If the 36 straddle is at-the-money, meaning the stock is also $36, then the 36 call is half the straddle or $1.50.
If the stock experiences the expected move to the upside those calls will double your money as they will be worth $3 of intrinsic value at expiry. In the language of odds, you get paid 1-1 on an 8% up move.
But if this is our specific bet, can we find an expression of the trade with more octane?
We load up the 38.5/39 call spread. It’s marked for $.12 and if NKE expires $39 or higher, it will be worth $.50.
You risk $.12 to make $.38 so you are getting just over 3-1 odds for the same $39 outcome.
Why does this sound so much better than getting 1-1 if we just buy the 36 call for $1.50?
Because the outright call continues to pay off if the stock climbs more than 8%. You are paying for the extra profit potential if NKE goes up 10, 15, 20 percent. Being vague leads you to pay for more margin of error. But if you’re specific, you can get extra leverage on being right. Of course, this is risky because if you’re wrong most scenarios lead to total loss.
Let’s deal with the risk
The setup of the session was that we have $5,000 of capital.
If we are wrong on the vertical spread and the stock doesn’t clear $39 then we likely lose our entire premium.
We get a 3-to-1 payoff when we are right.
This type of binary outcome is a perfect setup to use Kelly sizing. The nice thingabout Kelly is that it assumes you have a total loss when you are wrong, so the sizing itself is the risk management. You’ve already budgeted for the maximum risk.
We’ll get to the sizing in a second but just for building fluency in thinking probabilistically, let’s acknowledge a few heuristics:
1. We know from the derivation of the straddle approximation, that the straddle is 80% of 1 standard deviation. Therefore, it encompasses ~79% of the distribution assuming normality. Therefore the upper tail, where NKE goes up more than $3, is expected 21% of the time. To get 3-1 odds would require a minimum probability of 25% to be a positive EV bet. But, normality is doing a lot of lifting in that sentence. Earnings are more likely a binary outcome, and while it’s too simplistic to say there’s a 50/50 chance of the stock going up or down 8% (which would make this vertical spread seem irresistibly cheap), it’s not unreasonable to think the probability of that spread hitting is greater than 25% (or for that matter, for the symmetrical downside put spread for those with a bearish view).
In the name of walking you through sizing GIVEN you already have the bull view, we’ll presume you believe there’s a 40% chance the stock pops higher by the amount of the straddle or $3 to $39.
How much should we bet on the vertical spread if we are getting 3-1?
Kelly:
f* = p − q/b
where:
p = probability of winning
q = 1−p or probability of losing
b = the odds you’re getting → what you win divided by what you risk
Plug and chug:
f* = .40 – .60/3 = 20% or $1,000 of the allotted capital
Converting to contracts:
Each option contract is $.12 or $12 once we adjust for the 100 share multipler.
$1,000/$12 ~ 83 contracts
We don’t want to risk overbetting if we overestimate our p (i.e., probability of winning), so let’s size to about half Kelly and only buy 40 contracts.
We actually bought the vertical spreads live during the session.
Postscript
NKE ended up falling after earnings. The equivalent put spread on the downside would have been the 33.50/33 put spread, looking to get paid if NKE fell $3.
NKE closed Friday at $33.90, so the put spread would have expired worthless. At one point, the stock touched $32.09, so the spread would have been deep in-the-money, but because there was still some time value remaining, you wouldn’t have been able to capture its maximum value, although you probably could have flipped out of it and tripled your money, from $.12 to $.36
But it would have been hard to do since, with the stock sitting around $32 you have theta on your side, and you’d probably be thinking, “even if the stock rallies $1 or 3% intraday, I’ll have turned my $.12 into $.50.”
Alas, the stock ripped 6% intraday after bottoming, leaving the put spread worthless!
Notice how tricky this is for everyone. If you were short the spreads, you ended up collecting the full $.12 of premium, but the path would have greyed your hair as it looked like you’d face maximum loss or feel forced to cover only to later watch NKE stage a heroic rally.
Closing words
Options are surgical. I think of them like term life insurance. I have a specific thing I’m trying to mitigate or take advantage. If I buy them and lose, it may or may not have been a good outcome depending on what the counterfactual would have been. The term life analogy is an obvious demonstration of benchmarking to the right outcome.
The more tightly you can wrap an option expression around a terminal value thesis, the easier the risk management is because you can budget for total loss and let sizing do the heavy work.
Thinking probabilistically is a habit of mind. Reasoning between what’s in a price and your own confidence intervals around scenarios will develop not just discipline but it gives you metrics, even if they are finger-in-the-air estimates, to journal or track, allowing you to get a bit more calibrated with each rep.
Mat Cashman and I did a teach-in session to a large audience of conference-goers who use options. Here’s one idea from that session.
What does optionality actually buy you?
The simplest way to understand the value of buying an option for directional reasons is to benchmark it to the counterfactual: how do I perform relative to “if I just bought the stock?”
If you buy a call, you do better in the large up move OR the large down move vs just buying the stock. You have more upside leverage, and if the stock craters, you only lose your premium. The tradeoff is that the option strategy underperforms owning the shares on intermediate-sized moves. That’s why we say owning an option is “long volatility” even if someone resists thinking in such admittedly abstract terms.
The chart subtracts the P/L of 100 shares from the P/L of ~2.2 calls on the 1-year 110 strike, at expiry. The calls cost $2,194 and carry the same delta as shares that cost $10,000 (100 shares of a $100 stock).
🔨Exercise: Solve for what delta the calls are
The shares win between roughly $78 and $137. The gap is widest at $110, where the calls expire worthless and the shares are up $1,000, a $3,200 difference. Below $78, the calls lose $2,194 at most while the shares keep falling. Above $110 the calls behave like 217 shares instead of 100.
If you buy a call and lose your premium, you can’t evaluate if this is a good or bad outcome without considering the counterfactual.
The INTC calls and a possible Micron sympathy trade
This was one of the examples I discussed in the live stream session where we talked about combining sources of data to reason through trade ideas.
The print. On Sep 28th, there was notable size buying in INTC October 16 130 calls.
An odd expiry. It was a lot of premium spent on high gamma & theta calls which expire a week ahead of Intel’s October 23 earnings call. Why spend that much premium on a short-dated bet that rolls off before the stock’s own catalyst?
The sympathy hypothesis. Micron was reporting on the day I was talking after the close, Sep 30th. Maybe the buyer was using INTC as a semis proxy, owning short-dated calls to catch a sympathy move off Micron’s print. It’s one possible story, not something I can confirm.
Did the seller make a mistake? A seller may think “these don’t capture earnings, so they don’t deserve a premium,” or INTC shouldn’t move much in the quiet blackout period before an earnings report, ignoring possible spillover effects from MU’s bellwether report and guidance.
The next question. Is Micron earnings even a big deal? Look at Micron’s earnings straddle. If the implied move is small, the market isn’t expecting Micron to move much. Then the sympathy thesis loses most of its appeal, since there’s less distance for INTC to get dragged along. Turn out MU earnings were priced cheaply compared to the last few years.
Then look at what INTC vol already costs. Implied vol was middle-of-the-road but high cross-sectionally (ie relative to other stocks in the current market regime). Call skew was rich, though not extreme.
Put it together. Rich vol plus a firm call skew, the primary market for MU volatility saying earnings are going to be a non-event, and a possible story for putting the INTC buyer “on a hand” adds up to making me more inclined to want to fade the INTC call buyer. Possible trades are selling INTC straddles or, if you already own the stock, writing calls against it.
You don’t have to sell the 130s to fade the 130 buyer. Heavy buying in one strike lifts vol across that whole expiry. Strikes near each other move together, so the bid in the 130 calls shows up in the 125s, the 135s, and the at-the-money options too. Selling the straddle, or whatever structure fits your book, still takes the other side of this buyer. Don’t anchor on the exact strike that printed. Put-call parity ensures the entire surface gets richer.
Finally, an analogy I didn’t make at the conference, but notice how much like poker the chain of thought is. The cards you can see (ie the flop) are the measurables like IV rank or call skew. The quantity of options purchased is the bet size. The expiry is like thinking about the position they bet from (ie “under the gun” or right after the button to last position or “dealer”). One of the market-makers’ advantages is that they keep tabs on notable flow, like a poker bot which examines online players’ hand histories and tendencies.
One of my close friends runs an advisory for HS students applying to college. He wrote this guest post years ago: Moneyballing College Admissions.
He lives in the Bay Area and told me he’s seeing more parents complaining about the lack of acceleration options in the public school. He said outside CA it’s becoming far more common for high schoolers to take AP Calculus BC before senior year. My 8th graders’ goal is to take it by junior year, giving himself a chance to take Multivariate Calculus by senior year at a local college. This is already offered in some high schools across the nation, including some on the peninsula. Just this week, I talked to someone whose 9th grader was in Calc BC. That only sounds crazy if you haven’t been paying attention to what I’ve been sharing about Math Academy students (my 8th grader says his 5th grade little bro is already doing the same stuff his “advanced” math class is doing.)
In our local school district, we are finally seeing the pendulum swing the other way on math instruction. A few years ago, they got rid of “tracking” in 6th grade, forcing all kids, regardless of their interest/aptitude, into the same class. Well, the seeds of change are obvious in this survey I just filled out, as I don’t think we could have even had this conversation 5 years ago:
That #7 offers choices besides “supportive” and “very supportive” tells me these 2 articles are as important as I think they are.
Pamela Hobart argues “the academic acceleration community needs a viable brand”. Co-sign. Terrific post. I’ve had many of the same thoughts, especially as my kid’s schoolwork will leave me wondering, as Pamela has, “what are we even doing here?”
Everyone Wants a Child Like Eileen Gu (10 min read)
”almost nobody today wants the childhood that produced her.”
Violet Gordeljevic spitting truth in this one. I’ll share a few excerpts that resonated with me. You should read it to see what lands or doesn’t for you.
On Discipline and Exceptionalism
“So when people say she sounds and talks so disciplined, and that she knows her own mind—well, yes. That is what a decade of being asked to do your best looks like from the outside.”
“Nobody arrives at exceptional by accident. It doesn’t turn up later as a nice surprise because the child was left alone and bored for long enough.”
On Passion and Competence
“Passion doesn’t usually come first and then get supported. Usually exposure comes first, then competence (from actually having gone through practice), and passion tends to arrive somewhere after that, because human beings mostly love the things they are good at.”
“What is important to understand is that she could not have fallen in love with skiing, or Mandarin, or Olympiad maths, if nobody had ever put her in front of these things…A child who is never introduced to numbers will not reveal unusual mathematical ability. A child who never sits at an instrument will not discover she has an ear.”
On Modern Parenting and “Full Schedules”
“We didn’t lighten the load at all. We drive our children to more things than any generation in history. We just stopped asking them to become good at any of it.”
“If we are honest, the schedule is full but the demand is zero. It’s honestly like we have swapped skill for entertainment.”
“Underneath every version of this conversation sits the same line: I don’t need my child to be exceptional, I just want them to be happy. … But being the parent who stretches a child—who asks for one more go, who expects them to learn to write something properly or figure out that math equation, who holds the line at six when she wants to quit ballet—doing this is harder and considerably less pleasant than being her friend.”
On Warmth, Adversity, and Long-Term Value
“Protecting a child from effort is not the same as protecting a child.”
“What separates people with more demanding childhoods from those with easy, low demand ones, isn’t how much or even what was asked. It’s rather whether the asking sat inside enough warmth to be survivable.”
“The choice was never between pressure and love. It was always meant to be both.”
And then my 2 favorite:
“I have met a great many people who regret being allowed to quit”
“The thing is, children are not good judges of what they will be glad to be good at at thirty.”
One of my close friends runs an advisory for HS students applying to college. He wrote this guest post years ago: Moneyballing College Admissions.
He lives in the Bay Area and told me he’s seeing more parents complaining about the lack of acceleration options in the public school. He said outside CA it’s becoming far more common for high schoolers to take AP Calculus BC before senior year. My 8th graders’ goal is to take it by junior year, giving himself a chance to take Multivariate Calculus by senior year at a local college. This is already offered in some high schools across the nation, including some on the peninsula. Just this week, I talked to someone whose 9th grader was in Calc BC. That only sounds crazy if you haven’t been paying attention to what I’ve been sharing about Math Academy students (my 8th grader says his 5th grade little bro is already doing the same stuff his “advanced” math class is doing.)
In our local school district, we are finally seeing the pendulum swing the other way on math instruction. A few years ago, they got rid of “tracking” in 6th grade, forcing all kids, regardless of their interest/aptitude, into the same class. Well, the seeds of change are obvious in this survey I just filled out, as I don’t think we could have even had this conversation 5 years ago:
That #7 offers choices besides “supportive” and “very supportive” tells me these 2 articles are as important as I think they are.
Pamela Hobart argues “the academic acceleration community needs a viable brand”. Co-sign. Terrific post. I’ve had many of the same thoughts, especially as my kid’s schoolwork will leave me wondering, as Pamela has, “what are we even doing here?”
Everyone Wants a Child Like Eileen Gu (10 min read)
”almost nobody today wants the childhood that produced her.”
Violet Gordeljevic spitting truth in this one. I’ll share a few excerpts that resonated with me. You should read it to see what lands or doesn’t for you.
On Discipline and Exceptionalism
“So when people say she sounds and talks so disciplined, and that she knows her own mind—well, yes. That is what a decade of being asked to do your best looks like from the outside.”
“Nobody arrives at exceptional by accident. It doesn’t turn up later as a nice surprise because the child was left alone and bored for long enough.”
On Passion and Competence
“Passion doesn’t usually come first and then get supported. Usually exposure comes first, then competence (from actually having gone through practice), and passion tends to arrive somewhere after that, because human beings mostly love the things they are good at.”
“What is important to understand is that she could not have fallen in love with skiing, or Mandarin, or Olympiad maths, if nobody had ever put her in front of these things…A child who is never introduced to numbers will not reveal unusual mathematical ability. A child who never sits at an instrument will not discover she has an ear.”
On Modern Parenting and “Full Schedules”
“We didn’t lighten the load at all. We drive our children to more things than any generation in history. We just stopped asking them to become good at any of it.”
“If we are honest, the schedule is full but the demand is zero. It’s honestly like we have swapped skill for entertainment.”
“Underneath every version of this conversation sits the same line: I don’t need my child to be exceptional, I just want them to be happy. … But being the parent who stretches a child—who asks for one more go, who expects them to learn to write something properly or figure out that math equation, who holds the line at six when she wants to quit ballet—doing this is harder and considerably less pleasant than being her friend.”
On Warmth, Adversity, and Long-Term Value
“Protecting a child from effort is not the same as protecting a child.”
“What separates people with more demanding childhoods from those with easy, low demand ones, isn’t how much or even what was asked. It’s rather whether the asking sat inside enough warmth to be survivable.”
“The choice was never between pressure and love. It was always meant to be both.”
And then my 2 favorite:
“I have met a great many people who regret being allowed to quit”
“The thing is, children are not good judges of what they will be glad to be good at at thirty.”
Money Angle
Wrapping up Kelly
I know I’ve published a few articles on the Kelly Criterion recently, but if you prefer other mediums:
I was in Houston this week for the Robinhood Summit to help with 2 sessions.
A “trading lab” class where Mat Cashman and I gave a lesson on buying options.
I was also invited to play the role of “trader” on a panel featuring a few of the data providers on RH’s new agentic trading platform. You can think of it like an app store, where users can subscribe to have vetted providers’ data accessible within RH’s trading agent.
Since this panel was on the “main stage” the recording is available here:
The kids were able to watch the live stream before school
Money Angle For Masochists
2 examples from the Summit.
Mat Cashman and I did a teach-in session to a large audience of conference-goers who use options. Here’s one idea from that session.
What does optionality actually buy you?
The simplest way to understand the value of buying an option for directional reasons is to benchmark it to the counterfactual: how do I perform relative to “if I just bought the stock?”
If you buy a call, you do better in the large up move OR the large down move vs just buying the stock. You have more upside leverage, and if the stock craters, you only lose your premium. The tradeoff is that the option strategy underperforms owning the shares on intermediate-sized moves. That’s why we say owning an option is “long volatility” even if someone resists thinking in such admittedly abstract terms.
The chart subtracts the P/L of 100 shares from the P/L of ~2.2 calls on the 1-year 110 strike, at expiry. The calls cost $2,194 and carry the same delta as shares that cost $10,000 (100 shares of a $100 stock).
🔨Exercise: Solve for what delta the calls are
The shares win between roughly $78 and $137. The gap is widest at $110, where the calls expire worthless and the shares are up $1,000, a $3,200 difference. Below $78, the calls lose $2,194 at most while the shares keep falling. Above $110 the calls behave like 217 shares instead of 100.
If you buy a call and lose your premium, you can’t evaluate if this is a good or bad outcome without considering the counterfactual.
The INTC calls and a possible Micron sympathy trade
This was one of the examples I discussed in the live stream session where we talked about combining sources of data to reason through trade ideas.
The print. On Sep 28th, there was notable size buying in INTC October 16 130 calls.
An odd expiry. It was a lot of premium spent on high gamma & theta calls which expire a week ahead of Intel’s October 23 earnings call. Why spend that much premium on a short-dated bet that rolls off before the stock’s own catalyst?
The sympathy hypothesis. Micron was reporting on the day I was talking after the close, Sep 30th. Maybe the buyer was using INTC as a semis proxy, owning short-dated calls to catch a sympathy move off Micron’s print. It’s one possible story, not something I can confirm.
Did the seller make a mistake? A seller may think “these don’t capture earnings, so they don’t deserve a premium,” or INTC shouldn’t move much in the quiet blackout period before an earnings report, ignoring possible spillover effects from MU’s bellwether report and guidance.
The next question. Is Micron earnings even a big deal? Look at Micron’s earnings straddle. If the implied move is small, the market isn’t expecting Micron to move much. Then the sympathy thesis loses most of its appeal, since there’s less distance for INTC to get dragged along. Turn out MU earnings were priced cheaply compared to the last few years.
Then look at what INTC vol already costs. Implied vol was middle-of-the-road but high cross-sectionally (ie relative to other stocks in the current market regime). Call skew was rich, though not extreme.
Put it together. Rich vol plus a firm call skew, the primary market for MU volatility saying earnings are going to be a non-event, and a possible story for putting the INTC buyer “on a hand” adds up to making me more inclined to want to fade the INTC call buyer. Possible trades are selling INTC straddles or, if you already own the stock, writing calls against it.
You don’t have to sell the 130s to fade the 130 buyer. Heavy buying in one strike lifts vol across that whole expiry. Strikes near each other move together, so the bid in the 130 calls shows up in the 125s, the 135s, and the at-the-money options too. Selling the straddle, or whatever structure fits your book, still takes the other side of this buyer. Don’t anchor on the exact strike that printed. Put-call parity ensures the entire surface gets richer.
Finally, an analogy I didn’t make at the conference, but notice how much like poker the chain of thought is. The cards you can see (ie the flop) are the measurables like IV rank or call skew. The quantity of options purchased is the bet size. The expiry is like thinking about the position they bet from (ie “under the gun” or right after the button to last position or “dealer”). One of the market-makers’ advantages is that they keep tabs on notable flow, like a poker bot which examines online players’ hand histories and tendencies.
From My Actual Life
Once the RH Summit ended, I was rewarded with some nice personal moments. I took Mat Cashman to Camaraderie in the Heights area of Houston for an exceptional meal by super chef Shawn Gawle.
Cashman is a fitting name for a trader
Shawn is a friend of mine. He lived with Yinh and I for a couple of months when he first moved to SF to be the pastry chef at Saison. If you are in Houston, do not miss Camaraderie.
Finally getting to experience what I fully expected to live up to the hype.
Then on Friday, Yinh and I celebrated our 17th wedding anniversary. (We’ve been together for 23 years in total).
We both played hooky and went hiking up in Fairfax (Marin County) before an exceptional dinner at Vin in San Rafael. It’s both casual (we were still in hiking clothes) and worth going out of your way for. Unbelievable eating week. Might as well. I have a colonoscopy later this week, so I need to follow some dietary restrictions for the next few days.
We also did a bit of shopping in downtown San Rafael. Picked up some used vinyl at Red Devil and also went to one of the best game shops I’ve ever been in:
This Gamescape has a shared lineage with the classic one on Divisadero in SF but the San Rafael one is much larger. Their game wall is immense and ordered from least to most complex.
I picked up these 2:
Acquire is one of my favorite games and while I have it already, this newer edition has irresistible plastic buildings. I must support this.
Brass Birmingham has been at or near the top of the BGG rankings for nearly a decade, which is Nigel Richards’ level of domination. My favorite boardgame is the first edition of Brass released in 2007, but supposedly the 2018 Brass Birmingham is much better. Time to find out.
Stand at one spot on a curve, measure everything you can there, and use those measurements to guess the height somewhere else.
You’d do this when the curve is hard to compute everywhere but easy to measure at one point, or when you want to see what drives a change. A bond’s price is a familiar case: you know its yield and duration today and want to know what happens if rates move.
The guess is built in layers. Each layer uses one more thing you measured at the anchor. The first layer is a straight line. The second bends it. The third bends the bend. You stop when the layers stop mattering, or when you run out of measurements.
Worked example: guess x³ at x = 1.2 using only what you know at x = 1
The curve is y = x³. You are standing at x = 1. The true answer is 1.2³ = 1.728, but pretend you can’t compute it.
What you know at x = 1
Thing
How you get it
Value at x = 1
Height
x³
1
Slope
derivative, 3x²
3
How fast the slope changes
derivative of that, 6x
6
How fast that changes
derivative of that
6
Anything further
derivative of a constant
0
The walk. Destination minus start: 1.2 − 1 = 0.2. Call it h.
Layer 1: pretend the slope stays 3. Guess = height + slope × walk = 1 + 3 × 0.2 = 1.6. Off by 0.128.
Layer 2: the slope drifts. It goes up at 6 per unit, so over the walk it rises from 3 to 3 + 6 × 0.2 = 4.2. Use the average slope, 3.6, instead of 3. Guess = 1 + 3.6 × 0.2 = 1.72. Off by 0.008.
Written as a separate correction: the new piece is ½ × 6 × 0.2² = 0.12, added to the 1.6.
Layer 3: the drift rate drifts. Same move one level deeper. The correction is 6 × 0.2³ ÷ 6 = 0.008. Guess = 1.728. Off by exactly 0.
Layer 4 and beyond: the measurement is 0, so every further correction is 0.
The guess is now perfect at x = 1.2, and it’s perfect at every other x too. x³ only has three pieces of information in it. Use all three and you have rebuilt the function.
Try it at x = 3 (walk h = 2): 1 + 3(2) + 3(2)² + (2)³ = 1 + 6 + 12 + 8 = 27 = 3³. Still exact, even two units from the anchor.
Layer 1 is the tangent line. Layer 2 bends it into a parabola that hugs the curve near x = 1 but misses on both sides. Layer 3 (dashed) lies on top of the black x³ curve at every x, which is the whole point: for a polynomial, enough layers is exactly the function.
Why you divide by 2, then 6, then 24
Each layer’s measurement gets divided before it’s used: layer 1 by 1, layer 2 by 2, layer 3 by 6, layer 4 by 24. Those are 1!, 2!, 3!, 4!. Two ways to see why.
The averaging picture. Layer 2 used the average slope over the walk. The slope was a ramp going from 3 to 4.2, and the average of a ramp is halfway: that’s the ÷2. Layer 3 needs the average of something that grows like a parabola, and a parabola from zero spends most of the walk being small, so its average is only a third of its end value: another ÷3. Stack them: 2 × 3 = 6. One more level and the average of a cubic is a quarter: 2 × 3 × 4 = 24.
Check the parabola claim with numbers. Sample t² at t = 0.1, 0.2, …, 1.0: you get 0.01, 0.04, 0.09, 0.16, 0.25, 0.36, 0.49, 0.64, 0.81, 1.00. They add to 3.85, average 0.385, and with finer sampling it settles to ⅓.
The power-raising picture. Differentiating hⁿ gives n × hⁿ⁻¹: lowering a power by one multiplies by that power. So raising a power by one divides by it. To turn a constant measurement into a term in h³ you raise the power three times, paying ÷1, ÷2, ÷3 along the way. That product is 3! = 6.
Layer
Measurement (x³ at 1)
Divide by
Term
1
3
1
3h
2
6
2
3h²
3
6
6
h³
4
0
24
0
Worked example: ln x, where the corrections stop helping
Same game, anchor at x = 1. Height is ln 1 = 0. Slope is 1/x, so 1 at the anchor.
What you know at x = 1
Layer
Derivative
Value at 1
Divide by
Term
1
1/x
1
1
h
2
−1/x²
−1
2
−h²/2
3
2/x³
2
6
h³/3
4
−6/x⁴
−6
24
−h⁴/4
5
24/x⁵
24
120
h⁵/5
6
−120/x⁶
−120
720
−h⁶/6
The measurements never hit zero. They alternate sign and grow. So there is always another correction, forever. Whether the corrections help depends on how far you walk.
Short walk: x = 1.5, so h = 0.5. True value ln 1.5 = 0.4055.
Layers used
Guess
Gap
1
0.5
0.0945
2
0.375
0.0305
3
0.4167
0.0112
4
0.4010
0.0044
5
0.4073
0.0018
6
0.4047
0.0008
Each layer roughly halves the gap. Keep going and it heads to zero.
Long walk: x = 2.5, so h = 1.5. True value ln 2.5 = 0.9163.
Layers used
Guess
Gap
1
1.5
0.5837
2
0.375
0.5413
3
1.5
0.5837
4
0.2344
0.6819
5
1.7531
0.8368
6
−0.1453
1.0616
The gap gets worse. Each term is ±hᵏ/k, and with h = 1.5 the hᵏ grows faster than the k can shrink it: 1.5, 1.125, 1.125, 1.27, 1.52, 1.90, … The guess swings wider and wider around the truth.
Inside the band the colored curves pile onto the black one, each layer tighter than the last. Right of x = 2 they fan out: layer 3 shoots up, layer 4 dives, layer 5 shoots higher, layer 6 dives harder. Every extra layer swings further from ln x instead of closer.
The rule. For ln x about 1, the corrections help when |h| < 1 and hurt when |h| > 1. That distance, 1, is the radius of convergence. It’s set by where the function itself breaks: ln x blows up at x = 0, exactly one unit left of the anchor, and the series can’t reach further right than it can reach left.
x³ had no such limit because its corrections ran out before they could misbehave. That is the difference between a polynomial and everything else.
the curve you’re guessing; f(x) is its height at x
x³
x
“x”
the destination, any point on the horizontal axis
1.2
x₀
“x-nought”
the anchor, where you stood and took measurements
1
x − x₀
“the walk”
destination minus start, also written h
0.2
P
“the polynomial”
the guess; a polynomial because it’s a sum of powers of the walk
1 + 3h + 3h² + h³
n
“n”
how many layers you used; the highest power in the guess
3
Pₙ(x)
“P-n of x”
the guess using n layers, evaluated at x
P₃(1.2) = 1.728
k
“k”
the counter: which layer you’re on, running 0, 1, 2, … up to n
0, 1, 2, 3
Σₖ₌₀ⁿ
“sum from k = 0 to n”
add up the term for every k from 0 through n
four terms
f′, f″, f‴
“f-prime, double-prime, triple-prime”
first, second, third derivative: slope, rate of slope, rate of that
3x², 6x, 6
f⁽ᵏ⁾
“f-k”
the k-th derivative; f⁽⁰⁾ is f itself, f⁽¹⁾ is f′, and so on
f⁽²⁾ = 6x
f⁽ᵏ⁾(x₀)
“f-k at x-nought”
the k-th derivative evaluated at the anchor, a plain number
1, 3, 6, 6
k!
“k factorial”
1 × 2 × … × k, the divide-by column; 0! = 1
1, 1, 2, 6
(x − x₀)ᵏ
“the walk to the k”
the walk raised to the layer number
1, 0.2, 0.04, 0.008
f(x) − Pₙ(x)
“the gap”
true height minus guess
0
ξ
“xi” (Greek letter)
some unknown point between x₀ and x; used only in the error formula
somewhere in [1, 1.2]
So the k = 2 term of the sum is: take the second derivative (6x), evaluate at the anchor (6), divide by 2! (3), multiply by the walk squared (0.04). That’s 0.12, the layer-2 correction from the worked example.
How big is the gap? It’s controlled by the next measurement you didn’t use, taken somewhere along the walk:
f(x) − Pₙ(x) = f⁽ⁿ⁺¹⁾(ξ) / (n+1)! · (x − x₀)ⁿ⁺¹ for some ξ between x₀ and x
Read it as: the gap is small when the walk is short (the hⁿ⁺¹ is tiny), when the next derivative is tame, or when n is large enough that (n+1)! dominates. The gap is large when the walk is long and the higher derivatives are big, which is exactly what happened to ln x at x = 2.5.
Real-world uses
Four places you’ve met this without the name. Each is worked with numbers.
Bond prices: duration and convexity. A 10-year zero at a 4% yield is priced 100/1.04¹⁰ = 67.56. The anchor is 4%. The two measurements are duration (layer 1, slope) = 10/1.04 = 9.62 and convexity (layer 2) = 10 × 11/1.04² = 101.7.
Yield rises 1%, so the walk is 0.01:
Layers
Guess
True price
Gap
1 (duration only)
67.56 × (1 − 9.62 × 0.01) = 61.06
61.39
0.33
2 (add convexity)
61.06 + 67.56 × ½ × 101.7 × 0.01² = 61.40
61.39
0.01
Yield rises 3%, walk 0.03: duration alone says 48.07, convexity pulls it to 51.16, true is 50.83. The gap is 30 times larger than for the 1% move. That’s the long walk. Traders quote duration and convexity for exactly the reason your tool quotes delta and gamma.
How a calculator computes sin. There’s no sin key inside the chip; it sums the series about 0: x − x³/6 + x⁵/120 − x⁷/5040 + ….
sin(0.5): 0.5 − 0.0208 + 0.0003 = 0.4794. True value 0.4794. Three terms.
sin(3): 3 − 4.5 + 2.025 − 0.434 + 0.050 − 0.004 = 0.137. True value 0.141. Six terms and still off in the third decimal. The series converges everywhere, unlike ln x, but a long walk needs many more layers. Calculators dodge this by folding the input back to a small angle first, which is the “walk less” strategy.
Pendulum clocks. The textbook period T = 2π√(L/g) comes from replacing sin θ with θ, which is layer 1 of the sine series. The next layer says the true period is longer by a factor of about 1 + θ₀²/16, where θ₀ is the swing in radians.
Swing
θ₀ in radians
Correction
Period error if you ignore it
5°
0.087
1.0005
0.05%
20°
0.349
1.0076
0.8%
60°
1.047
1.069
7%
A clock built on layer 1 keeps time at small swings and drifts at big ones. Same story: the anchor is θ = 0 and the walk is the amplitude.
Compound growth. (1 + r)ⁿ about r = 0 is 1 + nr + n(n−1)r²/2 + …. For 5% over 10 years: layer 1 says 1.50, layer 2 says 1.50 + 45 × 0.0025 = 1.61, true is 1.63. Layer 1 alone is the “simple interest” mental shortcut, and the layer-2 term is exactly how much compounding beats it.
All four break the same way ln x did. Past some size of move the layers you kept stop describing the function, and the fix is either more layers, a shorter walk, or computing the real thing.
Quick reference
Recipe for any function about any anchor
Pick the anchor x₀. Compute the walk h = x − x₀.
Take derivatives of f until you have as many as you want, and evaluate each at x₀.
Divide the k-th one by k!, multiply by hᵏ.
Add them up. That’s the guess. The leftover is the gap.
If the derivatives hit zero, the guess becomes exact. If they don’t, check whether the terms are shrinking; if not, you walked too far.
The two examples side by side
x³ about 1
ln x about 1
Derivatives at anchor
1, 3, 6, 6, 0, 0, …
0, 1, −1, 2, −6, 24, …
k-th term
3h, 3h², h³, then 0
(−1)ᵏ⁺¹ hᵏ / k
Exact after
3 layers
never
Works for
every x
0 < x < 2 only
Why
polynomial: information runs out
ln x breaks at 0, one unit from the anchor
Factorials
k
k!
Average of tᵏ over [0, h]
1
1
h/2
2
2
h²/3
3
6
h³/4
4
24
h⁴/5
5
120
h⁵/6
Common series about 0, for reference
Function
Series
Converges for
eˣ
1 + x + x²/2 + x³/6 + …
all x
sin x
x − x³/6 + x⁵/120 − …
all x
cos x
1 − x²/2 + x⁴/24 − …
all x
1/(1−x)
1 + x + x² + x³ + …
|x| < 1
ln(1+x)
x − x²/2 + x³/3 − …
−1 < x ≤ 1
The last two have a radius because the function breaks at x = 1 or x = −1. The first three never break, so the series works everywhere even though it never terminates.
Key Insights
What
Why it matters
Anchor and walk
Everything is measured at one point; the guess only ever knows about that point
Layers = derivatives
Each derivative at the anchor buys one more correction; that is all the information you have
Divide by k!
Raising a power costs a division each time; averaging a ramp, parabola, cubic costs ÷2, ÷3, ÷4
Polynomials terminate
Derivatives hit zero, so finitely many layers rebuild the function exactly, everywhere
Radius of convergence
For everything else, past the distance to the nearest breakdown the layers make the guess worse
The gap formula
The error is the first term you dropped, evaluated somewhere on the walk
My second favorite thing after a piece of art, writing, music, movie, and sport that I love is watching someone else gush over things they love. It’s why the Lost in Vegas guys are so endearing, especially when the channel first started. You were watching hip-hop heads not only discover but unpeel the layers in Tool’s music.
One of my dad friends recently bought a 1979 Jeep to restore with his son and asked if my son Zak (13) would like to help. Zak’s “hell yea” beat the sound of my last syllable when I asked him. He started watching YouTube, and we landed on a channel where these obsessed details comb the U.S. looking for “barn finds” to clean for the owners. For free! It’s free because the channel has almost 2 million subscribers clamoring to watch an impossible cleaning.
I watched a few videos. I get it. Renewal of beauty, just like the objects of renewal never goes out of style. The process is as timelessly rewarding as the result.
Caitlin’s tweet is an appeal to our spirit of obsession and craft.
This essay not only feels good but is an important lens on the fast-approaching experiment of whether infinite electric monkeys will pound out Shakespeare.
Normally I might apologize for extensive excerpting but they are the point. I can do no better than this.
It opens:
I owe an apology to every English teacher I ever had. I always assumed that so-called “great” literature was a hoax, a punishment inflicted upon adolescents for the crime of being young. These books did not have anything special about them, and covering up that fact was simply a make-work exercise for former English majors, a sort of “jobs for snobs” program.
I was wrong about this. There is such thing as greatness. More specifically, there is such thing as thickness. Great works of fiction—for that matter, great works of any art—unfurl in response to your attention. The more time you spend with them, the more you get out of them. That kind of responsiveness is so addicting that it can lead people to do crazy things, like try to teach literature to high schoolers.
But thickness is tricky, because rewarding the careful reader often means repelling the casual one. And this is where I would like an apology in return from my English teachers, because while this might have been obvious to them, they never made it obvious to me.
I was presented with art and literature as if it was self-explanatory, and that everything wonderful about it was plainly visible from the outside. But those works were much more like dark, winding caves with treasure stashed inside of them. My teachers were like, “Right, well, into the cave you go!” and I was like “But there’s nothing in there” and they were like “Entering the cave is 30% of your grade” and so I took a few steps into the darkness and I was like “Just as I suspected: an empty cave” and then I came trudging back out and pretended that I saw something.
It goes on to describe spectacular examples of “thickness”. You won’t want to miss the The Garden of Earthly Delights painting and its unintended invitation to hear “butt music”.
Adam presents 4 qualities that make something thick with examples from a children’s book, the “ears” in Hamlet, and Penn (of Penn and Teller fame) eating fire.
There’s an amazing takedown of what passes for popular non-fiction.:
Reading a book like this feels like wandering through a Potemkin village. Touch any of the ideas, and they tip over.
He contrasts gilded examples of non-fiction with solid gold:
Thickness comes from surfacing a few facts well, and in such a way that you realize the existence of entire universes of additional facts that could be known….[Jane Jacobs book] is pointing out a fact that millions of people observe every day, but almost none of them notice.
He addresses an easily anticipated objection which if you’ll be familiar with if you have read any of deBoer on “poptimism”:
If you allow for the existence of secret treasures that can only be accessed with effort and analysis, then you empower the elitists and the snobs. “Buddy, don’t even talk to me until you’ve been in the cave!”
But look around. The snobs are in full retreat. We have swung the pendulum so far toward poptimism, toward the blinkered idea that all art is equal because all humans are equal, toward the ethos that guilty pleasures are simply pleasures, that I’m not sure if we can ever swing it back.
And finally, Adam addresses the elephant:
Erasing the line between the thick and the thin has left us defenseless against slop at the exact moment of its onslaught. Everyone can sense there’s something amiss with the prose that comes out of the machines, but we lack the language to talk about it, and so we’ve converged on the idea that slop simply means using too many em dashes, bullet points, and line breaks.
What separates substance from slop is thickness. Slop holds no secrets; it signifies nothing. Under scrutiny, it evaporates. All it can offer is bottomlessness—sure, there’s nothing on, but at least there are infinite channels!
That’s why I’m neither surprised nor dismayed when studies find that people prefer AI art to human art. Of course they do! In the short term, thinness prevails. When people are making snap judgments, they want pretty flowers, poems that rhyme, pleasing pablum, the simulacrum of thought. But none of these last…
About once a week, I get a pitch from some AI startup that wants to automate some part of my writing. The most recent one says it’s “built for credible thinkers who have a book’s worth of ideas but not the time it typically takes to write one”.
I’m sorry, but if you’re building or using a tool like this, then you’ve got slop for brains. There is no such thing as having a “book’s worth of ideas” that are all ready to go except for the small matter of choosing the right words and putting them in the right order.
I know exactly the feeling that these slop-trepeneurs are preying on, because I feel it all the time: I’ve got these thoughts in my head, and boy oh boy they’re good ones, all-timers, really, and it’s so annoying that I have to spend all this time making the words sound good, when the ideas behind the words are already so good!
But this is an illusion. The ideas are not already good. They need to be thickened. I understand why it’s tempting to force a machine do the hard part for you, but it can’t, and the hard part is the only part worth doing anyway.
My second favorite thing after a piece of art, writing, music, movie, and sport that I love is watching someone else gush over things they love. It’s why the Lost in Vegas guys are so endearing, especially when the channel first started. You were watching hip-hop heads not only discover but unpeel the layers in Tool’s music.
One of my dad friends recently bought a 1979 Jeep to restore with his son and asked if my son Zak (13) would like to help. Zak’s “hell yea” beat the sound of my last syllable when I asked him. He started watching YouTube, and we landed on a channel where these obsessed details comb the U.S. looking for “barn finds” to clean for the owners. For free! It’s free because the channel has almost 2 million subscribers clamoring to watch an impossible cleaning.
I watched a few videos. I get it. Renewal of beauty, just like the objects of renewal never goes out of style. The process is as timelessly rewarding as the result.
Caitlin’s tweet is an appeal to our spirit of obsession and craft.
This essay not only feels good but is an important lens on the fast-approaching experiment of whether infinite electric monkeys will pound out Shakespeare.
Normally I might apologize for extensive excerpting but they are the point. I can do no better than this.
It opens:
I owe an apology to every English teacher I ever had. I always assumed that so-called “great” literature was a hoax, a punishment inflicted upon adolescents for the crime of being young. These books did not have anything special about them, and covering up that fact was simply a make-work exercise for former English majors, a sort of “jobs for snobs” program.
I was wrong about this. There is such thing as greatness. More specifically, there is such thing as thickness. Great works of fiction—for that matter, great works of any art—unfurl in response to your attention. The more time you spend with them, the more you get out of them. That kind of responsiveness is so addicting that it can lead people to do crazy things, like try to teach literature to high schoolers.
But thickness is tricky, because rewarding the careful reader often means repelling the casual one. And this is where I would like an apology in return from my English teachers, because while this might have been obvious to them, they never made it obvious to me.
I was presented with art and literature as if it was self-explanatory, and that everything wonderful about it was plainly visible from the outside. But those works were much more like dark, winding caves with treasure stashed inside of them. My teachers were like, “Right, well, into the cave you go!” and I was like “But there’s nothing in there” and they were like “Entering the cave is 30% of your grade” and so I took a few steps into the darkness and I was like “Just as I suspected: an empty cave” and then I came trudging back out and pretended that I saw something.
It goes on to describe spectacular examples of “thickness”. You won’t want to miss the The Garden of Earthly Delights painting and its unintended invitation to hear “butt music”.
Adam presents 4 qualities that make something thick with examples from a children’s book, the “ears” in Hamlet, and Penn (of Penn and Teller fame) eating fire.
There’s an amazing takedown of what passes for popular non-fiction.:
Reading a book like this feels like wandering through a Potemkin village. Touch any of the ideas, and they tip over.
He contrasts gilded examples of non-fiction with solid gold:
Thickness comes from surfacing a few facts well, and in such a way that you realize the existence of entire universes of additional facts that could be known….[Jane Jacobs book] is pointing out a fact that millions of people observe every day, but almost none of them notice.
He addresses an easily anticipated objection, which you’ll be familiar with if you have read any of deBoer on “poptimism”:
If you allow for the existence of secret treasures that can only be accessed with effort and analysis, then you empower the elitists and the snobs. “Buddy, don’t even talk to me until you’ve been in the cave!”
But look around. The snobs are in full retreat. We have swung the pendulum so far toward poptimism, toward the blinkered idea that all art is equal because all humans are equal, toward the ethos that guilty pleasures are simply pleasures, that I’m not sure if we can ever swing it back.
And finally, Adam addresses the elephant:
Erasing the line between the thick and the thin has left us defenseless against slop at the exact moment of its onslaught. Everyone can sense there’s something amiss with the prose that comes out of the machines, but we lack the language to talk about it, and so we’ve converged on the idea that slop simply means using too many em dashes, bullet points, and line breaks.
What separates substance from slop is thickness. Slop holds no secrets; it signifies nothing. Under scrutiny, it evaporates. All it can offer is bottomlessness—sure, there’s nothing on, but at least there are infinite channels!
That’s why I’m neither surprised nor dismayed when studies find that people prefer AI art to human art. Of course they do! In the short term, thinness prevails. When people are making snap judgments, they want pretty flowers, poems that rhyme, pleasing pablum, the simulacrum of thought. But none of these last…
About once a week, I get a pitch from some AI startup that wants to automate some part of my writing. The most recent one says it’s “built for credible thinkers who have a book’s worth of ideas but not the time it typically takes to write one”.
I’m sorry, but if you’re building or using a tool like this, then you’ve got slop for brains. There is no such thing as having a “book’s worth of ideas” that are all ready to go except for the small matter of choosing the right words and putting them in the right order.
I know exactly the feeling that these slop-trepeneurs are preying on, because I feel it all the time: I’ve got these thoughts in my head, and boy oh boy they’re good ones, all-timers, really, and it’s so annoying that I have to spend all this time making the words sound good, when the ideas behind the words are already so good!
But this is an illusion. The ideas are not already good. They need to be thickened. I understand why it’s tempting to force a machine do the hard part for you, but it can’t, and the hard part is the only part worth doing anyway.
Money Angle
People seemed to like the post I wrote 2 weeks ago teaching readers how to compute Kelly optimal bet sizes in their heads. A lot of people reached out saying that despite learning Kelly in the past, this treatment not only made it clearer but also helped them appreciate how its approach informs risk-taking in wider contexts.
Panoptica graciously asked me to republish it under their own banner:
After this post you will be sizing bets in your head (Panoptica)
Just to put a bow on it, I’ll restate what I think are the most crucial lessons without dwelling on the formula:
Even educated people are terrible at sizing bets. It’s not because it’s so complex, but I guess it’s like squatting. It seems like you should just know how to do it, but it’s actually something you need to learn the mechanics of.
Overbetting is incinerating money. This is something that’s hard to appreciate until you see the math. The reason you size smaller is because of the asymmetry of being wrong on your edge. If you underbet, you slow your growth rate but slow risk even faster. At least you’re exchanging lower returns for a better risk/reward. But if you overbet, you lower your growth rate AND increase your risk even faster than you reduce your reward. Both the numerator and denominator of your risk/reward move in the wrong directions!
If interested, Matt does this neat meta series “Notes on Notes”. It’s a short chat about how and why a particular article comes together in the first place. You can watch it here:
To round out this Kelly sprint, I have 2 more bits that are again overtures to curious learners who may find math intimidating.
1) Slides
The first is a condensed slide version, which I hope makes this accessible. It was born out of teaching Kelly to my 8th grader at breakfast this past Tuesday. “Zak, I wanna see if I can teach you something neat in 5 minutes.” He rolled his eyes, but at least he humored me while scarfing down his cereal.
The derivation of the Kelly formula is so fun because as it rolls downhill, a number of concepts we talk about in this letter stick to it, so when we get to the end it feels like something grand, but it’s so damn compact.
Another teaching experiment. Let’s see if I can narrate the derivation in a way so that as you follow along it never feels “hard”. I want to prove that this is fun to do and while I don’t expect to convert everyone, I do think there’s a bunch of you who’d like to be able to learn this but feel blocked because you have gaps in your foundations or can’t remember HS math, or just lack some confidence.
Screw all that, I’ll lay my jacket over the puddles so you can see that it’s not so bad out here. Just come along.
Money Angle For Masochists
Deriving Kelly From Scratch
Stating the question
You have a bet. You win with probability p and lose with probability q = 1 − p. If you win, you get paid B times what you risked. If you lose, you lose what you risked. B is also just a return. So if you double your money on a bet, B=1=100% return.
You’re going to bet the same fraction f of your bankroll every time and let it ride.
What’s the f that yields the highest compounded return?
Step 1: What one bet does to your wealth
If your wealth is W and you bet the fraction f:
Win: W → W(1 + Bf) Lose: W → W(1 − f)
Example:
Your starting wealth is $100 and you bet 50% of it on a coin flip. Remember B =1 because when you win you make 100%. When you lose you always lose f which is your bet size.
Wₙ is where you end up. We want a per-bet growth rate.
You already know how to do this. If an investment grew by a factor of (1 + 8%)¹⁰ over 10 years, you take the 10th root to get the CAGR back. Same move here: take the nth root of this equation: Wₙ = W₀(1 + Bf)ʰ(1 − f)ⁿ⁻ʰ
Note that W₀ can be divided out, just as if your starting identity was 150 = 100(1+8%)¹⁰ and you turned that into 1.5 = 1.08¹⁰ before taking the 10th root to get to annual growth rate.
This equation took our simple total growth equation and turned it into a growth rate equation:
Step 4: Let probabilities take over
We can clean up our new growth equation with several handy notation substitutions.
Ps and Qs
h/n is the share of bets won
(n-h)/n is the share of bets lost
Over a long run, the fraction of bets you win settles down to your win probability:
h/n → p and (n − h)/n → q
G
Wₙ/W₀ is just a wealth multiple. If you made 200% on your investment portfolio over a decade, your wealth multiple is 3 because you started with $1 and ended up with $3.
We’ll call that wealth multiple G (for “gross multiple”)
So the per-bet growth factor becomes:
G = (1 + Bf)ᵖ(1 − f)ᑫ
G = 1.05 means your bankroll grows 5% per bet on a compounded basis. In traditional investing, we’d substitute the word “year” for “bet”. Each flip is like a year.
Our job is to find the f that makes G (the gross multiple of our wealth) as big as possible.
Step 5: Take the log to make the math easy
Maximizing G means taking its derivative with respect to f and setting it to zero.
when someone says “derivative”
It’s 2026, you don’t need to know how to actually differentiate an equation. You just have to awaken that part of your brain that knows:
a) a derivative is the slope of a function at a given point on a curve
b) when the slope of a curve is 0 this is a maximum or minimum
c) intuitively, we can reason that this is a growth curve is a hill with no bumps. Its slope starts positive and only ever gets smaller as you bet more, so it can hit zero exactly once, and when it does, you’re at the top. (the jargon version: this growth curve is concave: it bends downward everywhere from the max, with no inflection points)
We are still here:
G = (1 + Bf)ᵖ(1 − f)ᑫ
But G is a product of powers, which is nasty to differentiate.
But remember from simple pleasures, logs fix this. They are the inverse of exponentiation, allowing them to turn multiplication into addition!
They pull exponents down in front: log(xᵖ) = p · log(x).
We can prep that equation for differentiation by taking the log of both sides and using both of those sexy log features:
log(G) = p · log(1 + Bf) + q · log(1 − f)
Maximizing log(G) gives the same f as maximizing G, because log always rises when its input rises.
rapidtables.com
There’s another little bonus.
Log(G) is the log return per bet, the same quantity as ln(S₁/S₀).
It re-expresses a compounded, multiplicative growth factor as a continuously compounded rate, and rates add cleanly.
Step 6: Differentiate and set to zero
Once you know this is a derivative problem because you are trying to maximize then we can rely on crutches (Claude) for thing that’s hard to remember.
Namely that:
the derivative of log(x) is 1/x.
If there’s something inside the log, also multiply by the derivative of that inside piece (the chain rule)
So we have our equation:
log(G) = p · log(1 + Bf) + q · log(1 − f)
then we differentiate our terms to be added with respect to f:
p · log(1 + Bf) → pB/(1 + Bf) (the inside, 1 + Bf, has derivative B)
q · log(1 − f) → −q/(1 − f) (the inside, 1 − f, has derivative −1)
Set the sum equal to zero, which is where the growth curve is flat at its peak:
pB/(1 + Bf) − q/(1 − f) = 0
Step 7: Solve for f
Move the second term across and cross-multiply:
pB(1 − f) = q(1 + Bf)
pB − pBf = q + qBf
pB − q = pBf + qBf
pB − q = Bf(p + q)
But don’t turn your pattern recognition noggin off!
Since p + q = 1, this collapses to:
f = (pB − q) / B
f = p − q / B
Can you reproduce this right now on a blank sheet of paper?* Probably not. But go through it once the way you used to trace when you learned the motions for drawing comic book characters or flowers or in my case TMNT. After a single reproduction by hand, I assure you some neural pathways will re-open that have had construction signs in front of them for years.
And to tie this back to the beginning, that is the joy of thickness.
[Get your mind back over here, this is a family letter.]
*You will be able to reproduce it on your own after 2 or 3 attempts. I don’t know why the motion of the pencil is a 10x better instructor than reading it a bunch of times (generation effect maybe?). But this is one of those learning principles I take seriously and impress on my kids.
The world will seduce you with ease where you’d be better served by friction. It is a 21st-century skill to have a point of view on the difference.