Quantamental Research·New York, NY·--:--:--
← research
driver · crowding2026-09-16·24 min read·Anthony Huang

Factor or Crowd? Measuring AI Exposure and Crowding in US Equities, 2016–2026

Is AI a factor or a crowd? Exposure measured point-in-time from 22,139 10-K filings says: neither, exactly. The AI long-short is largely a style bundle when priced against the published factors — but those factors now contain the AI names, and against styles rebuilt without them the alpha returns. The median AI firm is not expensive; the cap-weighted AI book is. And since spring 2026 the most AI-intensive names have started to trade as a bloc.

Executive summary

  • 01Disclosure diffused, then changed character. AI mentions went from 2.8% of filings in 2015 to 91.2% in 2026 — but the Business section's share of those mentions fell from 84.6% to 29.4%. By 2026, 68.3% of filers discuss AI in Risk Factors against 37.9% in Business. Most of the diffusion is boilerplate about somebody else's AI.
  • 02The published factors have partly become the AI trade, so "AI is just old styles" is partly circular. Against FF5+momentum the long-short's alpha is +3.5%/yr (t = 1.34) with R² 54.3%. Against style factors rebuilt without AI-leg stocks, alpha is +5.6%/yr (t = 1.82) and R² falls to 22.5%. Sorting on Business-section language only, post-ChatGPT alpha is +10.8%/yr (t = 2.46).
  • 03The cohort's market-cap gain was not mostly re-rating. With TTM fundamentals and margins separated, the Nov-2022 AI cohort's +73.6 log points split into +31 sales, +18.5 margin and +24.1 multiple. Against the low-AI leg — which re-rated more (+28) — the AI cohort's excess gain is +14.7 sales and +25.9 margin against −3.9 multiple.
  • 04Crowded on concentration and comovement, not on the median multiple. The industry-adjusted median valuation spread is −0.02 (23rd percentile) while the float-weighted spread is +0.99 (82nd). Factor volatility is at its sample high (28.3%), and since spring 2026 the top-decile core comoves beyond size- and tech-matched peers (+0.057, 99th percentile) — exploratory, but it strengthens under every control I could think to impose.
  • 05Nothing predicts, and the one pattern that looked like it did was an artefact. Run-up reversal (−6.47pp per σ at 12 months, HAC t = −2.76) has a bootstrap p of 0.225 once the null reproduces Stambaugh bias; no cell survives Romano-Wolf (min p = 0.34).
  • 06History prices the tail, not the mean. Under the Greenwood-Shleifer-You design — a 40% drawdown from the running peak — two-year doublings crash 21% of the time (CI 1232) against a volatility-matched base rate of 11%, and 42% when the run-up also beats the market by 100%. Computers are inside a live episode; semiconductors peaked just under the bar.

This is the second draft. Referee comments changed four headline numbers, and the text says so at each point rather than quietly restating them: the re-rating share of the AI cohort's gain (§05), the spanning verdict (§04), the crash frequency (§08) and the run-up reversal (§07). Where a result moved, the first-draft number is shown alongside the corrected one.

Two hypotheses, stated so they can fail

A factor, in the sense of Fama and French (2015), needs two things: stocks sorted on the characteristic must share common return variation, and the long-short must either earn a premium or be spanned by factors that do. A crowd is harder. Stein (2009) framed the problem — when many investors hold the same position, none can see the others, and the price impact of a joint exit is in nobody's risk model — and Lou and Polk (2022) made it measurable: crowding leaves a fingerprint in excess return correlation among the stocks the crowd holds, beyond what common factors explain. Brown, Howard and Lundblad (2022) link that fingerprint directly to tail outcomes, which is where §08 ends up. Barberis, Shleifer and Wurgler (2005) show the same signature appearing mechanically on index inclusion, which is why every comovement number below is measured against a matched benchmark rather than against zero.

There is a trap in testing the first hypothesis in 2026, and the first draft of this piece fell into it. If the AI names are a fifth of the market and most of the big-growth corner of the value and investment factors, then "the factors span AI" and "AI became the factors" are the same regression. §04 separates them by rebuilding the styles without AI-leg stocks. The second trap is time: the AI trade is about four years old, which is not enough history to learn how crowded trades end. §08 borrows a century of survivorship-free industry data for that, following Greenwood, Shleifer and You (2019).

Measuring AI exposure, and separating talk from business

The universe is every NYSE- and Nasdaq-listed 10-K filer that ranked in the top 1,300 by dollar public float in any year from 2014 to 2025 (2,084 firms). For each I downloaded every 10-K primary document filed between January 2015 and August 2026 (22,139 filings, median 63,086 words), stripped the HTML and inline-XBRL header, and counted a fixed dictionary: seven phrases (artificial intelligence, machine learning, deep learning, neural network, large language model, natural language processing, computer vision) plus the bare tokens AI, GenAI and LLM, matched case-sensitively. Exposure is mentions per 10,000 words, dated to the filing date:

AIi,t=104×kcountk ⁣(10-Ki,τ)words ⁣(10-Ki,τ),τ=max{filing datet}\mathrm{AI}_{i,t}=10^{4}\times\frac{\sum_{k}\operatorname{count}_k\!\left(\text{10-K}_{i,\tau}\right)}{\operatorname{words}\!\left(\text{10-K}_{i,\tau}\right)},\qquad \tau=\max\{\text{filing date}\le t\}(1)

The referee's cheapest suggestion turned out to be the most valuable: a mention in Item 1A (Risk Factors) — “our competitors may deploy AI” — is not the same economic signal as a mention in Item 1 (Business). I re-read all 22,139 filings and split them at the section headings, which parse cleanly in 86.8% of documents, then counted each section separately. Figure 1 is the result, and it reframes the adoption curve. Mentions went from 2.8% of filings in 2015 to 91.2% in 2026, but the Business section's share of all AI words fell from 84.6% to 29.4%. In 2026, 68.3% of filers mention AI in Risk Factors against 37.9% in Business. The diffusion everyone cites is mostly firms writing about somebody else's AI. The vocabulary converged too: the bare “AI” token was 49.6% of dictionary hits in 2016 and 71.5% in 2025, while “machine learning” fell from 17.3% to 8.1%.

AI disclosure: how much, and in which section
any AI mentionin Risk Factors (Item 1A)in Business (Item 1)Business share of AI words
201620182020202220242026020406080304050607080% of filings mentioningBusiness share of AI words (%)
Source: SEC EDGAR (10-K full text, submissions, XBRL frames); author's calculations. Primary 10-K documents only; 1,926 filings in 2026 through August. Section split parses in 86.8% of filings.
Filing year10-Ksany AIin Businessin Risk FactorsGenAI/LLMmean intensityBusinessRiskBusiness share of words
20151,5432.8%1.4%0.3%0.3%0.020.070.0178.2%
20161,6053.1%1.9%0.5%0.4%0.020.100.0184.6%
20171,6576.3%4.1%1.4%0.4%0.060.310.0380.7%
20181,71710.9%7.1%2.4%0.3%0.110.550.0681%
20191,76815%9.7%3.7%0.3%0.180.960.1181.8%
20201,82918.4%12.4%5.4%0.3%0.241.260.1578.8%
20211,91822.2%15.6%6.8%0.3%0.341.530.2865.8%
20221,98826.6%18.9%8.7%0.3%0.391.970.3170%
20232,00730.9%19.7%13.7%3.1%0.542.420.6360.1%
20242,03160.3%25.2%41.3%21.1%1.414.263.2736.2%
20252,06578.7%31.4%57%32.7%2.316.415.9632.2%
2026 (Jan–Aug)1,92691.2%37.9%68.3%42.8%3.679.339.7929.4%

Table 1. AI disclosure by filing year. Intensity = dictionary mentions per 10,000 words (eq. 1); Business and Risk columns are the same measure computed inside Item 1 and Item 1A. First post-ChatGPT filing season shaded. Source: SEC EDGAR (10-K full text, submissions, XBRL frames); author's calculations.

Two limits of the measure, stated here rather than in the appendix. It updates once a year, in a period when the language moved quarter to quarter — earnings-call text would be timelier. And a dictionary cannot tell selling AI from buying it: a semiconductor firm and a retailer deploying chatbots can score alike. The section split is a partial fix, not a complete one. Table 2 shows the latest cross-section passes the face-validity check either way.

Largest AI-leg namesAI / 10kuniv. wt
NVDA · Electronic Equipment (semis)34.16.8%
GOOGL · Computer Software23.55.6%
MSFT · Computer Software32.55.1%
AMZN · Retail10.23.8%
AVGO · Electronic Equipment (semis)14.62.6%
META · Computer Software19.21.8%
LLY · Pharma5.31.7%
TSLA · Autos9.81.3%
MU · Electronic Equipment (semis)10.51.1%
V · Business Services11.80.9%
Most AI-intensive filersAI / 10kfloat $bn
EXLS · Business Services41.25
ESTC · Computer Software34.17
NVDA · Electronic Equipment (semis)34.14,628
MSFT · Computer Software32.53,459
ADBE · Computer Software32.587
IREN · Banks30.77
CRWV · Computer Software29.417
FROG · Computer Software29.18
NOW · Computer Software27.895
DOCN · Computer Software26.57

Table 2. Latest cross-section (2026-07). Industries are Fama-French 49 from SEC SIC codes, which are sticky and self-reported (bitcoin miners turned AI hosts still file as finance companies); Hoberg-Phillips text-based industries would be the better classifier. Apple is absent from the AI leg: its 10-K intensity falls below the 5.3-per-10k cutoff, a reminder that disclosure volume is not exposure. Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices; author's calculations.

Did disclosure predict the ChatGPT repricing? It depends what you control for

Take each firm's exposure from 10-Ks filed before 30 November 2022, cumulate abnormal returns from December 2022 through June 2023, and regress the cross-section on standardized exposure with size controls, clustering by industry. The first draft ran this with six-factor abnormal returns and industry fixed effects, found nothing, and called it a null. Both choices were wrong, and reversing them reverses the answer.

CARi=a+bz ⁣(log(1+AIi,Nov-22))+clogFVi+γind(i)+εi\mathrm{CAR}_{i}=a+b\,z\!\left(\log(1+\mathrm{AI}_{i,\,\text{Nov-22}})\right)+c\log \mathrm{FV}_i+\gamma_{\mathrm{ind}(i)}+\varepsilon_i(2)

On market-adjusted returns, a 1σ increase in pre-event exposure is worth +3.44pp (t = 2.78); on raw returns, +3.65pp (t = 2.85). On six-factor-adjusted returns it is −0.27pp (t = −0.3). The six-factor adjustment removes the result because in the first half of 2023 the growth and investment factors were the AI repricing — the same circularity §04 confronts. Industry fixed effects do the same thing for a different reason: the repricing hit whole industries (semis, software), so absorbing industry means absorbs the effect. The placebo window is flat (−0.57pp, t = −0.65).

Three further cuts keep the claim honest. Splitting the margins, neither the mention dummy (−0.03pp, t = −0.02) nor intensity among mentioners (+0.52pp, t = 0.3) is significant on six-factor returns. Business-section exposure does no better than total exposure (−0.29pp, t = −0.3), so the section split, valuable in §04, does not rescue this test. And because “the winners were a handful of names” is a claim about the tail rather than the mean, I test the tail directly: the odds of landing in the top 5% of outcomes rise 1.2× per σ of exposure (p = 0.167), and the 90th-percentile quantile slope is +1.87pp (t = 1.84) against +0.61pp at the median. Directionally right, statistically marginal.

First draft: “pre-ChatGPT disclosure did not predict the repricing (−0.3pp per σ, t = −0.3).” Corrected: disclosure predicted raw and market-adjusted abnormal returns (+3.44pp per σ, t = 2.78); it does not survive six-factor adjustment or industry fixed effects, and the honest reading is that those controls absorb the event rather than that the signal is empty. The safest statement remains the narrow one: pre-ChatGPT disclosure intensity was a weak ex-ante signal, and a text screen built in 2022 would have been a blunt instrument — as Eisfeldt, Schubert and Zhang (2023) find with labor-based exposure, what “AI exposure” means depends entirely on the instrument.

Abnormal return, Dec-2022 → Jun-2023, by pre-ChatGPT AI exposure
no mention1.5%[−3, 0.2] · 697Q1 (0.1–0.2)3.5%[−10.3, 3.3] · 59Q2 (0.2–0.3)+2.2%[−3, 7.2] · 58Q3 (0.3–0.8)6.7%[−12.2, −1.3] · 58Q4 (0.8–1.5)+1.7%[−4.7, 8.7] · 58Q5 (1.5–10.4)2.3%[−8.9, 4.4] · 59
Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices, Kenneth R. French Data Library (CRSP-based); author's calculations. Six-factor betas from 156 weekly returns before the event; bars are six-factor CARs, the conservative version. Tags: bootstrap 95% CI · n.
Specification (dependent = CAR, Dec-22 → Jun-23)b (pp per σ)tn
Market-adjusted, no industry FE+3.442.782.6%988
Raw return, no industry FE+3.652.852.5%988
Six-factor-adjusted, no industry FE+0.360.370.2%988
Six-factor-adjusted, industry FE0.27−0.3016.7%988
— value-weighted (WLS by float)+0.660.7925.4%988
— Business-section exposure only0.29−0.300.1%988
— mention dummy (extensive margin)0.03−0.020.1%988
Placebo May-22 → Nov-22 (six-factor)0.57−0.6514.6%988
Extended window to Dec-24 (six-factor)+1.070.6410%988

Table 3. Cross-sectional regressions of abnormal returns on standardized pre-event exposure (eq. 2), controlling for log float value; standard errors clustered by FF49 industry. Shaded rows are the lead specifications: with the AI names inside the factors, factor-adjusting the dependent variable removes the event being measured. Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices, Kenneth R. French Data Library (CRSP-based); author's calculations.

Is AI a factor? Not against the published ones — but they are no longer independent

The sort is explicit, because the tie structure matters enormously early in the sample:

Ht={i:AIi,t>Q80,t  AIi,t>0},Lt={i:AIi,tQ30,t}H_t=\{i:\mathrm{AI}_{i,t}>Q_{80,t}\ \wedge\ \mathrm{AI}_{i,t}>0\},\qquad L_t=\{i:\mathrm{AI}_{i,t}\le Q_{30,t}\}(3)

When 97% of firms score zero, Q80,t=0Q_{80,t}=0 and the “top quintile” is every firm with any mention, while the “bottom 30%” is every firm with none. The legs therefore drift: 25 long and 975 short in January 2016, 127/873 in January 2019, 200/706 at ChatGPT, and 200/300 today. Early on this is “mentioners minus the market”; today it is a genuine intensity sort. Any pre/post comparison inherits that non-stationarity, which is one more reason the break tests below matter. Both legs are float-value-weighted and held one month. Four variants guard the obvious objections: equal weighting (kills the megacap bet), within-industry sorting (kills the sector bet), a top-decile core, and a sort on Business-section intensity only.

Growth of $1: AI long-short portfolios vs the market, 2016–2026
AI − low-AI (value-weighted)Business-section sortwithin-industrymarket (CRSP VW)
2016201820202022202420261.002.003.004.00growth of $1
Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices, Kenneth R. French Data Library (CRSP-based); author's calculations. Long-short returns exclude financing; market = CRSP value-weighted total return.

Table 4 reports performance. The headline portfolio earned +5.9% a year (t = 1.44), +11.9% after ChatGPT with volatility rising from 9.7% to 18.4%. The best-behaved variant is the one the section split made possible: sorting on Business-section language alone earns +16.6% a year post-ChatGPT at 16.4% volatility (Sharpe 1.01, maximum drawdown −12.7%). Firms that describe AI in their business beat firms that merely disclose it. The equal-weighted version is the counterweight: +3.1% after ChatGPT. Its pre-ChatGPT alpha is the only conventionally significant one in the FF6 table (+5.4%/yr, t = 2.5), consistent with Babina, Fedyk, He and Hodson (2024) on AI investment and firm growth in the pre-generative era — a result the first draft passed over in silence.

Portfoliomean, fullvolSharpetmax DDmean, premean, postSharpe post
AI − low-AI, value-weighted+5.9%13.3%0.441.4431.2%+2.7%+11.9%0.65
AI − low-AI, business-section text only+7.4%12.9%0.581.8731.0%+2.5%+16.6%1.01
AI − low-AI, equal-weighted+4.5%9.8%0.461.4923.9%+5.2%+3.1%0.23
AI − low-AI, within-industry+3.3%6.3%0.531.7219.9%+0.9%+7.8%1.04
Top-decile AI − low-AI, value-weighted+7.9%14.9%0.531.7332.8%+2.8%+17.5%0.85
Market excess (FF)+13.0%15.7%0.832.6925.3%+12.3%+14.3%1.05

Table 4. Monthly long-short returns, annualized; 2016-012026-07 (n = 127); pre = through Nov-2022 (n = 83), post = Dec-2022 onward (n = 44). Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices, Kenneth R. French Data Library (CRSP-based); author's calculations.

Now the spanning question, and the correction that matters most. Against the published FF5 plus momentum, six factors explain 54.3% of the long-short's monthly return variation, rising to 70.4% after ChatGPT, with an insignificant alpha of +3.5%/yr (t = 1.34). The first draft read that as “AI is old styles in new clothes.” But by 2026 the AI leg is 44.4% of universe float value, so the published factors are not independent of the thing being tested. Rebuilding four style factors from non-AI stocks only — market, size, a price-to-sales value factor and momentum, each computed inside the non-AI universe — the picture changes: R² falls to 22.5% and alpha rises to +5.6%/yr (t = 1.82), +10.6% post-ChatGPT (t = 1.79). For the top-decile core it is +7.8% (t = 2.16) and +16.2% post (t = 2.33).

Two further tests keep me from over-claiming in the other direction. Freezing the pre-ChatGPT loadings and applying them to the post period leaves an out-of-sample alpha of +3%/yr (t = 0.49) — the frozen factor model prices the post period about as well as the fitted one, which argues against a pure “AI became the factors” story. And the loadings themselves did break: a joint Wald test on alpha and all six betas rejects stability (χ² = 19.29, 7 df, p = 0.007), even though the alpha shift alone is −0.7pp (t = -0.12) with a minimum detectable effect of 15.3pp — that test could not have found a break of any plausible size. So the correct sentence is “no detectable break in alpha, and a clear one in loadings,” not “no structural break.”

Portfolio · periodα %/yrt(α)MKTSMBHMLRMWCMAUMD
AI − low (VW) full+3.51.34+0.24−0.34−0.26−0.31−0.51−0.0754.3%
pre+3.61.23+0.13−0.17−0.13−0.16−0.51−0.0248.5%
post+2.90.62+0.41−0.26−0.71−0.22−0.47−0.0670.4%
Business-section sort full+4.71.87+0.24−0.20−0.13−0.22−0.65+0.0350.6%
pre+3.21.15+0.18−0.15+0.03−0.18−0.71+0.0847.9%
post+10.82.46+0.29+0.10−0.77−0.09−0.44+0.0865.4%
AI − low (EW) full+3.51.58+0.15−0.22−0.28−0.33−0.23−0.0954.5%
pre+5.42.50+0.10−0.03−0.26−0.13−0.21−0.0258%
post3.1−0.66+0.19−0.41−0.29−0.49−0.32−0.1861.9%
Within-industry full+2.11.15+0.11−0.11−0.10−0.09−0.07−0.0521.2%
pre0.1−0.05+0.05+0.08−0.15+0.07+0.09−0.0611.7%
post+2.81.28+0.19−0.22−0.12−0.07−0.28−0.0258.9%
Top decile − low full+5.31.78+0.24−0.46−0.33−0.28−0.52−0.0753.9%
pre+3.71.11+0.10−0.23−0.20−0.11−0.48−0.0347.5%
post+7.61.56+0.45−0.37−0.86−0.13−0.49−0.0271.8%
vs ex-AI styles full+5.61.82+0.03+0.58−0.64−0.04ex-AI factors22.5%
pre+2.50.92+0.07+0.25−0.56−0.13ex-AI factors26.2%
post+10.61.79−0.16+1.12−0.77+0.12ex-AI factors26.1%
vs ex-AI styles (top decile) full+7.82.16+0.01+0.48−0.74−0.07ex-AI factors22.8%
pre+2.90.94+0.05+0.14−0.65−0.18ex-AI factors29%
post+16.22.33−0.19+1.01−0.88+0.14ex-AI factors24.9%

Table 5. Spanning regressions, monthly, Newey-West HAC (lag 3). Top block: published Fama-French five plus momentum. Bottom block: MKTx/SMBx/VALx/MOMx rebuilt from non-AI-leg stocks inside this universe (columns reuse the MKT/SMB/HML/RMW positions). Bright = |t| ≥ 2. Source: Kenneth R. French Data Library (CRSP-based), SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices; author's calculations.

One loading deserves explanation rather than a shrug. The AI leg loads negatively on RMW (−0.31), which looks absurd for a basket containing some of the most profitable firms in the world. Two reasons: Fama-French operating profitability expenses R&D, so research-heavy firms score poorly on it by construction; and the short leg is full of stable, cash-generative, low-growth firms that score well. The CMA loading (−0.51, t = −2.6) says the leg is tilted toward aggressive investment — asset growth broadly, which includes acquisitions and working capital, not only the data-center capex it is tempting to name.

First draft: “six factors explain 70% of the AI long-short, so AI is a bundle of old styles, with no significant alpha.” Corrected: the published factors explain 70.4% of its monthly return variation post-ChatGPT — which is not the same as “70% of the portfolio is old styles” — and that number is partly circular, because the AI names are now inside those factors. Against ex-AI styles the alpha is +5.6%/yr (t = 1.82). The defensible claim is “insufficient evidence for a distinct, priced AI premium in 127 months,” not “AI is not a factor.”

Hedge the styles you can name before calling anything AI alpha — but do not assume the standard factor suite is a clean hedge, because a fifth of the market's cap now sits on the AI side of it. The Business-section sort is the version worth tracking: same idea, better signal-to-noise (+16.6%/yr post-ChatGPT, Sharpe 1.01).

What the AI cohort's market-cap gain was made of

This section decomposes the change in aggregate market capitalization of a fixed cohort. That is not a shareholder return: issuance, buybacks and acquisitions move it too, so the numbers below describe where the market value went, not what an investor earned. With that label fixed, two measurement problems from the first draft remain to be corrected. Revenue was the latest annual figure, up to 18 months stale against a current price — which mechanically reclassifies fundamental growth as re-rating, and does so hardest for the fastest-growing cohort. And price-to-sales hides margin expansion, which is a fundamental. So: TTM revenue and TTM operating income from quarterly XBRL (median staleness now 32 days, coverage 82% of universe firms), and a three-way split:

ΔlogiMVi=ΔlogiSalesisales growth+ΔlogiOpInciiSalesimargin+ΔlogiMViiOpIncimultiple\Delta\log \textstyle\sum_i \mathrm{MV}_i=\underbrace{\Delta\log \textstyle\sum_i \mathrm{Sales}_i}_{\text{sales growth}}+\underbrace{\Delta\log\frac{\sum_i \mathrm{OpInc}_i}{\sum_i \mathrm{Sales}_i}}_{\text{margin}}+\underbrace{\Delta\log\frac{\sum_i \mathrm{MV}_i}{\sum_i \mathrm{OpInc}_i}}_{\text{multiple}}(4)

The result reverses the first draft's headline. The AI cohort's market value rose +73.6 log points (firm-bootstrap CI [45.1, 99]): +31 from sales, +18.5 from margin (CI [−6.6, 46.8]) and +24.1 from a higher multiple on operating income (CI [1.2, 50.4]). Aggregate operating margin went from 15% to 18.1%. The low-AI cohort rose +36.9, and — the number that settles it — its multiple expanded more (+28) while its margin fell (−7.4). Relative to the low-AI leg, the AI cohort's excess +36.7 log points decompose into +14.7 sales, +25.9 margin and −3.9 multiple. On this measure the AI cohort's outperformance was earned by fundamentals, not by re-rating.

Two caveats keep the reversal from being over-read. The intervals are wide — the multiple term's CI spans [1.2, 50.4] — because a cohort aggregate is dominated by a few firms: five names account for 81% of the AI cohort's dollar gain (CI [58, 93]) against 53% for the low-AI leg. And a shift-share split of the aggregate price-to-sales ratio attributes 74% of its change to within-firm re-rating and 11% to mix shift toward high-multiple names, so the aggregate is not merely a composition artefact.

Cohort (fixed at Nov-2022)firmsΔ log MVsalesmarginmultipleop margin then → nowP/SP/OpInctop-5 share of $ gain
AI leg115+73.6+31+18.5+24.115% → 18.1%3.39× → 5.2×22.6× → 28.7×81% [58, 93]
Low-AI leg356+36.9+16.3−7.4+2812.9% → 12%1.98× → 2.44×15.3× → 20.3×53% [23, 73]
AI minus low-AI+36.7+14.7+25.9−3.9

Table 6. Aggregate log decomposition (eq. 4), log points ×100, Nov-2022 → 2026-07; cohort fixed at Nov-2022, firms with valid market value, TTM revenue and positive TTM operating income at both ends. Market value = unadjusted price × cover-page shares. Intervals: firm bootstrap (2,000 resamples). This is a market-capitalization decomposition, not a return decomposition: it excludes dividends and is affected by issuance and buybacks. Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices; author's calculations.

First draft: “the AI cohort's value rose 68.8 log points: 28.2 from sales and 40.6 from a higher price-to-sales multiple — 59% re-rating.” That used the latest annual revenue against a current price and a multiple that hides margins. On TTM fundamentals the same price-to-sales framing would now read +42.6 of +73.6; separating margin puts the true multiple term at +24.1, and the AI cohort's excess gain over the low-AI leg becomes mostly margin and sales. The first draft's most quotable sentence was wrong.

Is it a crowd? Concentration, comovement — and which valuation you pick

Four gauges, each against an explicit benchmark. Comovement follows Lou and Polk: at each month-end, strip six factors from 52 weeks of returns and average the pairwise residual correlations inside a group, using only pairs from different industries so industry news cannot masquerade as crowding, and differencing against a benchmark matched on float-cap decile:

CoAIt=ρ(ei,ej)i,jCtind(i)ind(j)ρ(ei,ej)i,jBtind(i)ind(j),ei=rirfβ^if\mathrm{CoAI}_t=\overline{\rho}\big(e_i,e_j\big)_{\substack{i,j\in C_t\\ \mathrm{ind}(i)\ne \mathrm{ind}(j)}}-\overline{\rho}\big(e_i,e_j\big)_{\substack{i,j\in B_t\\ \mathrm{ind}(i)\ne \mathrm{ind}(j)}},\qquad e_i=r_i-r_f-\hat\beta_i' f(5)

Start with valuation, where the first draft contradicted itself. The median AI firm's industry-adjusted price-to-sales spread is −0.02 log points (23rd percentile) — unremarkable. The float-weighted spread is +0.99 (82nd percentile) and the aggregate spread +0.94 (84th). Both are right; they answer different questions. Since the portfolio being tested is cap-weighted, the cap-weighted number is the relevant one for crowding, and the sharper sentence is: the typical AI firm is not expensive relative to its industry; the AI book is, because the expense is concentrated in the megacap core. That also reconciles §05, where the aggregate multiple rose while the median firm did not re-rate.

Concentration needs its own foil, because US megacap concentration is not by itself an AI fact. The AI leg holds 44.4% of universe float value against 24.7% before ChatGPT, and the ten largest AI names hold 30.8%. But the ten largest firms in the universe — AI or not — hold 37.1%, so the AI-specific increment is the difference, not the level: essentially, most of the market's top 10 are the AI leg. The effective number of names in the AI leg (1/HHI) is 14, against 15 pre-ChatGPT. Read honestly, this gauge says the AI trade inherits the market's concentration rather than creating it — the direct test is holdings overlap from 13F filings, which is the pre-specified next step.

Gauge2026-071y agopre-ChatGPT avgz (real-time)percentile
Excess residual comovement (AI vs size-matched)−0.002−0.0070.0020.0552nd
AI leg, cross-industry pairs0.0150.0140.01374th
size-matched non-AI benchmark0.0170.0210.01159th
Excess comovement, market-model residuals−0.069−0.048−0.0141.2812th
Excess comovement, top-decile AI core0.0410.0000.005+3.3299th
top-decile core, cross-industry pairs0.0560.0170.015100th
vs size- AND tech-matched benchmark0.0570.0020.005+4.0599th
tech-factor residuals, tech-matched0.0640.0060.005+4.78100th
Valuation spread, median industry-adj. log P/S−0.02−0.010.070.5523rd
Valuation spread, median raw log P/S0.400.590.490.5836th
Valuation spread, float-weighted log P/S0.991.200.53+1.1882nd
Valuation spread, aggregate log P/S0.940.990.69+1.1284th
AI-leg share of universe float-cap44.4%43.0%24.7%+1.1795th
Top-10 AI names' share of universe30.8%29.3%14.8%+1.7398th
top-10 universe names (any exposure)37.1%36.1%22.9%+1.8595th
Effective number of names in the AI leg1413150.2529th
AI factor trailing 24m return8%27%14%0.3232nd
AI factor volatility (26w, ann.)28.3%20.1%9.8%+3.49100th
Crowding composite, mean of four z's0.060.270.0556th
re-standardized (real-time z)0.280.80−1.7369th
Composite with float-weighted valuation leg0.191.07−1.4448th

Table 7. Crowding dashboard as of 2026-07. ρ = mean pairwise correlation of weekly six-factor residuals, 52-week window, cross-industry pairs only. z is expanding-window (real-time, so it uses only data through each month); the percentile is against the full 2016–2026 history — the two conventions differ, which is why a 95th-percentile reading can carry z ≈ 1.2 for a trending series. Valuation coverage 77% of universe firms. Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices, Kenneth R. French Data Library (CRSP-based); author's calculations.

Now the gauge at a genuine extreme. On the broad AI leg there is no excess comovement (−0.002, 52nd percentile): 200 names spanning software to retail do not trade as a bloc. The top-decile core does. Its cross-industry residual correlation is 0.056 against 0.015 for size-matched peers, an excess of +0.041 — the sample high. The referee's objection is the obvious one: the core is 92% tech (48 Computer Software, 20 Business Services, 12 Electronic Equipment (semis), 9 Computers), and software, IT services, semis and hardware are economic neighbours, so a common AI-disruption headline would produce this signature with no crowded ownership at all. Three tests, and the result survives all three. Matching the benchmark on tech membership as well as size raises the excess to +0.057. Adding a tech-sector factor to the residualization raises it to +0.064. And the Forbes-Rigobon adjustment for correlations estimated in a high-volatility window leaves 0.054.

The honest qualifier is about the choice of cutoff, not the estimate. The top decile was chosen after the broad leg showed nothing, so its bootstrap interval understates the real uncertainty. Table 8 therefore reports a pre-specified grid: the excess is +0.090 at the top 5% (t = 4.73), +0.056 at 10% (t = 3.45), +0.022 at 15% and +0.004 at 20% (t = 0.38). A monotone gradient in exposure concentration, not a knife-edge at one cutoff, and the maximum t across the grid is 4.73 at the top 5%. Over the same six months the core rose +13.3% value-weighted, so this is a joint rally, not a joint liquidation.

Core cutoff (top % by exposure)namesexcess ρ, size-matched95% CItexcess ρ, size+tech-matchedt
top 5%50+0.090[0.052, 0.129]4.73+0.0955.14
top 10%100+0.056[0.027, 0.09]3.45+0.0634.29
top 15%149+0.022[−0.005, 0.049]1.58+0.0282.22
top 20%199+0.004[−0.015, 0.022]0.38+0.0080.84

Table 8. Excess cross-industry residual comovement of the AI core at four pre-specified cutoffs, latest 52-week window. CIs and t's from a 4-week block bootstrap over weeks (400 resamples), which covers sampling noise inside the window but not the choice of cutoff — hence the grid. Benchmarks are drawn from non-core stocks matched on float-cap decile, and in the last columns on tech-industry membership as well (10 draws). Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices, Kenneth R. French Data Library (CRSP-based); author's calculations.

Cross-industry residual comovement: the AI core vs its matched benchmark
top-decile AI coresize-matched benchmarkbroad AI leg (top quintile)
2016201820202022202420260.0000.0200.0400.060mean pairwise ρ
Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices, Kenneth R. French Data Library (CRSP-based); author's calculations. Mean pairwise correlation of 52-week six-factor residuals, cross-industry pairs. Benchmark = size-decile-matched non-core stocks (10 draws).
Concentration: the AI leg against the market's own top ten
AI leg (%)top-10 universe names (%)top-10 AI names (%)factor vol, 26w (%)
201620182020202220242026102030405051015202530share of universe float (%)factor vol (%)
Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices; author's calculations. Left: shares of universe float value. Right: annualized volatility of weekly AI long-short returns, trailing 26 weeks.

The composite averages four real-time z-scores — comovement, industry-adjusted median valuation, concentration and run-up. Because those legs are negatively correlated (appendix Table 13), their mean has a standard deviation of only 0.44, so the first draft's “+0.3σ” was mislabelled; re-standardized, the composite reads +0.28σ (69th percentile). Swapping the float-weighted valuation leg in gives +0.19σ (48th). Either way the composite is unremarkable, and it is the least useful object on this page: its legs disagree by construction, and the two that matter — concentration and core comovement — are clearer read separately.

The binding constraints are the cap-weighted multiple, an effective breadth of about 14 names, record factor volatility (28.3%), and a core that has started moving together. Stress tests should assume the core's correlation goes to one on the way down rather than to its 52-week average — and should not take comfort from the median AI firm's ordinary multiple, because the median firm is not what a cap-weighted sleeve owns.

Does crowding predict? Not once the null is built properly

Forward 3-, 6- and 12-month factor returns, the forward 12-month maximum drawdown and forward 12-month volatility, each regressed on each standardized gauge. Two things make the naive version misleading, and the first draft only handled one. Overlapping windows leave about n/hn/h independent observations. And — Stambaugh (1999) — a persistent regressor whose innovations correlate with contemporaneous returns biases the slope; for a trailing run-up that bias is negative, which manufactures exactly the reversal the first draft reported as its most consistent pattern. So p-values now come from a null bootstrap that reproduces both: the gauge is simulated from its own AR(1) with block-resampled innovations paired to the return innovations, under β=0\beta=0. Romano-Wolf step-down then controls family-wise error across the grid, exploiting the dependence that made Bonferroni absurdly conservative.

GaugeForward outcomeβ (pp/σ)HAC tp (null boot)p (Romano-Wolf)MDEn eff
Excess comovement (broad)fwd 3m AIX0.38−0.580.61311.8541
fwd 6m AIX0.60−0.620.6412.6920
fwd 12m AIX+0.570.480.76513.329
fwd 12m max drawdown+2.763.300.0670.6942.359
fwd 12m factor vol2.43−2.650.1660.8242.579
Excess comovement (core)fwd 3m AIX+0.390.440.67212.4441
fwd 6m AIX0.17−0.110.92614.3120
fwd 12m AIX+0.520.250.86915.849
fwd 12m max drawdown+3.342.740.1120.8243.419
fwd 12m factor vol2.91−3.100.090.7432.639
Valuation spread (VW)fwd 3m AIX0.18−0.180.85112.7141
fwd 6m AIX1.90−1.340.3310.9593.9820
fwd 12m AIX3.77−1.870.2460.8965.669
fwd 12m max drawdown4.12−4.510.0180.342.569
fwd 12m factor vol+3.342.700.1180.8243.479
Float-cap sharefwd 3m AIX0.05−0.070.95512.0141
fwd 6m AIX0.56−0.520.70913.0120
fwd 12m AIX0.50−0.270.86715.279
fwd 12m max drawdown3.37−2.730.1410.8243.469
fwd 12m factor vol+2.472.160.2580.8963.219
24m factor run-upfwd 3m AIX2.27−2.110.3670.8963.0133
fwd 6m AIX3.52−2.110.3610.8964.6716
fwd 12m AIX6.47−2.760.2250.8246.577
fwd 12m max drawdown3.23−3.060.1260.7452.967
fwd 12m factor vol+0.450.380.82913.287
Crowding compositefwd 3m AIX1.42−1.720.4030.9052.3141
fwd 6m AIX2.54−1.960.3020.8963.6320
fwd 12m AIX4.21−2.040.2970.8965.779
fwd 12m max drawdown0.30−0.250.86513.419
fwd 12m factor vol1.17−2.800.0970.8171.189

Table 9. Predictive regressions of forward AI-factor outcomes on standardized crowding gauges, 2016-012026-07. Returns and drawdowns in pp per 1σ (a drawdown coefficient < 0 = deeper drawdowns); volatility in points. HAC lag = horizon. p from 2,000 draws of a Stambaugh-style null (AR(1) gauge, block-resampled innovations paired to return innovations, β = 0); Romano-Wolf step-down across all 30 cells. MDE = 2.8 × HAC standard error. Source: SEC EDGAR (10-K full text, submissions, XBRL frames), Yahoo Finance daily prices, Kenneth R. French Data Library (CRSP-based); author's calculations.

The run-up reversal does not survive its own null. The point estimate is unchanged —−6.47pp per σ at twelve months, HAC t = -2.76, which in the first draft carried a pairs-bootstrap p of 0.02 — but against a null that reproduces Stambaugh bias its p is 0.225. A persistent trailing-return regressor produces slopes that size routinely when nothing is there. The strongest surviving single-test cell is valuation spread (vw) predicting fwd 12m max drawdown (p = 0.018), and after Romano-Wolf the smallest family-wise p in the whole grid is 0.34. With about 7 independent observations at the twelve-month horizon and detectable effects of 3.36.6pp per σ, this sample cannot answer the question. That is not a finding about crowding; it is a finding about the sample, and it is why §08 goes looking for more data instead of more p-values.

First draft: “the most consistent pattern is reversal after run-ups (−6.5pp per σ, bootstrap CI excluding zero), suggestive but not surviving Bonferroni.” Corrected: the pattern is what the null itself generates (p = 0.225). The negative sign was mostly econometrics, not crowding.

A century of run-ups: what happens after an industry doubles

The Fama-French 49 value-weighted industry portfolios run from 1926-07 to 2026-07, built from CRSP, including every firm that ever listed. The first draft got the experiment wrong in a way that mattered. Following Greenwood, Shleifer and You, a run-up is a two-year industry return above 100% together with a five-year return above 50% — the long-horizon filter is what keeps rebounds from a market-wide crash out — and a crash is a 40% drawdown from the running peak within the next two years. The first draft anchored the crash on the episode-month price instead, which only fires if an industry surrenders the entire run-up and more. An industry that rallies 60% and gives it all back is a 37% drawdown from peak and no crash at all by the entry-anchored rule; that is precisely the case bubble studies care about.

Redone properly (Table 10): 224 episodes across 56 calendar years crash 21% of the time (year-block bootstrap CI 1232). The entry-anchored rule from the first draft gives 12% on the same episodes — the definition, not the data, produced that number. The right comparison is not the 9.8% unconditional rate either: run-up industries are volatile industries, where a 40% drawdown is mechanically likelier, so Table 10 also reports a base rate computed over industry-months in the same trailing-volatility deciles, which is 11%. Against that foil, a doubling roughly doubles crash risk rather than tripling it. Requiring the run-up also to beat the market by 100% — the definition closest to a true bubble — gives 42% (CI 2259, vol-matched 14%), and a 150% two-year run-up gives 51%. Excluding 1998–2001 leaves 19%, so the dot-com cluster is not driving it; the post-1945 subsample gives 18%.

The mean is a different story, and it is the part of Greenwood-Shleifer-You that survives here cleanly: forward 24-month returns net of the market after a baseline run-up average +0.1pp (CI [−3.7, 4.3]). Only the 150% bucket tilts negative (−8.8pp, CI [−16.6, 2.5]). Doubling does not predict low average returns; it predicts a fatter left tail.

First draft: “100%+ two-year run-ups, raw and net of market, crash 26% of the time against a 7.9% base rate, 46% above 150%.” Two things were wrong. The crash was anchored on the entry price rather than the running peak, which misses exactly the episodes bubble research is about, and there was no five-year filter, so the episode set differed. And the unconditional base rate was the wrong foil, because run-up industries are volatile industries. Corrected: 21% against a volatility-matched 11%, rising to 42% when the run-up also beats the market by 100%.

Probability of a 40% drawdown from peak within two years
unconditional+10%41,157 ind-months100% 2y + 50% 5y+21%[12, 32] · 224+ also > 100% net of market+42%[22, 59] · 4850% 2y + 50% 5y+12%[7, 17] · 736150% 2y + 50% 5y+51%[30, 69] · 55100% 2y only (no 5y filter)+19%[11, 29] · 288100% 2y + 50% 5y, since 1945+18%[10, 27] · 201100% 2y + 50% 5y, ex 1998–2001+19%[10, 29] · 211
Source: Kenneth R. French Data Library (CRSP-based); author's calculations. Tags: year-block bootstrap 95% CI · episodes. Unconditional rate across all industry-months with ≥10 firms: 9.8%; volatility-matched rates in Table 10.
Run-up definitionepisodesyearsP(crash), peak-anchored95% CIvol-matched baseP(crash), entry-anchoredfwd 24m netnet 95% CI
GSY baseline: 100% 2y + 50% 5y2245621%[12, 32]11%12%+0.1%[−3.7, 4.3]
— also > 100% net of market482942%[22, 59]14%29%1.0%[−13.2, 14.1]
50% 2y + 50% 5y7368112%[7, 17]9%7%+1.7%[−1.1, 4.9]
150% 2y + 50% 5y552651%[30, 69]14%35%8.8%[−16.6, 2.5]
100% 2y only (no 5y filter)2885919%[11, 29]11%11%0.1%[−4.4, 4.9]
100% 2y + 50% 5y, since 19452014818%[10, 27]11%10%+0.3%[−3.8, 5.3]
100% 2y + 50% 5y, ex 1998–20012115319%[10, 29]10%11%+0.0%[−3.7, 4.6]
All industry-months (base rate)41,1579.8%6.7%+1.9%

Table 10. Fama-French 49 value-weighted industries, 1926-072026-07, industries with ≥10 firms. Peak-anchored crash = a 40% fall from the running peak within 24 months (Greenwood-Shleifer-You); entry-anchored = 40% below the episode-month level (the first draft's definition). Vol-matched base = crash rate among all industry-months in the same trailing-volatility deciles as the episodes. Intervals: calendar-year block bootstrap. Source: Kenneth R. French Data Library (CRSP-based); author's calculations.

With the right crash definition, the shape of a run-up starts to matter — another first-draft null that does not survive. Among the 224 baseline episodes (48 crashes), crashed episodes had larger run-ups (1.27 vs 1.13, permutation p = 0.001) and higher trailing volatility (0.24 vs 0.19, p = 0); acceleration and issuance still do not separate them. A pre-specified four-variable logit gives volatility t = 2.27 and run-up size t = 1.88, with an in-sample AUC of 0.69 and — the test that matters — 0.64 leave-one-out. Modest, but better than the coin flip the first draft reported under the entry-anchored definition.

Characteristic at episode startcrasheddid notperm. podds ratio / σlogit p (yr-clustered)
Run-up size (24m raw)+1.27+1.130.0011.620.034
Realized volatility (12m)+0.24+0.1901.680.014
Acceleration (last 12m − prior 12m)+0.08+0.120.6650.930.705
Δ log number of firms (24m)+0.01+0.030.6020.920.621
Δ log share of market cap (24m)+0.27+0.270.98410.985
log B/M vs median industry0.210.040.0540.730.149

Table 11. Baseline episodes (100% two-year + 50% five-year). Means by outcome in decimals (1.27 = 127%). Permutation p from 5,000 label shuffles. Four-variable logit (run-up, volatility, acceleration, issuance): in-sample AUC 0.69, leave-one-out 0.64. Turnover and firm age, two of the attributes Greenwood-Shleifer-You use, need stock-level CRSP data and are not replicated here. Source: Kenneth R. French Data Library (CRSP-based); author's calculations.

Where does today's AI complex sit? Semiconductors (22.6% of CRSP market value) peaked at a 156% two-year run-up within the last three years, 97% net of the market, and the June-2024 episode has since returned +88% with a peak-to-trough drawdown of −16% — a qualifying episode that did not crash. The trailing run-up has cooled to 79%. Computer Software peaked at 102% and is now +31%. The industry inside a live episode today is Computers: 142% raw, 103% net, accelerating (+70pp), dated 2026-04, 3 months in. Its fitted crash probability from the Table 11 logit is 39%, against the 21% unconditional episode rate — a number to hold loosely given an AUC of 0.64. Treat the Computers reading with care for a second reason: at 1.7% of market value it is a small portfolio whose membership depends on CRSP's SIC assignment of the largest hardware names, so it is a thinner signal than the semiconductor one it is often conflated with. The dot-com precedent is the reason to care about definitions: Software's December-1998 episode fell −51% from its peak (a crash) while ending the window only −10% lower (not a crash by the entry-anchored rule), and semiconductors' 1999 episode fell −73% from peak.

Trailing 24-month industry return today: the run-up league table
Computers+142%net +103 · pk 148Precious Metals+101%net +63 · pk 378Machinery+79%net +40 · pk 126Electronic Equipment (semis)+79%net +41 · pk 156Electrical Equipment+69%net +31 · pk 127Aircraft+61%net +23 · pk 92Banks+55%net +16 · pk 95Non-metallic Mining+47%net +8 · pk 101Trading+47%net +8 · pk 93Healthcare+44%net +6 · pk 57Agriculture+40%net +1 · pk 61Construction+40%net +1 · pk 178
Source: Kenneth R. French Data Library (CRSP-based); author's calculations. As of 2026-07. Tags: net of market · peak 24m run-up within the last 36 months.
Episodestart2y run-up5yvolpeak aftermax DD from peak24m returncrash?
Computers1995-07102%76%19%+91%15%+91%no
Computers1998-07108%366%29%+226%20%+210%no
Computers2000-08308%708%42%−12%81%81%yes
Computers2010-12102%65%23%+34%14%+15%no
Computers2026-04111%188%28%+28%2%+26% (3m, open)open
Computer Software1991-12106%255%25%+21%17%+20%no
Computer Software1996-04110%287%25%+79%13%+79%no
Computer Software1998-12125%431%33%+82%51%10%yes
Computer Software2021-08105%260%22%+2%39%10%no
Computer Software2024-12102%136%19%+25%20%+17% (19m, open)open
Electronic Equipment (semis)1993-08107%150%18%+84%7%+82%no
Electronic Equipment (semis)1997-01112%351%23%+66%23%+66%no
Electronic Equipment (semis)1999-10106%366%35%+70%73%47%yes
Electronic Equipment (semis)2011-02107%9%22%+2%22%4%no
Electronic Equipment (semis)2020-11120%258%43%+50%30%+19%no
Electronic Equipment (semis)2024-06124%374%22%+93%16%+88%no
Building Materials2024-09106%175%23%+17%19%+12% (22m, open)open
Machinery2026-06126%198%32%−21%21%−21% (1m, open)open
Electrical Equipment2026-04106%73%35%+10%20%−12% (3m, open)open
Autos2024-12155%391%48%+14%31%−13% (19m, open)open
Precious Metals2025-08110%32%38%+87%32%+28% (11m, open)open
Non-metallic Mining2026-02101%171%34%−8%17%−15% (5m, open)open

Table 12. Technology-industry episodes since 1990 (two-year filter, so live episodes appear) plus every episode still inside its 24-month window. Open episodes report the path so far. Source: Kenneth R. French Data Library (CRSP-based); author's calculations.

What this means for a portfolio

  • Screen on the Business section, not the whole filing. By 2026 91.2% of large filers mention AI, but only 29.4% of AI words sit in Item 1. The Business-section sort is the version with post-ChatGPT alpha (+10.8%/yr, t = 2.46); the all-text sort is half boilerplate.
  • Do not assume the standard factor suite hedges this. The published value and investment factors now contain the AI names; against styles rebuilt without them the long-short still earns +5.6%/yr. A “factor-neutral” AI book may be neutral to a benchmark that is itself long AI.
  • Size for breadth, not for the median multiple. Effective breadth is about 14 names, factor vol is at its sample high, and the cap-weighted valuation spread (+0.99) is nothing like the median firm's (−0.02).
  • Treat the tail as fatter than base rates, and size accordingly. A doubling raises the two-year crash probability to roughly 21% against 11% for equally volatile industries, without predictably lower average returns. That argues for lower risk budgets and explicit drawdown stress tests in the AI sleeve.
  • Watch core comovement; it is the one gauge at an extreme. Monitor it as a state variable — it survives tech-matching and a volatility adjustment — but it is young, exploratory, and not yet shown to predict anything.

One recommendation from the first draft is withdrawn. It advised buying convexity — puts and collars — on the strength of the historical crash frequency. That does not follow. A physical crash probability says nothing about whether options are cheap; high-run-up, high-volatility industries carry elevated implied volatility and downside skew, and the market may already price a 21% two-year crash probability, or more. Comparing physical and risk-neutral tail probabilities needs option data this piece does not have. What the evidence supports is a position-sizing and stress-testing conclusion, not an options trade.

Conclusion

The cleanest summary of this evidence is narrower than “AI is not a factor but is a crowd.” AI exposure behaves like a concentrated style complex: its returns load heavily on growth, size, investment and market beta, and those styles have themselves become partly synonymous with the trade, so the published factor suite can neither price it cleanly nor hedge it cleanly. There is not enough evidence for a distinct, priced AI premium in 127 months — but against styles rebuilt without AI names the residual is larger than the first draft implied, and the Business-section version of the signal is stronger still. The cohort's gains since ChatGPT came more from sales and margins than from multiple expansion, which is the opposite of what the first draft concluded from stale annual revenue. On crowding, the median AI firm is not expensive, the cap-weighted book is, breadth is about 14 names, and since spring 2026 the most AI-intensive names have begun to move together beyond what size, industry and a tech factor explain — a Lou-Polk signature, though ownership data, not return correlations, is what would prove it. Nothing here predicts the factor's returns, and the sample cannot: the one pattern that looked predictive was a Stambaugh artefact. What a century of industry data adds is the tail: doubling roughly doubles the probability of a 40% drawdown relative to equally volatile industries, and leaves the mean alone. For an allocator that is the operative fact — not because the AI trade is mispriced, but because its distribution is wider than the base rate implies, and one industry in the complex is inside a live episode as I write.

Pre-specified and timestamped by this page's publication date, so later specification changes are auditable: (i) replace the dictionary with revenue-segment exposure from XBRL and re-run §03; (ii) add holdings-based crowding from the SEC 13F data sets — ownership breadth, overlap and common institutional ownership — which is the direct measure this piece approximates with return correlations; (iii) compare the §08 physical crash frequency with option-implied tail probabilities, the missing half of the §09 argument; (iv) re-test §07 in 2028, when the sample doubles.

Data, method & limitations

Sources. SEC EDGAR: company-ticker-exchange map; submissions API; 22,139 10-K primary documents filed 2015-01 → 2026-08 (median 63,086 words), each also parsed into Item 1 and Item 1A (86.8% parse rate); XBRL frames for dei:EntityPublicFloat, dei:EntityCommonStockSharesOutstanding, and quarterly/annual us-gaap revenue, operating income and net income. Revenue tags are taken in priority order (RevenueFromContractWithCustomerExcludingAssessedTax, then Revenues, then SalesRevenueNet) rather than by maximum, so gross and net definitions are not mixed. TTM series sum four quarters with a 60-day availability lag and impute a missing fourth quarter as annual minus the first three. Yahoo Finance: daily adjusted and split-adjusted closes plus split events for 2,080 of 2,081 tickers. Kenneth R. French Data Library (202607 CRSP build): FF5 and momentum (daily, monthly), FF3 from 1926, 49 industry portfolios (daily, monthly, firm counts, average size, BE/ME), SIC definitions.

Data validation. The value-weighted return of the 1,000-stock universe correlates 0.998 with the CRSP value-weighted market (beta 0.98, tracking error 1.14%/yr, mean difference +0.18%/yr). Floats: 89.9% verified against price × shares; 251 filings rescaled for an exact 1,000× tagging error; 112 dropped. Monthly returns outside (−95%, +500%) and weekly outside (−90%, +300%) are treated as bad prints; preferred-stock ticker lines excluded.

Portfolios. Month-end sorts, next-month holding, float-value weights (cover-page float rolled forward on split-adjusted price). H and L as in eq. (3); top decile above the 90th percentile; the Business-section sort applies the same rule to Item 1 intensity; within-industry sorts inside each FF49 industry with ≥5 universe firms and aggregates by industry float. Ex-AI style factors (MKTx, SMBx, VALx, MOMx) are built from non-AI-leg stocks in the same universe: market excess, a median size split, a 30/70 price-to-sales value spread and a 12-1 momentum spread. They are proxies, not French-library replicas — there is no ex-AI RMW or CMA here, because operating profitability and asset growth are not in this data set.

Inference. Spanning: Newey-West HAC (lag 3); joint break test by Wald on alpha plus six interaction terms. Predictive tests: HAC at the horizon, p-values from 2,000 Stambaugh-style null simulations, Romano-Wolf step-down across 30 cells. Event study clusters by FF49 industry; group intervals are iid bootstraps. Core comovement: 4-week block bootstrap over weeks (400 resamples) at four pre-specified cutoffs, with a Forbes-Rigobon volatility adjustment reported alongside. Run-up statistics: calendar-year block bootstraps, because episodes cluster in time. Crowding z-scores are expanding-window; percentiles are full-sample.

excessvalIndcapHrunup
excess1.000.52−0.64−0.16
valInd0.521.00−0.840.04
capH−0.64−0.841.000.24
runup−0.160.040.241.00

Table 13. Correlation of the four composite legs (n = 104 months). The strong negative correlations are why the equal-weight mean of the four z-scores has a standard deviation of only 0.44, and why the composite is re-standardized before being quoted in σ.

Limitations. (1) Survivorship. The text universe contains only firms listed today. The 0.998 CRSP correlation shows the aggregate value-weighted index is close to the real one, but it does not bound survivorship bias in characteristic-sorted long-shorts, which distort if failed firms sat systematically on one side of the signal; equal-weighted results are the most exposed. A CRSP-linked rebuild is the fix. (2) Foreign private issuers file 20-F, so TSMC and ASML are absent. (3) SIC codes are sticky and self-reported; Hoberg-Phillips text-based industries would be better. (4) The dictionary cannot separate selling AI from using or fearing it; the section split is a partial remedy, and exposure updates only annually. (5) Valuation covers 77% of universe firms and price-to-sales is weakly comparable across banks, retailers and software even after industry adjustment. (6) The core-comovement cutoff was chosen after seeing the broad-leg result; Table 8's grid is the honest presentation. (7) Industry portfolios lack turnover and firm age. (8) Concentration measures market structure as much as AI crowding; holdings data is the right instrument. Reproducible via analysis/ai_crowding.py (first run downloads ~22,000 filings twice — once for totals, once for sections — at EDGAR's fair-access rate).

References. Babina, T., A. Fedyk, A. He & J. Hodson (2024), “Artificial Intelligence, Firm Growth, and Product Innovation,” JFE. Barberis, N., A. Shleifer & J. Wurgler (2005), “Comovement,” JFE. Brown, G., T. Howard & C. Lundblad (2022), “Crowded Trades and Tail Risk,” RFS. Carhart, M. (1997), JF. Cohen, L., C. Malloy & Q. Nguyen (2020), “Lazy Prices,” JF. Eisfeldt, A., G. Schubert & M. B. Zhang (2023), “Generative AI and Firm Values,” NBER WP 31222. Fama, E. & K. French (2015), JFE. Forbes, K. & R. Rigobon (2002), “No Contagion, Only Interdependence,” JF. Greenwood, R., A. Shleifer & Y. You (2019), “Bubbles for Fama,” JFE. Hoberg, G. & G. Phillips (2016), “Text-Based Network Industries,” JPE. Lou, D. & C. Polk (2022), “Comomentum,” RFS. Loughran, T. & B. McDonald (2011), JF. Newey, W. & K. West (1987), Econometrica. Romano, J. & M. Wolf (2005), “Stepwise Multiple Testing as Formalized Data Snooping,” Econometrica. Stambaugh, R. (1999), “Predictive Regressions,” JFE. Stein, J. (2009), JF.

This is research, not investment advice.