Quantitative Research

Does the signal survive size and sector controls?

The monthly panel version of this factor died exactly here: neutralize by size and sector and the spread dissolved. The event study and the portfolio built on it had never been put through the same filter. A finding published without passing the test that killed the previous finding is not audited — it is protected.

Four specifications, fixed before the script was run — the same discipline the monthly panel uses, for the same reason: run enough variants and one will clear the significance bar by chance. Inside each stratum the event return has the stratum mean subtracted from it, then the same buy-versus-sell Welch test runs on the pooled, demeaned set. If microcaps simply return more than large caps, that difference leaves with the mean, and only the buy-versus-sell difference within each stratum survives.

It survives all four. Neutralizing by size costs 4% of the spread. Neutralizing by sector costs nothing at all. The double control leaves the spread essentially unchanged from the unneutralized baseline. This is the opposite of what happened to the monthly panel, and it is the strongest evidence so far that the event-study result is not a size or sector sort wearing a disguise.

First, what requiring a market cap costs

Measuring size point-in-time means every event needs a shares-outstanding figure filed before it. Roughly a quarter of events do not have one. That restriction is not neutral, and it moves in the counterintuitive direction — so it belongs before the results, not in a footnote after them.

Fewer events, a smaller spread, a larger t. The discarded events have the bigger spread and the far weaker t-statistic: they are noisy companies without XBRL share counts. Requiring a market cap drops much more variance than signal. That is why the baseline on this page reads higher than the t on the event-study page — it is a different sample, not a different result. Comparing the two numbers without this table in hand leads straight to the wrong conclusion.

Inside each stratum

Diagnostic, not result. With three size terciles and twelve sectors, one of them clearing |t| = 2 by chance is exactly what you should expect — read the direction of the pattern, not any individual row. Only cells with at least 20 events on each side are shown; a cell with 30 buys and 2 sells is noise shaped like a spread.

The spread falls monotonically with size — and still clears the bar comfortably in the largest tercile. Consistent with the liquidity test: the signal is strongest where the market is least efficient, but it does not live only there.
All twelve sectors show a positive spread. The strongest are healthcare and industrials; the weakest are real estate and utilities — sectors where informed insider trading is plausibly less relevant to begin with.

Alpha, or just beta?

The diversified portfolio is long-only, so its gross return contains the drift of the whole market. Each day’s return is measured against the same day’s benchmark return over the identical one-day window, and everything is recomputed on the excess. Two benchmarks, decided in advance: SPY as the market control, and IWM (Russell 2000) as the real size control — if the book is mostly microcaps, judging it against SPY flatters it.

The excess survives both, and survives IWM better than SPY. So the result is not explained by being loaded with small caps during a decade that was kind to small caps — in these windows IWM returned less than SPY, and the portfolio still beat it by more. Position sizing uses the original bankroll as a fixed base with an additive balance; the accumulated balance is never reinvested.
Two numbers not to confuse with the other pages. The gross figure here is computed on the market-cap subset, so it is not the total shown on the diversified-portfolio page. And the benchmark percentage is the sum of one-day window returns on the days the book is actually invested — not the multi-year buy-and-hold shown elsewhere on this site. The portfolio holds for one day at a time, so the honest benchmark is the market on those same days, not being invested for seventeen straight years.

Sanity checks

Run on every export and published whether they pass or fail. The first one exists because the naive version of it failed: it compared this page’s baseline against the event-study page’s t-statistic and flagged a bug that was not there. It now verifies that the baseline reproduces the event study’s own function on the same subset, exactly.

What this does not settle. Slippage is still a flat 0.1% of price rather than a function of position size, there are no capacity limits, and the survivorship bias of the universe — SEC registrants active today — is untouched. These controls answer “the edge is not an artifact of size or sector.” They do not answer “this is deployable.”