The pitch has three ingredients, and each one is individually true.
First, information asymmetry is real. A CEO of a clinical-stage biotech knows things the market does not: how enrollment is going, what the safety committee said, how the FDA meeting felt. Insider-trading law forbids trading on material non-public information, but conviction is legal, and a large open-market buy is conviction made visible.
Second, the disclosure is public. In the US, Form 4 filings hit EDGAR within two business days. In Europe, MAR Article 19 notifications reach the regulator within three. Anyone can read them. The trade seems reproducible by construction: watch the filings, follow the buys.
Third, the examples exist. Small biotechs really do go from 5 dollars to 25 dollars on a positive readout, and sometimes insiders really did buy in the months before. Each time it happens, the chart circulates, and the lesson everyone draws is the same: the signal works, you just have to follow it.
Put together, you get the strategy that a hedge-fund friend of ours pitched to us, almost word for word: "insiders sell for a thousand reasons but only buy when they think the stock is undervalued; the alpha lives in illiquid micro-caps where institutional algos will not go."
Insiders sell for a thousand reasons but only buy when they think the stock is undervalued. The alpha lives in illiquid micro-caps where institutional algos will not go.
It is a good story: coherent, built from true components, and sturdy under casual inspection. That is exactly what makes it dangerous, because the only way to find the flaw is to run the numbers, and almost nobody who shares the screenshot ever does.
We did. What follows is the test, published in full in our public audit, losing windows included.
The proposal we tested was not a vague vibe. It was a concrete five-criterion screen, built for the US market:
Each criterion encodes a real intuition. Together they describe the archetypal dream trade: a tiny, unloved company, an insider writing a personally material check, in a sector where knowledge is concentrated, with management shrinking the float. If insider alpha exists anywhere, the story says, it must be here.
The story is testable. So we tested it.
Backtests die of silent shortcuts, so we state up front what our data can and cannot support. Of the five criteria, only three are cleanly testable with our data model:
| Criterion | Testable | Why |
|---|---|---|
| Small cap under 500M | Partial | Market cap is a current snapshot, not point-in-time. Using it is a survivorship-optimistic proxy: names that crashed into the bucket after the trade get swept in, and companies that delisted to zero are gone. The real result is worse than shown. |
| Amount over 1M | Yes | Filing amount, currency-normalised. |
| Position raised over 10% | No | We do not store insider holdings history. |
| Biotech / health | Yes | Sector tag on the company. |
| Gold / mining | No | We currently tag zero gold or mining companies. Cannot test. |
| Cannibal (buybacks) | No | No share-count history to measure a 2-3% annual shrink. |
| After-market intraday entry | No | Daily end-of-day data only. Our standard next-session entry already is the "enter next morning" rule. |
| Free-market buys only | Yes | We already exclude grants, option exercises, tax withholding, gifts and derivatives. |
So the combo we actually tested is: small-cap AND material AND biotech AND free-market, run separately per region. Two of the five criteria and one of the two "best sectors" are structurally untestable with our data, and we say so rather than pretend otherwise.
One of these limitations deserves a highlight, because it cuts in a direction most people do not expect. Using today's market cap to define "small cap" is not a neutral approximation. It is a gift to the strategy being tested. Companies that took the worst path, dilution into oblivion, delisting, bankruptcy, are simply absent from a sample keyed on current data. Whatever the backtest says, the true historical performance of the micro-cap screen is worse. Keep that in mind for everything that follows: every damning number below is the optimistic version.
The construction is deliberately boring. Each month we build an equal-weight book of every name matching the screen, hold each position for our standard window, and net out a realistic round-trip trading cost. Returns are winsorised to tame single-name blow-ups, which, note, makes the bad results below understatements, since the worst tails are capped.
We report an annualised Sharpe ratio, a bootstrap 95% confidence interval, a maximum drawdown, and a deflated Sharpe that penalises for the number of variants tried, the Bailey and Lopez de Prado correction. Nothing is cherry-picked to the best window: the test covers 2022 to 2026Q1, a stretch that contains both a brutal drawdown regime and a strong recovery.
Two regions were run side by side: the US market, which is the recipe's home turf, and our European universe. That regional split turned out to be where the most interesting result was hiding.
Start with the broadest reasonable pool on the US tape: all free-market insider buys, the ones that survive criterion five. Over the full window, that pool has an annualised Sharpe of exactly 0.00. Not a typo. The raw follow-every-insider-buy strategy on the US tape earns you volatility and nothing else. This alone contradicts the folk belief that copying US insider buys is a free lunch, and it is consistent with what we have published elsewhere about the US tape inverting the insider signal.
Now stack the friend's filters on top, one at a time, and watch what each one does:
| US pool, step by step | Sharpe (2022 to 2026Q1) |
|---|---|
| Free-market insider buys | 0.00 |
| plus biotech only | -0.53 |
| plus small-cap under 500M | -1.30 (drawdown -98%, near-total ruin) |
| plus materiality over 1M (full combo) | negative mean return, only 38 trades in 4 years |
Read that table slowly, because it is the whole article in four rows.
The biotech filter, the heart of the information-asymmetry thesis, does not concentrate the alpha. It concentrates the losses: from 0.00 to -0.53. The small-cap filter, the heart of the illiquidity thesis, is supposed to be where the edge hides from the algos. Instead it takes the Sharpe to -1.30 and the maximum drawdown to -98%. A -98% drawdown is not a risk statistic, it is an obituary: a portfolio that faithfully followed US micro-cap biotech insider buys through this window was, at its low, holding on to roughly a fiftieth of its peak value.
And the final filter, materiality over 1M, the skin-in-the-game test that should isolate the truest conviction? It shrinks the whole strategy to 38 trades in four years, with a negative mean return. Thirty-eight trades is not a strategy, it is an anecdote generator. Even if the mean had been positive, no statistical test on Earth would distinguish it from luck at that sample size.
This is the general shape of overfitting by intuition, and it is worth naming because it does not feel like overfitting from the inside. Each filter has a story. Each story is plausible. And each successive filter carves the sample down toward a corner of the market where the underlying businesses are structurally fragile, until what remains is a handful of trades whose outcome is dominated by variance. Stacking conviction filters does not distill signal. It distills concentration risk.
Here is the result that genuinely surprised us, and the reason this test earned a public audit rather than a quiet grave.
Cut the raw insider-buy pool into health versus non-health, per region, over the whole 2022 to 2026Q1 window:
| Pool | US | Europe |
|---|---|---|
| Health (biotech) | Sharpe -0.53 (drawdown -86%) | Sharpe +0.14 (drawdown -73%) |
| Non-health | Sharpe +0.23 | Sharpe -0.88 |
The biotech tilt does opposite things on the two tapes. In the US, health names drag the insider signal deep underwater while the rest of the market at least keeps its head up. In Europe, the pattern flips: the overall pool is weak, and the health subset is the relative bright spot. Same sector, same species of filing, same construction, opposite effect.
We have documented for years that the aggregate insider-buy signal inverts on the US tape, and it is one of the structural reasons our live universe is European. What this test shows is that the inversion extends into the sector dimension. It is not just that US insider buys in aggregate carry no edge; the specific corner that the folk wisdom celebrates most loudly, biotech, is where the US signal is most reliably wrong.
Why would the same sector behave so differently across the Atlantic? The honest answer is that our data shows the what much more cleanly than the why, but the structural differences are not mysterious. The US hosts the deepest pool of pre-revenue, single-asset, binary-outcome biotechs in the world, financed by continuous share issuance and covered by an entire retail ecosystem that trades the filings. Europe's listed health sector skews toward a different mix, and its insider filings emerge from a different disclosure culture. Whatever the mechanism, the empirical result is stark: importing a US-built biotech intuition into a European universe, or vice versa, is not a detail. It flips the sign.
And now the part we must say just as loudly, because honesty that only runs in one direction is marketing. The European biotech number is NOT a tradeable edge on its own. Its confidence interval straddles zero. Its deflated Sharpe is negative. It carries a -73% drawdown from sector concentration, and it leans on a single sub-window. It is a relative tilt inside a losing pool, not standalone alpha, and we do not trade it. We report it as evidence for the regional exclusion, nothing more.
For completeness: in Europe the friend's full stacked recipe rescues a weak pool to roughly flat, and the complete combo turns nominally positive, on 26 trades across 17 active months with a confidence interval spanning zero. That is noise wearing a positive sign, not an edge, and treating it as anything else would be exactly the mistake this article is about.
A backtest tells you that something fails. It takes a little more work to see why. Four mechanisms, each visible in the structure of the data, do most of the damage.
Survivorship bias curates the highlight reel. The 5-to-25 charts that keep the story alive are real. What never circulates is the denominator: for every micro-cap biotech that quintupled after insider buys, the same screen caught many others that faded, diluted, or delisted. The delisted ones do not just lose money, they vanish from casual datasets entirely, which is why we flagged above that even our own test is survivorship-optimistic. The screenshot economy runs on numerators. Backtests are how you force the denominator back into the picture, and the denominator is where this strategy dies. We wrote more broadly about this failure mode in our piece on survivorship bias in insider studies.
Dilution is the business model. A pre-revenue biotech does not fund trials out of profits; it funds them by selling stock, through secondaries, at-the-market programs and converts. That means the share count grinds upward and every holder's slice thins between catalysts. An insider's buy can be simultaneously sincere and doomed: management can genuinely believe in the science while the financing treadmill guarantees that believers are diluted along the way. Note the dark irony versus the recipe's own criterion four: the strategy asks for cannibals that shrink their share count, in a sector whose defining financial habit is the opposite.
Binary events break signal-following. Micro-cap biotech returns concentrate around a few dates: trial readouts, FDA decisions, partnership announcements. Between those dates the stock drifts and bleeds; on those dates it gaps, sometimes 50% or more, in either direction, with no chance to manage the position. A signal follower who enters after a filing is not riding a trend, they are holding a lottery ticket with a negative average payout, as the -98% drawdown attests. The insider may hold through the binary event with a cost basis and a time horizon a follower cannot replicate.
Illiquidity taxes both doors. The very feature that supposedly protects the edge, institutional absence, means thin books, wide spreads, and price impact on entry and exit. Retail flow chasing a publicised filing can itself move the price, manufacturing the short-lived pop that makes the signal look better in anecdotes than it performs in a costed backtest. Our test charges a realistic round-trip cost; real execution in these names is often worse.
None of these four mechanisms is exotic, and all of them are public knowledge. The trap is that a story built from true ingredients can still describe a losing trade, and only the boring machinery of a point-in-time backtest reveals it.
Here is the uncomfortable meta-result. This is the fourth outside conviction filter we have tested and retired, after a small-cap tilt, a conviction and position-size overlay, and a fundamental-quality screen. Four ideas, four plausible stories, four different proposers, one identical outcome: stacking intuitive filters on the insider signal does not manufacture alpha. It concentrates the book into small, high-variance, negative-drift corners of the market.
The pattern deserves a name, because you will meet it far beyond insider signals. Call it narrative compounding: each added filter makes the story better and the strategy worse. The story improves because each criterion adds texture and apparent rigour, five criteria feel more sophisticated than one. The strategy worsens because each criterion shrinks the sample, amplifies the role of luck, and, in this case, steers the book straight into the most structurally fragile issuers on the market.
What protects you is procedure, and it fits in four lines:
If one sentence survives from this article, let it be this one: the burden of proof does not scale with how good the story sounds. If anything, the relationship runs the other way.
Testing a strategy and publishing the corpse would be empty theatre if nothing changed in the product. Here is how the result is wired into what we ship.
The live universe excludes the trap by construction. Our production selection operates on a European universe. That is a standing design decision, taken and re-confirmed on evidence like the tests above: the US tape inverts the aggregate insider-buy signal, and this backtest shows the inversion extends to the very corner, micro-cap biotech, that the folk strategy celebrates. The full reasoning lives on our methodology page, and the complete numbers, including everything quoted here, are in the public audit.
Sector concentration is treated as risk, not conviction. A -73% drawdown on the European health subset, the relatively good one, is a reminder that a sector tilt is a risk position even when its Sharpe is positive. Our live selection does not chase sector stories; it scores filings on evidence that survived out-of-sample testing.
The claim stays modest, and audited. We do not claim the European biotech tilt as an edge, because its confidence interval does not support that claim. What we publish as our live track record is on the performance page, with its own out-of-sample accounting. The failed tests, this one included, stay published alongside it.
The door stays open, on evidence. Two of the friend's five criteria were untestable with our current data, and we tagged zero gold or mining companies, so the sector half of his asymmetry thesis remains untested rather than refuted. If that data ever lands in our model, we will run the test and publish it, whichever way it comes out. That is the difference between an exclusion and a prejudice: an exclusion has a reopening condition.
If you want the broader context on how small caps behave inside our European universe, where the story is different from the US micro-cap corner, we have covered it in our small-cap edge analysis. And for the specific sector dynamics around drug approvals, see insider buying before FDA approvals.
An insider buy at a tiny biotech is the most tellable trade in finance: asymmetric information, public filings, and spectacular worked examples. Tested point-in-time over 2022 to 2026Q1, net of costs, the recipe fails everywhere it is supposed to shine. On the US tape every conviction filter made the pool worse, bottoming at a Sharpe of -1.30 with a -98% drawdown in the micro-cap biotech cell, the worst intersection we have ever measured, and the test was rigged in the strategy's favour by survivorship. The sector tilt turned out to be regionally inverted, damaging in the US and mildly helpful in relative terms in Europe, which is one more reason our live universe is European and one more entry in a growing file: conviction filters that improve the story and destroy the returns. A good story is not an edge. Test before you believe, and publish what the test says, especially when it says no.
No. It means the aggregate follow-everything approach earns nothing on the US tape, and less than nothing in its most celebrated corner, micro-cap biotech. Signal quality is heterogeneous: it depends on region, filing type and context, which is why our live selection is built on a European universe with evidence-tested criteria rather than on a folk screen. The live results, good and bad, are on the performance page.
Because a negative-expectation pool still produces spectacular individual winners, and only the winners get shared. That is survivorship bias. Our test shows that holding the whole screened pool through 2022 to 2026Q1 produced a -98% maximum drawdown in the US micro-cap biotech cell. The screenshots you have seen are the numerator of a fraction whose denominator quietly delisted.
No, and we want to be precise about why: the +0.14 Sharpe on the European health subset has a confidence interval that straddles zero, a negative deflated Sharpe, and a -73% drawdown. It is a relative tilt inside a weak pool, interesting as evidence of the regional inversion, and not tradeable on its own. We publish it as context for excluding the US tape, not as an alpha claim.
Parts of it. Two of the five criteria (position increase over 10%, buyback cannibals) and the gold and mining sector leg were untestable with our data model, so they are untested rather than refuted. Point-in-time market caps would make the small-cap test cleaner, though since the current construction is survivorship-optimistic, cleaner data should make the result worse, not better. If the data lands, we will rerun and publish.
This is not investment advice.
DEUTZ drew a fresh insider-buying cluster in August, with Patricia Geibel-Conrad adding EUR 103,114 after a sharp defens...
JPMorgan turned bullish on BNP Paribas, but the bank's latest move sits inside a strong sector tape, a fresh rating affi...
Airbus trades near €195 as labor friction, engine bottlenecks and a new space JV frame the stock. The insider record sta...
Sanofi sits near recent lows after a strong Q2 and vaccine updates. The stock has no fresh insider buy to lean on, and t...
Nordnet’s co-CTOs filed matched buys and sells on 31 August as the Nordic broker keeps growing, while Avanza remains the...
Sea Ltd fell 4.93% on August 31 as Shopee kept growing and four insiders sold. Here is what the filings add, and what th...