Straud criticises statistical work. It is built to be judged the same way.
This page publishes its calibration rather than asserting it.
Contents
Every figure below is produced by a test suite that runs on every build, across macOS, Linux and Windows. None of it is a marketing number.
Every finding is labelled with one of five categories. In full:
Statistical validity — deflated Sharpe against declared trials, minimum backtest length, effective sample size under clustering, power to detect the claimed effect, and overlapping positions counted as independent events.
Data integrity — look-ahead, columns that are a frozen assumption rather than a measurement, archive coverage holes, survivorship, events fabricated by absence, and declared gross and net returns that are identical.
Robustness — probability of backtest overfitting, placebo cutoffs, parameter plateaus rather than spikes, era consistency, P&L concentrated in a few clusters or entities, recurrence-dependent edges, and whether a losing counterparty can be named.
Execution realism — cost sweeps and break-even headroom, holds crossing funding settlements, borrow against the observed band, and short availability under Reg SHO.
Capacity — trade size against accessible liquidity, and capacity headroom.
Both independence axes. Events cluster in time and by entity. These are different defects and neither fix addresses the other, so Straud reports n events, n dates and n entities — and tells you when the two disagree.
A flag is only worth acting on if it is rare. These are the two corpora that establish it: strategies built to be clean, which Straud should pass and does.
| Corpus | Clean strategies | Falsely flagged |
|---|---|---|
| Synthetic, swept across Sharpe and sample size | 45 | 0 |
| Real market returns, three markets | 15 | 0 |
Zero false positives in 45 clean cases bounds the true rate at 6.4% with 95% confidence. That bound, not the zero, is the honest claim.
Defects are planted deliberately, so the right answer is known in advance: selection bias, clustered observations, parameter spikes, look-ahead, fabricated cutoffs, costs that kill the edge, capacity limits, absence-defined events.
| Corpus | Planted defects | Caught |
|---|---|---|
| Synthetic | 28 | 27 |
| Real market returns | 25 | 25 |
Generated data encodes its author's assumptions about what returns look like — and those assumptions are exactly what decides whether a check false-positives. The noise therefore comes from real markets instead.
| Market | Returns | Instruments | Skew | Kurtosis |
|---|---|---|---|---|
| US small-cap equities | 6.5M | 6,819 | +5.5 | 105 |
| Crypto perpetuals | 114k | 116 | +2.8 | 30 |
| FX (ECB reference rates) | 179k | 29 | +4.8 | 275 |
Gaussian kurtosis is 3.0. Adding crypto alone exposed a 40% false-positive rate that equities had not — three genuine defects, fixed. Every market added has found something the previous ones could not.
40 of 40 cases correct across all three.
Every corpus above uses injected effects. That establishes correct behaviour on planted defects — not that Straud says anything useful about work produced by real researchers under real incentives.
So: 204 published equity anomalies (Open Source Asset Pricing), each with its original paper's sample window and returns running long past publication. Straud sees only the in-sample window — exactly what the original author had.
| Result | |
|---|---|
| Post-publication decay reproduced | 53% (published literature: ~58%) |
| Anomalies Straud flagged, later return | +0.22%/mo |
| Anomalies Straud passed, later return | +0.41%/mo |
| Significance | Welch t +2.76, permutation p 0.0040 |
Flagged anomalies did measurably worse after publication than those it passed. The flagged, passed and significance rows are measured at 50 declared trials; the decay figure does not depend on a trial count.
Straud's deflated Sharpe does not beat the in-sample t-statistic you already have: AUC 0.741 against a naive baseline of 0.744.
That comparison moves with the declared count: at 200 trials it matches the baseline at 0.744, and at 1,000 it passes it at 0.748 — which is the same point from the other side. The deflation is only worth what the trial count is worth.
The reason is mathematical, not a defect: at a constant trial count the deflated Sharpe is a monotone re-expression of the Sharpe. Deflation only adds information when the trial count reflects the search that actually happened — which is what the trial ledger is for, and why this page claims no more than that.
It is published here because a page of only favourable results is exactly the kind of artifact Straud exists to catch.
A finding that infers a cause from a pattern can be right about its measurement and wrong about what the measurement means. So it cannot reach Critical on its own: it needs corroboration from an independent check, and without one it is reported at Serious and says so.
A finding that measures the claim directly is decisive by itself — a placebo cutoff matching your real one, an assumed cost below the venue's published fee, a trial-adjusted Sharpe of 0.001. There is no inferential gap to argue with.
This distinction exists because of a real mistake. An axis disagreement was once read as a fake sample when the true cause was that the edge lived in recurring names — the measurement was correct and the conclusion was not. Requiring corroboration for inferential checks costs nothing in detection: every planted defect that reaches Critical does so through a direct check.
| Test | Result |
|---|---|
| Attacks on the validator's own inputs | 10 / 10 held — 5 caught, 4 surfaced, 1 correctly silent |
| Hostile and malformed inputs | 59 probes |
| Garbage given a confident clean verdict | 0 |
A clean verdict on garbage is the worst thing Straud could produce, because it is indistinguishable from a clean verdict on a real strategy.