Skip to content
EN
Send my files

Methodology

How we audit without making anything up

The principle that governs everything else: when a piece of data is missing, the gap is declared — it is not filled in with an optimistic guess. It sounds obvious and it is almost never done, because a report with gaps sells worse than a full one.

01The five pillars

The five pillars

The score comes out of five blocks with different weights. If one of them cannot be measured with the data supplied, its weight is shared out among the ones that can, and the report says which were missing. Below half the criterion we issue no score at all: grading with less than half the evidence would mean inventing the other half.

PillarWeightWhat it measures
Overfitting35 %PBO, Sharpe corrected for search, walk-forward and sensitivity. This is the differentiator: it measures whether you chose well or got lucky.
Risk20 %Maximum drawdown on the real equity, and probability of ruin.
Microstructure20 %How realistic the costs are: spread, commission, financing — and what happens when you triple them.
Stress15 %The shape of the loss tail, expected loss in the worst cases, and reordering of the history.
Value against doing nothing10 %Whether the strategy beats buying the asset and waiting.
02Rejection rules

The six rules that make a “rejected” believable

Without these, saying an idea does not work is just an opinion. It is the same protocol we apply to our own hypotheses before applying it to anybody else’s.

  1. Write the experiment down BEFORE running it: what is being tested, what is expected, and what would count as success. Fixing the criterion before seeing the result is the only thing that stops the goalposts moving afterwards.
  2. At least 100 trades per cell. With fewer we record the result, but we do not conclude from it.
  3. Neighbouring threshold: a real signal survives at the values either side. If it only works at the exact number you tried, it is noise shaped like a signal.
  4. Positive both in and out of sample, never only in the combined total.
  5. A pretty result in a side script does not count until it runs in the real engine.
  6. The frozen out-of-sample period is the final judge, not another adjustment on data that has already been spent.
03Certificate

The certificate can be checked without trusting us

Every report carries an identifier and an Ed25519 digital signature, with the public key published. Anyone can verify that the document came from here and that not a single figure has been touched since.

And it is reproducible: starting from the same trade history, the analysis produces exactly the same results and the same identifier. If two audits of the same file disagree, one of them is wrong — and it can be demonstrated which.

04What we don’t do

What we do NOT do

We do not run your bot. We do not tell you how to improve it, nor hand you a list of indicators to try: that would make the auditor part of your search, and every additional test on the same data consumes your statistical luck. It is the same conflict of interest that separated auditing from accounting consultancy after Enron.

Nor do we give investment advice. A favourable verdict says that a program holds up statistically on the history supplied. It does not say it is going to make money.