How It's Scored

The Fleet, Measured

Every engine is put through the same test on the same charts and graded four different ways, then checked against how its money actually behaves. No engine gets to pick its own exam. Here's what each grade means, in plain language.

The short version

An engine is only as good as the worst corner of the market it has to trade. So we score each one across twelve market conditions, take the average, and then subtract how much it wobbles between them. A great average with wild swings loses to a slightly lower average that's steady. That penalty — average minus wobble — is the number every grade is built on.

Four grades look at the same engine from four angles:

SQSx — signal quality SQSr — is the data trustworthy DQSx — real-money behavior FQS — the combined verdict

The test bench — 12 cells

Four stocks with very different personalities, each on three speeds. Twelve situations in total. An engine has to appear in all twelve to be ranked — no cherry-picking the charts where it happens to shine.

AMZN — steadyMSFT — steadyNFLX — jumpyMSTR — wild swings
1-hour2-hour4-hour
The idea under every grade

Average minus wobble

"Reliable-good beats occasionally-great."

Say an engine scores well on the calm stocks and terribly on the wild one. Its average might still look fine — but that average is hiding a landmine. We measure the spread of the twelve scores and subtract it from the average. An engine that scores 70 everywhere keeps almost all of it. An engine that averages 70 by swinging between 95 and 45 gets most of that gap taken away.

Two engines, same average Engine A — twelve scores all near 68
average 68 · wobble 3 → final 65

Engine B — half at 88, half at 48
average 68 · wobble 20 → final 48
Why it matters: the market will eventually hand you Engine B's bad half. The score punishes that risk before your money finds it.
— THE FOUR GRADES —
Lens 1

SQSx — Signal Quality

"How good are the buy and sell marks themselves?"

This grade ignores position sizing entirely. It takes only the engine's raw buy/sell signals and runs them through one identical, standardized trade simulation — the same fixed bet every time, for every engine. That's deliberate: it isolates signal timing from everything else. A clever sizing scheme can't rescue a bad signal here, and a buggy one can't sink a good one. It's the purest measure of "are these marks in the right places?"

The signal quality itself is built from four sub-measures, blended together:

These four are combined by multiplying, not averaging — so an engine can't paper over a zero. Weak on any one axis drags the whole cell down. You have to be at least decent at all four.

Read it as: the quality of the engine's eyesight — does it spot real turning points, cleanly, without a lot of false alarms.
Lens 2

SQSr — Reliability-Adjusted

"Can we actually trust that signal-quality number?"

A score built on only a handful of trades isn't as trustworthy as one built on hundreds. And a score that came almost entirely from a single lucky stock isn't as trustworthy as one earned evenly. SQSr takes the signal-quality score and applies two honesty discounts:

An engine with clean, plentiful, well-spread trades takes almost no discount — its SQSr sits right next to its SQSx. An engine that scored high on a thin or lopsided sample watches the two numbers separate. That gap is the tell.

Read it as: the signal grade after a skeptic has audited it. Small gap between SQSx and SQSr = believable. Big gap = handle with care.
Lens 3

DQSx — Deployment Quality

"Forget the idealized test — how does the real book behave?"

The first two grades run a standardized simulation so engines can be compared apples-to-apples. DQSx does the opposite: it reads the engine's actual trading book — its real sizing, its real drawdowns, its real profit and loss — and asks whether that book is one you'd want to hold. It's built from four money-health checks:

Same rule as before — these four are multiplied, so a weakness anywhere pulls it down, and the twelve-cell result gets the average-minus-wobble treatment.

Read it as: would a careful person be comfortable running this book with real money. An engine can have sharp eyesight (high SQSx) and still keep an uncomfortable book (lower DQSx), or the reverse.
Lens 4 — the headline

FQS — The Combined Verdict

"One number that respects both the signal and the book."

FQS fuses the trustworthy-signal grade (SQSr) with the real-book grade (DQSx), cell by cell, then applies average-minus-wobble across all twelve. It's the number the fleet is ranked by, because it refuses the two easy ways to look good: a beautiful signal on an ugly book, or a comfortable book built on a flimsy signal. To score high on FQS an engine has to earn both, in every corner of the test.

FQS cell = √( SQSr × DQSx ) → board = average − wobble

The square root is just the fair way to blend two scores so neither can dominate — a good number on one side can't fully cover a bad number on the other.

Read it as: the closest thing to a single "how good is this engine, all things considered" — trustworthy signal and livable book, proven steady across twelve market moods.
— AND THE MONEY —
The reality check

The Money Ratios

"Do the grades line up with the cash?"

Grades can drift from intuition, so the fleet board also shows four plain money ratios pulled straight from each engine's real book — the same numbers you'd read off its performance table:

Two familiar numbers are deliberately left off the board:

Raw dollars — because a bigger bet size makes any engine's dollar figures look larger without making it any better. Ratios compare fairly; dollar totals don't.

Win rate — because it's easy to game. An engine can win 95% of the time by taking tiny wins and letting a few losses run, or by using a loss-deferral rule that turns most losses into slivers. A high win rate can hide a bad book, so it doesn't earn a column.

Why both grades and ratios: the grades are careful and comparable; the ratios are blunt and real. When they agree, you can trust the engine. When they disagree, that disagreement is the first thing worth investigating.
Choosing the best

How the podium is decided

The board has eight columns in total — the four grades and the four money ratios. To pick the top engines, we count how many of those eight columns an engine finishes in the top three of. That count is its dominance. The rules:

The standard we're after: dominant where it counts, weak nowhere, proven on the market you actually intend to trade.
Why simulate a fixed bet in the first place?

To compare engines fairly, the signal grades run every engine through the exact same trade simulator with the exact same fixed bet — a $100 buy, selling a quarter of the position at a time. Stripping out the sizing scheme means two engines are judged purely on where their signals fire, not on who has the fancier money-management on top. The real sizing still gets judged — that's what the deployment grade and the money ratios are for.

What do the letter grades or high/low labels mean?

Scores are read against a fixed, stock-market-calibrated scale so they mean the same thing over time. On crypto, that same scale reads harsher — the scores look lower simply because the scale wasn't loosened for a rougher market. That's on purpose: one honest ruler, not a friendlier one for the hard tests. Compare an engine to the rest of the fleet, not to the raw letter.

Why does "consistency" keep coming up?

Because lumpy profits are a hidden risk. An engine that makes its money in a few giant windfalls, with long dry stretches between, is harder to live with than one that pays steadily — even if the totals match. Wildly-swinging stocks make even payouts nearly impossible, which is why consistency is the hardest axis in the fleet to score well on, and why it's watched so closely.

Where do the numbers come from, and how current are they?

Every score on the fleet board is computed from a single results database and refreshed whenever an engine's results are committed. Nothing on the board is typed in by hand — it's generated from the same source of truth each time, so the site can't drift out of step with the actual record.

← Back to the Fleet board