The findings we can name: 145 flags across 26
audited backtests
2026-09-04. Every figure quoted here is a field in a
published audit JSON, and the listed audits are linked in place.
Lizard carries seven named findings,
CryonX carries two, and
Logan carries ten. All three went
through the same engine on the same kind of tester report. Across our
26 default audits that engine has raised 145 flags under
22 distinct names, and the list of names it is able to raise at
all is finite, short and published. This page is that list.
What a flag is
A flag is not a sentence we write. It is a record with five fields
and it is produced by a rule that either fires or does not.
- A dimension. One of six, namely data quality, structure,
costs, concentration, regime and prop fit.
- An id. A short stable name such as K-8020 or
S-GRID. The prop fit ids carry the profile name inside them, so
one rule evaluated against three rule sets can appear under three
different ids.
- A severity. One of info, caution or red. There is a fourth
level called ok and no flag ever carries it, because a rule that finds
nothing writes nothing.
- A title and a detail. Both are generated from the measured
numbers of that audit, which is why two audits with the same id rarely
read the same.
The lamp on an audit page is then the highest severity among the
flags of its dimension, and every lamp starts at ok. That single
sentence is the whole aggregation rule, and it has a consequence worth
holding on to. An ok lamp does not mean we looked and found the
dimension healthy. It means none of the named rules for that dimension
fired.
The twenty two names, ordered by how often they appear
The count is the number of the 26 default audits that carry the id.
The 145 flags break down into 45 red, 59 caution and
41 info.
- P-iqcapital_classic-CONS, 19 audits. The consistency rule of
the prop profile is breached in at least one year. This is the single
most common finding in the catalog.
- K-8020, 18 audits. Eighty percent of the profit was made on
a small number of days.
- R-NEGYEARS, 16 audits. Two or more calendar years closed
negative.
- S-STACK, 15 audits. More than one position open at the same
time, either as heavy stacking or as plain parallel entries.
- S-HEDGE, 12 audits. Long and short held at the same
moment.
- DQ-SILENT-HEAD, 11 audits. The robot refused to trade the
opening stretch of the window it was tested on.
- S-GRID, 7 audits. Entries added against an open losing
basket, which is the averaging signature.
- C-SHARE, 6 audits. Costs eat a large share of the gross
profit.
- C-SWAP, 6 audits. Financing costs exceed commission.
- R-ONEYEAR, 5 audits. One year carries more than 60 percent
of the whole result.
- DQ-FAILS, 5 audits. More than half a percent of intended
entries never filled in the simulation.
- DQ-TICKS, 4 audits. Less than half the window ran on real
ticks.
- P-ftmo_challenge-SIZING, 3 audits. Death free only at a
quarter of the tested size or smaller, measured against that
profile.
- P-iqcapital_classic-MPL, 3 audits. The per position loss
limit is breached even at one third size.
- DQ-RANDOM, 3 audits. The tester log contains randomization
prints, so a single run understates the spread of outcomes.
- R-LONGBETA, 3 audits. The profit leans on the long side
while the instrument itself rose.
- S-MARTINGALE, 2 audits. Volume rises after losses more
often than after wins.
- P-iqcapital_classic-SIZING, 2 audits. As above, against the
IQ Capital rule set.
- P-generic_6pct_trailing-SIZING, 2 audits. As above, against
the generic trailing rule set.
- P-iqcapital_classic-DEATH, 1 audit. No sizing down to one
eighth survives the drawdown floor.
- P-generic_6pct_trailing-DEATH, 1 audit. The same finding
against the second profile, and it is the same audit,
Scalping Robot Pro.
- EQ-FLOAT, 1 audit. The floating drawdown ran far deeper than
the closed trade curve shows.
Two things in that list are worth a second look. The most common
finding is not about the trading at all, it is about a funding rule,
and the rarest finding is the one that measures risk while a position
is still open.
The thinnest and the thickest audit
CryonX and
Gold Snap carry two flags each
and no default audit carries fewer. Logan
carries ten, and so does Pulse Engine, which is measured but not
listed. No default audit carries none at all.
Logan is worth naming because five of its ten flags are prop fit
findings across the three profiles. A single structural property, in
its case sizing that cannot be scaled down far enough, is counted once
per rule set it collides with. Reading the raw flag count as a ranking
would therefore punish it three times for one trait, which is exactly
why the audit page shows six lamps and not one number.
Thirteen checks have never fired, and most of them cannot
Counting each prop fit rule once rather than once per profile, the
engine defines 32 named checks. 19 of them have appeared
in a published audit. 13 have not, and for most of them the
reason is our own test setup rather than the robots.
- The missing log check. It fires when no tester log was
supplied. All 26 default audits record a log, so it has never
had an opportunity.
- The zero delay check. It fires when the test ran without
execution delay. All 26 ran at 10 milliseconds, which is
the fixed value of our protocol.
- The two pairing checks. They fire when exits could not be
matched to entries from the log. All 33 stored audits record a
pairing confidence of exact.
- The three integrity checks. They fire on a failed balance
reconstruction, on deals out of chronological order and on duplicated
deal numbers. No default audit records any of the three.
- The zero commission check. It fires when commission is zero
and no retrofit was applied. 22 of 26 carry a retrofit,
and the remaining four carry commission natively at -22.14,
-5.59, -697.86 and -41.43 USD.
- The zero swap check. It fires when swap is zero although
positions were held over a rollover. Exactly two runs show a swap of
zero, Lizard and
Prop Firm Gold EA, and both
show 0 overnight trades. Not paying financing you never owed is
not a modelling gap.
- The carrier check. It fires when one sub strategy produces
the result while the others are ballast. 17 of 26 audits
detect more than one sub strategy, and Logan detects 76, so the
rule had plenty of chances and stayed quiet in all of them.
- The overnight ban and the weekend ban. Both fire only
against a profile that forbids the holding. All three profiles we run
allow overnight and weekend positions, so these two can never fire
until a fourth profile is added.
- The data head check. It fires when the window opens before
the symbol has history. No default audit records that
classification.
That is the honest shape of a clean data quality lamp in this
catalog. It is partly a statement about the report and partly a
statement about the fixed conditions we test under.
Run these checks yourself
- Ask what the absence of a warning was allowed to mean. A
review with no complaints has either checked something and found it
clean or never had a rule for it. The two look identical from the
outside and only a published list of checks tells them apart.
- Count the findings before you weigh them. Ten flags in one
audit and two in another is not a ranking when five of the ten are the
same trait measured against three rule sets.
- Read the detail line, not the title. Two audits under the
same id can differ by orders of magnitude, because the numbers in the
detail come from that run alone.
Every figure above is public in the audit files linked in place, and
the category pages list the verdict lamps for free. Your own tester
report? The browser check is free.
The honest limits. A finite list of checks
finds a finite list of problems. Nothing here searches for a defect we
have not thought of, and a robot whose weakness has no rule in this
catalog will pass every one of them. The thresholds inside the rules,
such as half a percent of failed entries or 60 percent of profit in one
year, are conventions we chose and published rather than results. Two
of our 26 default audits, Quantum Queen and Quantum Queen X, are the
same measured run published under two slugs, so the counts above cover
25 distinct runs. And every one of these flags describes a backtest
over a window that ends in August 2026. None of them is a statement
about what these robots will do next.