The findings we can name: 145 flags across 26 audited backtests

2026-09-04. Every figure quoted here is a field in a published audit JSON, and the listed audits are linked in place.

Lizard carries seven named findings, CryonX carries two, and Logan carries ten. All three went through the same engine on the same kind of tester report. Across our 26 default audits that engine has raised 145 flags under 22 distinct names, and the list of names it is able to raise at all is finite, short and published. This page is that list.

What a flag is

A flag is not a sentence we write. It is a record with five fields and it is produced by a rule that either fires or does not.

The lamp on an audit page is then the highest severity among the flags of its dimension, and every lamp starts at ok. That single sentence is the whole aggregation rule, and it has a consequence worth holding on to. An ok lamp does not mean we looked and found the dimension healthy. It means none of the named rules for that dimension fired.

The twenty two names, ordered by how often they appear

The count is the number of the 26 default audits that carry the id. The 145 flags break down into 45 red, 59 caution and 41 info.

Two things in that list are worth a second look. The most common finding is not about the trading at all, it is about a funding rule, and the rarest finding is the one that measures risk while a position is still open.

The thinnest and the thickest audit

CryonX and Gold Snap carry two flags each and no default audit carries fewer. Logan carries ten, and so does Pulse Engine, which is measured but not listed. No default audit carries none at all.

Logan is worth naming because five of its ten flags are prop fit findings across the three profiles. A single structural property, in its case sizing that cannot be scaled down far enough, is counted once per rule set it collides with. Reading the raw flag count as a ranking would therefore punish it three times for one trait, which is exactly why the audit page shows six lamps and not one number.

Thirteen checks have never fired, and most of them cannot

Counting each prop fit rule once rather than once per profile, the engine defines 32 named checks. 19 of them have appeared in a published audit. 13 have not, and for most of them the reason is our own test setup rather than the robots.

That is the honest shape of a clean data quality lamp in this catalog. It is partly a statement about the report and partly a statement about the fixed conditions we test under.

Run these checks yourself

Every figure above is public in the audit files linked in place, and the category pages list the verdict lamps for free. Your own tester report? The browser check is free.

The honest limits. A finite list of checks finds a finite list of problems. Nothing here searches for a defect we have not thought of, and a robot whose weakness has no rule in this catalog will pass every one of them. The thresholds inside the rules, such as half a percent of failed entries or 60 percent of profit in one year, are conventions we chose and published rather than results. Two of our 26 default audits, Quantum Queen and Quantum Queen X, are the same measured run published under two slugs, so the counts above cover 25 distinct runs. And every one of these flags describes a backtest over a window that ends in August 2026. None of them is a statement about what these robots will do next.