A backtest is one ordering of one set of trades. Our engine takes that set apart and puts it back together 1,000 times, in three different ways, and then shows where the real run lands in the result. Two audits make the point on their own. Lizard finished 183.99 USD ahead after 21,524 trades, and 221 of its 1,000 resampled versions end in the red. Gold Trade Pro finished 13,446.15 USD ahead, and not one of its 1,000 resampled versions loses money. Same method, same catalog, two entirely different kinds of profit.
The permutation keeps every trade result exactly as it is and shuffles only the order, so each path ends on the same total as the real run. The bootstrap draws the same number of trades from the same set with replacement, so the total moves from path to path. The block bootstrap leaves the trade level entirely and works on the daily result series, drawn in blocks of five trading days and wrapped around the end of the series, which keeps a week of clustered results together. All three run 1,000 paths from a fixed seed of 42, and an audit with fewer than 50 trades gets no resampling at all, because percentiles from a smaller sample would be false precision.
The three differ in one thing only, which is how much of the original structure they keep, and that is what makes them readable as a set. The permutation destroys the order and keeps the total. The bootstrap destroys both. The block bootstrap keeps whatever happened inside a trading week and destroys the arrangement of the weeks.
The bootstrap reports one number that answers the question most directly. It counts the paths whose total is zero or below, and divides by 1,000. Thirteen of our 26 default audits report 0.000, meaning no draw from their own trade set turns the run negative. Three report 1.000, and they are the only three of the 26 whose run ends below its starting balance. Logan ends at -24,334.65 USD, Smart Gold Hunter at -1,418.55 USD and Scalping Robot Pro at -53,234.68 USD. Those runs lose so evenly that no rearrangement of their own trades rescues them.
The ten audits in between are the interesting ones. Lizard sits at 0.221 on a net result of 183.99 USD. OilVector X sits at 0.115 on 35.23 USD. Quantum Emperor reaches 0.244 on 7.59 USD, and it is one of four audits still measured on a vendor report with a 1,000 USD deposit rather than in our own world. Those four are unlisted, so they are named here and not linked. A profit that a reshuffle of its own trades turns negative once in four or five attempts is a thin profit, whatever the equity curve looks like.
The same block prints the interval that explains it. Lizard's bootstrap puts expectancy per trade at -0.01 USD at the 5th percentile, 0.01 at the median and 0.03 at the 95th. The plausible range straddles zero, so the sign of the result is not settled by 21,524 trades. Gold Trade Pro's interval runs from 1.89 to 2.39 USD per trade and never comes near zero, and its profit factor stays above 1.65 even at the 5th percentile. That is what a result looks like when the trade set carries it rather than the ordering.
The audit page puts the observed drawdown next to the resampled ranges under the heading Resampled max drawdown range versus observed. The block bootstrap is the fair comparison, because it resamples the same end of day balance series with the same formula that produces the published max drawdown. In 18 of the 23 audits with a measurable end of day drawdown, the real drawdown is deeper than the median of that resampled range. Every audit is compared only against its own paths, so the mixture of test setups in the catalog does not enter the count.
In five of them the real drawdown is deeper than the 99th percentile. Adaptive Gold Scalper reaches 0.24 percent against a 99th percentile of 0.13. Gold House reaches 0.60 against 0.22, Gold Snap 0.54 against 0.08, Gold Trade Pro 1.38 against 1.17 and Lizard 2.55 against 1.19. The three audits left out of the comparison are Quantum Queen, Quantum Queen X and Quantum Athena X, whose deepest end of day episode is 0.46 USD for the two Queens and 0.30 USD for Quantum Athena X. On a 100,000 USD balance that is a drawdown of 0.0 percent on both sides of the comparison, so there is nothing to compare.
The two figures do not share a denominator, which is worth naming before anyone leans on a narrow gap. The observed drawdown is measured against the running peak of the balance curve, the resampled ones against the start balance. In the 22 audits that document their deepest episode, that peak sits between 99,998.99 and 108,403.86 USD against a start balance of 100,000, so the two conventions differ by at most 8.4 percent of the reading. None of the five audits above the 99th percentile moves, and three others sit close enough for the convention to decide the side. Gold Atlas reads 0.46 percent against its peak and 0.50 against the start balance, with a resampled median of 0.47. CryonX and TwisterPro Scalper land on their own 99th percentile in the published reading and just above it in the other.
Lizard's deepest end of day episode peaks on 2003-05-13, reaches its trough on 2007-11-08 and is not recovered until 2026-06-18. The depth is 2,545.79 USD, the descent takes 1,172 trading days, the account stays under water for 6,027 of them, and the worst single trade inside the episode loses 2.42 USD.
That is the mechanism in two sentences. A drawdown assembled from thousands of small losses across 1,172 trading days cannot survive a reshuffle in blocks of five days, because the reshuffle breaks the long sequence into pieces and separates them with profitable weeks. What the audit prints as a 2.55 percent drawdown is therefore not a bad day but an arrangement, and the resampling is what makes that visible.
Both other methods say the same thing from their own side. The permutation puts Lizard's longest losing streak at a median of 25 trades and 37 at the 99th percentile, which is what a random order of these results produces. The block bootstrap keeps the weeks intact and puts time under water at a median of 3,264 of 6,067 days, with the 99th percentile at all 6,067.
Resampling never leaves the trade set. Every path is built from results this EA produced in this window, so a market state that never appeared in the window cannot appear in a path either. The method measures how fragile a record is, not what a strategy will do next.
Paths are not stopped when the account dies, and the audits carry that note themselves. It is why Logan's block bootstrap reaches 118.04 percent at the 99th percentile. Anything beyond 100 percent means the simulated account was wiped out and kept trading, so the figure is a measure of severity and not a balance anyone could reach.
The seed is fixed at 42, so a rebuild reproduces every percentile on this page. That makes the numbers checkable. It does not make them a forecast.
Two questions get most of the value out of the section. The first is how many of the 1,000 bootstrap paths end in a loss, because that answers whether the profit survives its own trade set. The second is where the observed drawdown falls inside the block bootstrap range, because that answers whether the drawdown you were shown is what these days produce in a different order. Everything else in the section is detail behind those two questions.