TradingEvolve · Knowledge Base

To read your rating, first read the standard

A rating answers "which tier are you in"; the knowledge base answers "why is it calculated that way, and what should you train next". Here we publish the ability framework and the professional metric definitions — not teaching material, but the public standard we ourselves use to judge trading ability.

5 layersAbility model
30Metrics
11+Glossary entries

Why we publish this standard

Most "trader leaderboards" rank performance, not ability. Performance = ability × luck × market environment — without stripping out the latter two, everyone is a genius in a bull market. We publish our measurement criteria openly so that (1) rating results can be independently scrutinised, and (2) you know which direction to train.

Section 1 · Ability framework

A five-layer model, bottom-up; each layer is a prerequisite for the one above. Lower layers are "harder" (hard to fake), upper layers "softer" (depend on interpretation)

L1 Survival & risk

Whether you stay alive. Everyone is willing to take risk; the scarce skill is controlling it — returns are soft evidence, risk is hard evidence. Key metrics: max drawdown, drawdown recovery time, tail loss, losing-streak tolerance, leverage and concentration, risk of ruin.

L2 Return quality

How much is left after luck is removed. Absolute return is meaningless; returns must be read against the volatility and drawdowns they required. Key metrics: Sharpe / Sortino / Calmar, excess over benchmark, return consistency, trimmed test, sample size and confidence.

L3 System efficiency

Whether you have a positive-expectancy approach. Not "it made money this time" but "it makes money over time". Key metrics: expectancy, reward-to-risk, profit factor, win-rate × R:R pairing, opportunity utilisation, cost erosion, turnover match.

L4 Execution discipline

Whether you can actually do what you know. This layer comes entirely from in-platform behaviour — the most controllable, most trainable, and most capable of forming a feedback loop. Key metrics: stop-loss execution rate, plan consistency, position sizing consistency, loss-chasing rate, chasing extremes, overtrading.

L5 Cognitive fit

Whether you truly understand what you are doing. The most fundamental and hardest to measure — questionnaires can test "do you know it", not "can you do it". Key metrics: market mechanics knowledge, strategy-instrument-timeframe fit, review loop, adaptability, behaviour stability in extreme markets.

Three dividing lines to remember

Ability ≠ performance

Strip out market beta and luck first; what remains is ability. A ranking that skips this only ranks "who caught the good times".

Process ≠ outcome

Good decisions can lose money; bad decisions can make money. Judging only outcomes rewards gambling.

Measurable ≠ fully measurable

Not every ability can be measured directly. Measuring too aggressively causes false positives; not measuring causes blind spots — hence layering and confidence labels.

⚠️ The win-rate trap: a 90% win rate with 10:1 R:R (one loss wipes everything out) is usually weaker than a 30% win rate with 3:1 R:R. Simply never using a stop-loss can push a win rate to 95% — any rating that sells itself on a high win rate is not credible.

Section 2 · Professional metric glossary

Common metrics in four categories, each with "how it's calculated / how to use it / where its blind spots are"

Sharpe ratio

[How it's calculated] Excess return ÷ total volatility. [How to use it] Measures return per unit of volatility — the most widely used risk-adjusted metric. [Blind spots] It penalises upside volatility too, so trend strategies look worse than they are, and it is highly sensitive to the sample period.

Sortino ratio

[How it's calculated] Excess return ÷ downside volatility. [How to use it] Penalises only losing volatility, which matches trader intuition better. [Blind spots] Downside deviation is unstable with small samples, and definitions vary across platforms — compare only within the same methodology.

Calmar ratio

[How it's calculated] Annualised return ÷ max drawdown. [How to use it] Directly answers "how many times the drawdown did we earn". [Blind spots] It only looks at the single worst drawdown, so one extreme event can dominate; strategies with small but long drawdowns get overstated.

Max drawdown (MDD)

[How it's calculated] The largest peak-to-trough decline in equity. [How to use it] The most intuitive evidence of "couldn't take it" — the core metric of the survival layer. [Blind spots] It has no time dimension: losing 30% over three months versus three years are entirely different experiences, so always read it alongside drawdown recovery time.

Drawdown recovery time

[How it's calculated] Time needed to return from the trough to the previous peak. [How to use it] Harder than drawdown depth — only recovering counts, and professional institutions require it. [Blind spots] In extreme cases recovery may never happen; pair it with "did it make a new high".

Profit factor (PF)

[How it's calculated] Gross profit ÷ gross loss. [How to use it] A single number containing both win rate and R:R; above 1.5 is good, above 2 is excellent. [Blind spots] Sensitive to one very large trade, and easily inflated by a single outlier when the sample is small.

Expectancy (ER / R-multiple)

[How it's calculated] Win rate × average win − loss rate × average loss, expressed in R-multiples. [How to use it] Once expressed in R, traders with different capital sizes become comparable. [Blind spots] Depends on sample size, and R calculated with different stop-loss conventions is not directly comparable.

Reward-to-risk (R:R)

[How it's calculated] Average win ÷ average loss. [How to use it] Must always be read together with win rate; either one alone misleads. [Blind spots] High R:R often comes with a low win rate and long losing streaks, which is a real test of discipline.

Trimmed test

[How it's calculated] Remove the most profitable trades (e.g. top 5%) and check whether the result is still profitable. [How to use it] Distinguishes "the system makes money" from "I was right once" — the hardest robustness check. [Blind spots] The trim ratio must use a fixed convention, otherwise it can be presented selectively.

Cost erosion rate

[How it's calculated] (Fees + funding + slippage) ÷ gross profit. [How to use it] Essential for high-frequency and derivatives trading — funding and slippage can consume the entire profit. [Blind spots] Live slippage is hard to reconstruct precisely, and backtests often assume costs optimistically.

Loss-chasing rate

[How it's calculated] The proportion of re-entries shortly after a loss where position size significantly exceeds the trader's average. [How to use it] Captures the substance of "revenge doubling down" without penalising normal re-entry. [Blind spots] Threshold choice affects results and must be calibrated to each trader's cycle.

Ulcer Index

[How it's calculated] A root-mean-square measure combining drawdown depth and duration. [How to use it] Merges "how deep" and "how long" into one number, closer to real pain than MDD. [Blind spots] More complex to compute and harder to explain intuitively, so usually a supporting rather than primary metric.

Theoretical basis and references (selected)

This framework is not invented here — every layer traces back to checkable literature or industry practice. Selected references for further reading:

Sharpe, W.F. (1966). Mutual Fund Performance. Journal of Business 39(1) · Sortino, F. & Price, L. (1994). Performance Measurement in a Downside Risk Framework. Journal of Investing · Young, T.W. (1991). Calmar Ratio: A Smarter Tool. Futures Magazine · Martin, P. (1987). Ulcer Index · Van K. Tharp (1998). Trade Your Way to Financial Freedom. McGraw-Hill · Pardo, R. (2008). The Evaluation and Optimization of Trading Strategies. Wiley · Vince, R. (1990). Portfolio Management Formulas. Wiley · Kahneman, D. & Tversky, A. (1979). Prospect Theory. Econometrica 47(2) · Barber, B. & Odean, T. (2000). Trading Is Hazardous to Your Wealth. Journal of Finance 55(2) · Shefrin, H. & Statman, M. (1985). The Disposition to Sell Winners Too Early and Losers Too Long. Journal of Finance 40(3)

These references are cited solely to attribute the methodological sources and do not constitute an endorsement of any third-party method, product or institution. The five-layer division is our own synthesis; specific band thresholds are internal platform conventions and subject to what is actually shown in the product.

When to come back to this page

You don't understand why your rating movedYou want to know how a specific metric is calculatedYou want to find which layer your weak spot is inYou want to check whether our methodology holds up