← Back to Blog
2026-09-18 Coaching

What Trading Coaches Should Track for Their Students (Screenshots Are Not Data)

Every coach eventually notices the same thing: the trades a student brings to a review are not a random sample of the trades they took. That is not a motivation problem you can solve by asking for more screenshots. It is a sampling problem, and the fix is structural.

Why Screenshot Reviews Do Not Scale

Start with the charitable reading, because it is also the correct one: students are not hiding their bad trades. They are showing you the trades they remember. Those are two different biases and only the second one is universal.

Memory of a trade tracks the size of its outcome and the strength of the feeling attached to it. The 8R runner gets remembered. The full stop-out that hurt gets remembered. What does not get remembered is the body of the distribution: the 0.3R scratch, the trade closed early out of boredom, the entry taken twenty minutes outside the session the curriculum specifies, the position that was a quarter larger than planned for no reason the student could reconstruct afterwards. Those trades are forgettable precisely because nothing dramatic happened, and they are also where the drift lives.

So a screenshot review gives you a sample conditioned on outcome magnitude. You see the tails and never the middle. Every conclusion you draw from it is drawn from the wrong distribution, and the error does not shrink as the student sends you more screenshots - it just gets better documented.

Three further problems ride along with it:

None of this is an argument against looking at charts with a student. It is an argument against charts being the only instrument you have. The chart review is where you teach; it is a terrible place to find out what needs teaching.

The Four Numbers That Tell You Where a Student Actually Is

You do not need a dashboard of forty metrics. You need four numbers that answer four different questions, read in the same order every week.

Number What it answers What it cannot tell you
Expectancy
mean R
Is there anything here at all? Whether it came from one trade or forty
Risk adherence
stdev of planned risk %
Are they trading the size they said? Whether the sizing was justified
Tag coverage
% of trades labelled
How much activity is on-programme? Whether the labels are honest
Rolling SQN
(mean R ÷ stdev R) × √n
Is the edge separable from noise yet? Almost anything, below about 30 trades

Expectancy is the mean of the R series, and it is first because it is the only one of the four that speaks to whether the student has anything at all. It is also the one most likely to be built on five trades, so never read it without the trade count sitting beside it.

Risk adherence is the one coaches skip, and it is the one that matters most. It gets its own section below, because there is a piece of arithmetic in it that surprises people.

Tag coverage is a process measurement disguised as a data-quality measurement. If a third of a student's trades carry no setup label from your curriculum, that is not laziness about journaling - that is a third of their trading happening outside the thing you are teaching them. The untagged pile is usually the most interesting object in the account.

Rolling SQN combines mean and spread into one figure, and it comes with a warning label. The formula carries a square-root-of-sample-size factor, which means the same unchanged system scores higher the longer it is traded. A student whose SQN rose from 0.7 to 1.9 over a quarter may have improved, or may simply have logged more trades. Compare a student only against their own history at a similar trade count, and never against another student's figure computed over a different number of trades.

Worth knowing about the tool, including ours. Most journals - SignalDeck included - will compute and display SQN as soon as there are two trades with any dispersion between them. There is no minimum-sample guard stopping the number from rendering. A confident-looking score on nine trades is a number the software is happy to print and you should not be happy to read. Treat the trade count as part of the metric, not as metadata.

Diagnosing Process Failure vs Edge Failure

This is the single most valuable thing a coach can do, and almost nothing on a standard dashboard helps with it. A student is down over sixty trades. There are two completely different diagnoses and they look identical in the equity curve:

The prescriptions are opposites. Coach discipline at a student whose edge is broken and you will spend six months making them very disciplined at losing money. Rewrite the strategy for a student whose problem was execution and you have removed the only thing that was working.

The data cut that separates them is simple and almost nobody runs it: partition the trades by rule compliance, not by outcome, then compute expectancy on each subset separately.

Compliant trades Non-compliant trades Diagnosis
+0.31R −0.62R Process failure. The edge is real. Coach execution.
−0.18R −0.24R Edge failure. Discipline is not the problem. Rebuild the setup.
−0.15R +0.29R The improvisation is beating the curriculum. Find out why.

That third row is uncomfortable and it is the reason to run the cut honestly. A student who is quietly outperforming the thing you taught them is telling you something about your own rules, and the only way to hear it is to have run the partition before you knew which way it would come out.

Two constraints make or break this. First, compliance has to be recorded when the trade is logged, not reconstructed at review time. Nobody can honestly answer "did I follow my rules?" about a trade from three weeks ago while looking at whether it worked. The answer will be contaminated by the outcome every time. This is what a pre-trade entry checklist is for: in SignalDeck the checklist items attach to the individual trade, and a per-user preference can hard-gate the Log Trade button until every item is ticked, so the record is made at the only moment it can be made truthfully. The post-mortem fields on a closed trade - setup quality, stop management, verdict, training gaps, and a free-text mentor feedback field - carry the rest.

Second, both subsets will be small. Splitting sixty trades into forty-two compliant and eighteen non-compliant leaves you with two noisy averages, and the difference between them is noisier still. This cut is directional evidence for a conversation, not a verdict. It earns its place because the alternative - guessing - has no error bars at all.

Risk Adherence Is the Leading Indicator

Students do not usually blow up because their strategy stopped working. They blow up because their position size stopped being a calculation and started being a reaction. That failure arrives months before the edge fails, and catching it early is most of the job.

Here is the part that surprises people, and it is the same arithmetic that makes R-multiples so useful everywhere else: the R series cannot see position-size drift at all.

An R-multiple is the outcome divided by the risk taken on that trade. A student who plans to risk 1% and takes a full stop-out logs −1R. A student who sizes up to 3% on the same setup, in a fit of impatience, and takes a full stop-out also logs −1R. The normalisation that makes trades comparable is the same normalisation that erases the single most dangerous thing the student did.

Trade Planned risk Result in R Result in % of account
11.0% −1.0R−1.0%
21.0% −1.0R−1.0%
32.5% −1.0R−2.5%
43.0% −1.0R−3.0%
Total −4.0R −7.5%

Four identical-looking losses in the R column. Nearly double the planned damage in the column that decides whether the account survives. If your review reads only the R series, this student looks like someone having a normal bad run, and you will find out otherwise when the drawdown limit is breached.

So track planned risk as a percentage of equity as its own series, and read two things from it. The first is its dispersion: a student running fixed-R position sizing should produce an almost flat line, and any visible spread is the finding. The second is its correlation with the previous trade's outcome, which is where the diagnosis lives. Sizing up after a loss is revenge. Sizing up after a win is overconfidence. Sizing down after a loss is fear, and it is the one that quietly caps a student's upside while looking like prudence. All three are invisible anywhere except this column.

One related leak worth checking in the same pass: whether the stop that defines R is the one the student set before entry. R is anchored to the original stop, so a student who widens a stop mid-trade has increased their real risk without the R series recording it. A journal that keeps the original stop as the anchor will show that as a trade whose realised loss exceeds −1R, which is a clean, searchable signature for a habit that is otherwise very hard to catch.

Building a Shared Setup Taxonomy

Everything above assumes a student's labels mean what you think they mean. Across a cohort, that assumption fails immediately unless someone designs against it.

Leave the setup field open as free text and within a few weeks it contains four spellings of the same pattern, a couple of single-letter entries made to get past the form, and at least one misspelling that will silently split a setup's statistics in two forever. Every per-setup number you compute after that is wrong in a way that does not announce itself.

A taxonomy that survives a cohort has four properties:

It helps to separate the two kinds of label. One strategy per trade - the thing being scored, drawn from a closed list you control - and as many tags as are useful for the conditions around it: session, news proximity, emotional state, whether it was a re-entry. SignalDeck models them exactly this way, one strategy and many tags per trade, because they answer different questions: the strategy label is what per-setup expectancy is computed on, and the tags are what you slice it by.

Fix the taxonomy before the cohort starts. Retagging three hundred historical trades is the task that never gets done, and a taxonomy changed mid-programme gives you two incomparable halves of a dataset instead of one.

Reviewing a Cohort Rather Than an Individual

Once every student's trades carry the same setup labels and the same R normalisation, the most valuable view stops being per-student and becomes per-setup across the whole cohort.

Suppose eleven of your twelve students have negative expectancy on the same breakout setup. Reviewed one student at a time, that is eleven separate conversations about discipline, and eleven students who conclude they are the problem. Reviewed across the cohort, it is one finding about your teaching - or about the regime the market has been in since you taught it.

The shape of the distribution tells you which:

SignalDeck's team view was built for this shape of question. A team gets a shared trade feed, a leaderboard, per-member monthly R contribution, shared post-mortems, per-member journaling streaks and a team-level streak. The journaling streak is worth more attention than it sounds: the student who stops logging is the one you most need to see, and their absence from the data is itself the earliest signal you will get.

Three honest limits on that feature as it stands today.

  • A team is symmetric. Every member can see every other member's trades - there is no one-way coach view. For a cohort learning in the open that is the feature; for a confidential one-to-one mentorship it is not, and you should know which one you are buying.
  • There is no elevated coach role today. Membership carries a role field, but no part of the product grants a coach permissions a student does not also have.
  • Teams are not self-serve. A cohort is provisioned by us rather than created from the app, which is why the invitation below is to talk to us rather than to click a button.

Team features are available on every tier, including the free one, which is capped at membership of a single team.

Verified Track Records for Graduating Students

The last thing a good programme produces is a student who can prove something about themselves to somebody else - a prop firm, an allocator, a first client. This is where the screenshot problem comes back in its most expensive form.

A track record is credible in proportion to what its publisher could not choose. A screenshot of a P&L figure is a claim with no such property: it is selected, un-auditable, and trivially the best day of the year. A profile computed from a complete trade log is a stronger claim, because the numbers that matter are derived rather than asserted - the trade count, the win rate, the equity curve, and above all the maximum drawdown, which is the number nobody would ever volunteer.

A student on SignalDeck can enable a public profile hosted on sgnldk.com and share the link. The owner controls three things and only three: whether the profile is on at all, which strategies appear on it, and whether dollar P&L is hidden while the R-multiples stay visible. That last option is genuinely useful - account size is nobody's business and R-multiples are the more informative number anyway.

The honesty constraint is the second of those three, and it has to be said plainly: because the owner can filter the profile down to selected strategies, a public profile is a verified claim about the strategies shown, not an audit of the whole account. That is a real limitation of using one as proof, and a post about cherry-picking would be a poor place to pretend otherwise. Two practical consequences follow. If you are reading someone's profile, check the trade count and the period before you treat it as evidence. If you are a student publishing one, publish the account unfiltered - a profile that shows the losing strategy alongside the winning one is worth several times more than one that does not, for exactly the reason this whole post exists.

Two administrative notes. Public profiles are a paid-tier feature, and the page returns a 404 if the subscription lapses - so a link on a CV can go dead. And the profile belongs to the student, not to the programme, which is the correct arrangement: it is portable evidence they keep, which makes it worth more to them and, indirectly, a better advertisement for you than any testimonial you could write.

A graduating student with two hundred logged trades, a visible drawdown, a stable tag distribution and an honest record of the trades that did not work is a better result than one with a screenshot of an 8R day. It is also considerably harder to fake, which is the entire point.

None of the above requires a student to journal more than they already do. It requires the journal to be complete rather than curated, the labels to be shared rather than personal, and the planned risk to be written down before the trade rather than inferred after it. Those three properties are what turn a pile of student screenshots into something you can actually coach from.

Frequently Asked Questions

How can a trading coach track student progress?

By reading the student's whole trade log rather than the trades they bring to a review. A screenshot sample is conditioned on what the student remembers, and memorability tracks outcome size, so the trades you are shown are systematically the tails of the distribution and never the body. The practical setup is a journal the student logs into as a habit, a shared setup taxonomy so their labels match your curriculum, and a small fixed set of numbers you read the same way every week. What matters is that the sample is complete and the results are expressed in R-multiples, which makes trades comparable across different symbols, account sizes and position sizes.

What metrics show whether a student is improving?

Four, read together. Expectancy, the mean R-multiple, tells you whether there is anything there at all. Risk adherence, the dispersion of planned risk as a percentage of account equity, tells you whether they are trading the size they said they would - and this one is invisible in the R series, because R normalises each outcome by the risk taken on that trade. Tag coverage, the share of trades carrying a setup label from your curriculum, tells you how much of their activity is on-programme. Rolling SQN combines mean and spread, but it carries a square-root-of-sample-size factor, so it climbs on trade count alone and should only be compared against the same student's own history at a similar number of trades.

How do I tell if a student's problem is discipline or strategy?

Partition their trades by rule compliance rather than by outcome, then compute expectancy separately on each subset. If the compliant trades are positive and the non-compliant trades are negative, the strategy works and the problem is execution. If both subsets are negative, discipline is not the issue and no amount of coaching on process will fix it - the edge is the thing that is broken. If the compliant subset is negative and the improvised trades are positive, your curriculum is the weaker of the two, which is worth knowing. This only works if compliance is recorded at the moment the trade is logged; it cannot be reconstructed honestly weeks later, and both subsets will be small, so treat the result as directional rather than conclusive.

Can students share verified results?

They can share a profile computed from their log, which is a stronger claim than a screenshot but is not an audit. On SignalDeck a student can enable a public profile at sgnldk.com that derives trade count, win rate, maximum drawdown and the equity curve from the underlying trades. The owner controls three things: whether the profile is on, which strategies appear on it, and whether dollar P&L is hidden while R-multiples remain. Because strategy filtering exists, a public profile is a claim about the strategies shown and not about the whole account, so read the trade count and the period before treating one as evidence, and publish unfiltered if you want yours to carry weight. Public profiles are a paid-tier feature and the page returns a 404 if the subscription lapses.

You cannot coach what you cannot see

Every student can start on the free tier and keep their own log and their own track record. If you run a cohort or an academy and want them reviewed together, come and talk to us about it - cohorts are set up by hand at this stage. Free during beta; Pro is $30/mo and Elite $50/mo when billing launches.

Related reading