Skip to content
All research notes

What a strategy test needs to tell you.

A useful test explains the rules, trading costs, evaluation period and conditions under which the result changes.

A trading idea becomes a research question when its rules can be stated clearly enough to test. The goal is to understand its behaviour across time and assumptions, including the conditions in which it fails.

01

Start with a rule specification

Write down what produces an entry, what invalidates it, how a position is sized and how it exits. Specify the instrument, sampling frequency, eligible hours and any dependency on economic releases.

If a rule cannot be implemented without a judgement made after seeing the outcome, the experiment is underspecified. Clarifying that rule is research work; an attractive result should not substitute for it.

02

Evaluate in time order

The model-development period and evaluation period serve different purposes. Development data supports parameter choices and fitting. Later evaluation data asks whether those choices carry into a period that was not used to select them.

For time-dependent observations, the split should preserve the decision sequence. The scikit-learn documentation provides a time-series splitting approach in which training precedes each subsequent test fold. The exact spacing and any required gap depend on the experiment’s prediction horizon.

03

State the execution assumptions

A simulation needs assumptions about spread, commission, slippage and order handling. A fill at the recorded mid-price is different from an executable quote, especially around an economic release or a thin trading period.

Evaluate how the result changes when costs are less favourable. A strategy that depends on one optimistic fill assumption needs further investigation. The cost model should be saved with the experiment rather than applied only after a headline metric has been chosen.

04

Report more than one number

Our intended research report includes net results after modelled costs, drawdown, turnover, the number of observations and trades, exposure and the evaluation dates. These fields help explain how a result was produced.

Break the evaluation into meaningful periods where possible. One aggregate result can hide long inactive intervals, concentration in a few events or poor behaviour during a different market environment.

05

Ask whether the result is stable

Test nearby parameter choices and reasonable changes in inputs or costs. Record the comparisons that were attempted, including those that performed poorly. Repeatedly choosing the best result without accounting for that search can make an experiment appear stronger than it is.

An unchanged evaluation sample is useful only while it remains independent of selection. Once its results start shaping the next model revision, a later untouched period is needed to evaluate that revision.

06

Separate research from deployment

A historical experiment does not exercise all the failure conditions of an operating trading bot. Connectivity, rejected orders, stale feeds, duplicate signals and restart behaviour belong to a separate execution evaluation.

Our development approach moves from historical testing to simulation and paper trading before any assessment for live use. Promotion should follow an explicit review of data, strategy behaviour and execution controls.

THE RESEARCH PRINCIPLE

A useful research result includes enough information to question it, rerun it and explain its limits.

References & further reading

scikit-learn: time-series cross-validation

This note describes research principles and development objectives. It is not a report of a validated model or live trading performance.