What the tool needs to handle
A backtest loads historical data, evaluates your rules, models trades, and calculates returns. The implementation needs to keep track of what information was available at each decision and when a simulated order could have filled.
Those details need review whether you write the code yourself or generate it. A one-period timing error or missing cost can change the result without causing the program to fail.
The no-code path
Start with a complete rule: the asset, candle interval, entry, exit, position size, and costs. A phrase such as 'buy a 3% dip' leaves open what the dip is measured against and when to sell. Resolve those choices before running the test.
Review the generated code and a sample of trades against that description. Keep the data range and assumptions with the result so you can compare later revisions on the same basis.
What to watch out for
Check these assumptions before interpreting the result:
- Make sure real costs (fees, slippage) are included, or your results will be too rosy.
- Check the data range covers different market conditions, not one lucky bull run.
- Confirm the tool isn't using future information - a real backtest only ever sees the past at each step.
- Vary your inputs. If performance depends on one exact setting, investigate whether the rule was overfit to the test period.
How Premiss does it
Premiss generates and runs a Python backtest from your description. You can inspect the code, trade record, and performance metrics, then revise the rule and compare another run. Use those records to check the strategy's interpretation, costs, and behavior across the historical period.