Simulated data: why and how#
v1 rolls out with a simulated FX pair. The market prices (and therefore every P&L and risk figure derived from them) are generated, not observed. This document explains why, how the data is produced, how it is labelled, and what it can and cannot be used for.
Code: src/trade_engine/quant/simulate_fx_gbm.py. Data: data/raw/EURUSD.csv
(canonical v1 artefact) and data/generated/ (regenerated output).
Why simulate#
Reason |
Detail |
|---|---|
No data licence or feed dependency |
v1 has no vendor connection; simulated data lets the whole front-to-back flow (load, book, P&L, risk, reports) run end to end with no external dependency. |
Deterministic and reproducible |
A fixed seed yields the same history every time, so tests, demos and bug reports are repeatable. |
Known ground truth |
The generating parameters are known (drift, volatility), so calculations can be sanity-checked against theory. For example, 1-day 95% VaR of a simulated return series should be close to \(1.645 \times \sigma/\sqrt{252}\). |
Safe sandbox |
P&L, limits and risk logic can be exercised with no real positions or real money implied. |
Plentiful history |
Ten years of daily bars on demand, enough for historical VaR to have a meaningful window. |
Simulation is a stand-in for market data during v1, not a claim about the market.
What is simulated#
Only daily OHLC prices for one FX pair. Trades are entered by hand (or from the
seed script); the sample trade prices in services/seeding.py were chosen near
the simulated series and are equally hypothetical.
The model#
Spot follows geometric Brownian motion with the interest-rate differential as drift, the standard Garman-Kohlhagen setting for an FX underlying:
The daily close path uses the exact solution, with \(Z_t \sim N(0,1)\):
Why GBM: it keeps the rate lognormal (never negative), ties the drift to covered interest parity (no arbitrage), and needs one volatility parameter, which is appropriate for a limited v1 dataset rather than a stochastic-volatility or jump model.
Parameter |
Default |
Meaning |
|---|---|---|
|
1.0850 |
Starting spot (the first bar’s open) |
|
0.045 |
USD risk-free rate, annualised |
|
0.030 |
EUR risk-free rate, annualised |
|
0.07 |
Annualised volatility |
|
10 |
History length; \(\Delta t = 1/252\), so 2,520 bars |
|
2016-01-04 |
First bar |
|
42 |
Random seed |
How OHLC is built#
Only closes come from the model. The rest is layered on top:
open[t] = close[t-1](no overnight gap process); the first open iss0.highandlowwiden the[open, close]range by a half-day-volatility fraction:high = max(open, close) * (1 + |z| * 0.5 * sigma * sqrt(dt))andlow = min(open, close) * (1 - |z'| * 0.5 * sigma * sqrt(dt)), with independent standard normal draws, so high and low always bracket open and close and ranges scale with the volatility parameter.Dates are business days (Monday to Friday,
pandas.bdate_range), with no holiday calendar.Values are rounded to 5 decimals. Rounding is monotonic, so the OHLC ordering survives it.
The close path is drawn from default_rng(seed) and the OHLC noise from
default_rng(seed + 1), so the two are independent but both reproducible.
Price basis#
The simulated series is a single path with no spread, so it represents a MID
price. Bid and ask would be derived as mid -/+ spread / 2 with an explicit,
recorded spread parameter (for example in pips), and labelled SIMULATED with
that spread in source. A constant spread is a simplification: real spreads
widen around news and at rollover. See Market database (market.db) (Bid, ask and mid).
Reproducibility#
data/raw/EURUSD.csv is byte-identical to the output of the generator with its
default arguments (checked by comparing SHA-256 hashes). NumPy does not
guarantee that Generator streams are identical across releases, so treat the
committed CSV as the canonical v1 artefact and write regenerated data to
data/generated/ rather than overwriting it.
Generating data#
python -m trade_engine.quant.simulate_fx_gbm --years 10 --seed 42
Output goes to data/generated/<symbol>.csv by default (--output to change).
The command prints a reminder that the file is simulated.
Labelling#
Simulated data must never be mistaken for observed data. The rules are enforced, not just documented:
Every
prices_eodrow hasdata_origin(OBSERVEDorSIMULATED), a closed set checked by the database. There is no default.The loader requires
originexplicitly (--originon the CLI,originin the web API) and refuses to load anything underdata/generated/asOBSERVED.sourcerecords the generator or vendor detail, for examplegbm-seed-42.data/raw/EURUSD.csvis simulated even though it lives underraw/(that folder holds original input files). It must be loaded with--origin SIMULATED.
python scripts/load_market_data.py --symbol EURUSD --csv data/raw/EURUSD.csv --origin SIMULATED --source gbm-seed-42
GET /market shows data_origin on each bar. P&L and risk
rows do not yet carry the label, so a figure computed from simulated prices
is not marked as simulated in those tables (see
Schema hardening roadmap, item 13).
What simulated data can and cannot be used for#
Suitable
Developing and testing data loading, booking, P&L, risk and reporting.
Checking calculations against known parameters.
Demonstrating the system.
Not suitable
Any conclusion about real markets or real risk. Historical VaR on this data is simply a function of
sigma: with constant volatility and independent normal shocks, the 5th percentile of daily returns is about \(-1.645 \times 0.07/\sqrt{252} \approx -0.73\%\). Real FX has volatility clustering and fat tails, which this model lacks.Strategy backtesting or forecasting. A random walk has no exploitable structure by construction, so any apparent edge on it is overfitting or a look-ahead bug. Backtests need observed data (or a model deliberately built to contain the effect being tested), with walk-forward separation of training and evaluation periods.
Model limitations: constant volatility, no gaps, no jumps, no fat tails, no volatility clustering, no weekends or holidays, no intraday structure, and high/low generated independently of any real intraday path.
Moving to real data#
Load real prices with
--origin OBSERVEDand a--sourcenaming the vendor.Do not load them over the same symbol and dates as simulated rows. The loader skips dates that already exist for a symbol, so existing
SIMULATEDrows would silently remain. Until price versioning exists (roadmap item 9), use a different symbol or a fresh database.Re-run EOD. Results derived from observed prices are only trustworthy once the origin label propagates to P&L and risk (roadmap item 13).
Other asset classes#
Simulation is implemented for FX only. Crypto would need higher volatility, fat-tailed shocks and 24/7 calendars; bonds would need a yield-curve model with pull-to-par; equities would need dividends and corporate actions. None of that is built; see Multi-asset design.