Simulated data: why and how#

v1 rolls out with a simulated FX pair. The market prices (and therefore every P&L and risk figure derived from them) are generated, not observed. This document explains why, how the data is produced, how it is labelled, and what it can and cannot be used for.

Code: src/trade_engine/quant/simulate_fx_gbm.py. Data: data/raw/EURUSD.csv (canonical v1 artefact) and data/generated/ (regenerated output).

Why simulate#

Reason

Detail

No data licence or feed dependency

v1 has no vendor connection; simulated data lets the whole front-to-back flow (load, book, P&L, risk, reports) run end to end with no external dependency.

Deterministic and reproducible

A fixed seed yields the same history every time, so tests, demos and bug reports are repeatable.

Known ground truth

The generating parameters are known (drift, volatility), so calculations can be sanity-checked against theory. For example, 1-day 95% VaR of a simulated return series should be close to \(1.645 \times \sigma/\sqrt{252}\).

Safe sandbox

P&L, limits and risk logic can be exercised with no real positions or real money implied.

Plentiful history

Ten years of daily bars on demand, enough for historical VaR to have a meaningful window.

Simulation is a stand-in for market data during v1, not a claim about the market.

What is simulated#

Only daily OHLC prices for one FX pair. Trades are entered by hand (or from the seed script); the sample trade prices in services/seeding.py were chosen near the simulated series and are equally hypothetical.

The model#

Spot follows geometric Brownian motion with the interest-rate differential as drift, the standard Garman-Kohlhagen setting for an FX underlying:

\[dS_t = (r_d - r_f)\,S_t\,dt + \sigma\,S_t\,dW_t\]

The daily close path uses the exact solution, with \(Z_t \sim N(0,1)\):

\[S_{t+1} = S_t \exp\!\left((r_d - r_f - \tfrac{1}{2}\sigma^2)\,\Delta t + \sigma\sqrt{\Delta t}\,Z_t\right)\]

Why GBM: it keeps the rate lognormal (never negative), ties the drift to covered interest parity (no arbitrage), and needs one volatility parameter, which is appropriate for a limited v1 dataset rather than a stochastic-volatility or jump model.

Parameter

Default

Meaning

s0

1.0850

Starting spot (the first bar’s open)

r_domestic

0.045

USD risk-free rate, annualised

r_foreign

0.030

EUR risk-free rate, annualised

sigma

0.07

Annualised volatility

years

10

History length; \(\Delta t = 1/252\), so 2,520 bars

start_date

2016-01-04

First bar

seed

42

Random seed

How OHLC is built#

Only closes come from the model. The rest is layered on top:

  • open[t] = close[t-1] (no overnight gap process); the first open is s0.

  • high and low widen the [open, close] range by a half-day-volatility fraction: high = max(open, close) * (1 + |z| * 0.5 * sigma * sqrt(dt)) and low = min(open, close) * (1 - |z'| * 0.5 * sigma * sqrt(dt)), with independent standard normal draws, so high and low always bracket open and close and ranges scale with the volatility parameter.

  • Dates are business days (Monday to Friday, pandas.bdate_range), with no holiday calendar.

  • Values are rounded to 5 decimals. Rounding is monotonic, so the OHLC ordering survives it.

The close path is drawn from default_rng(seed) and the OHLC noise from default_rng(seed + 1), so the two are independent but both reproducible.

Price basis#

The simulated series is a single path with no spread, so it represents a MID price. Bid and ask would be derived as mid -/+ spread / 2 with an explicit, recorded spread parameter (for example in pips), and labelled SIMULATED with that spread in source. A constant spread is a simplification: real spreads widen around news and at rollover. See Market database (market.db) (Bid, ask and mid).

Reproducibility#

data/raw/EURUSD.csv is byte-identical to the output of the generator with its default arguments (checked by comparing SHA-256 hashes). NumPy does not guarantee that Generator streams are identical across releases, so treat the committed CSV as the canonical v1 artefact and write regenerated data to data/generated/ rather than overwriting it.

Generating data#

python -m trade_engine.quant.simulate_fx_gbm --years 10 --seed 42

Output goes to data/generated/<symbol>.csv by default (--output to change). The command prints a reminder that the file is simulated.

Labelling#

Simulated data must never be mistaken for observed data. The rules are enforced, not just documented:

  • Every prices_eod row has data_origin (OBSERVED or SIMULATED), a closed set checked by the database. There is no default.

  • The loader requires origin explicitly (--origin on the CLI, origin in the web API) and refuses to load anything under data/generated/ as OBSERVED.

  • source records the generator or vendor detail, for example gbm-seed-42.

  • data/raw/EURUSD.csv is simulated even though it lives under raw/ (that folder holds original input files). It must be loaded with --origin SIMULATED.

python scripts/load_market_data.py --symbol EURUSD --csv data/raw/EURUSD.csv --origin SIMULATED --source gbm-seed-42

GET /market shows data_origin on each bar. P&L and risk rows do not yet carry the label, so a figure computed from simulated prices is not marked as simulated in those tables (see Schema hardening roadmap, item 13).

What simulated data can and cannot be used for#

Suitable

  • Developing and testing data loading, booking, P&L, risk and reporting.

  • Checking calculations against known parameters.

  • Demonstrating the system.

Not suitable

  • Any conclusion about real markets or real risk. Historical VaR on this data is simply a function of sigma: with constant volatility and independent normal shocks, the 5th percentile of daily returns is about \(-1.645 \times 0.07/\sqrt{252} \approx -0.73\%\). Real FX has volatility clustering and fat tails, which this model lacks.

  • Strategy backtesting or forecasting. A random walk has no exploitable structure by construction, so any apparent edge on it is overfitting or a look-ahead bug. Backtests need observed data (or a model deliberately built to contain the effect being tested), with walk-forward separation of training and evaluation periods.

Model limitations: constant volatility, no gaps, no jumps, no fat tails, no volatility clustering, no weekends or holidays, no intraday structure, and high/low generated independently of any real intraday path.

Moving to real data#

  1. Load real prices with --origin OBSERVED and a --source naming the vendor.

  2. Do not load them over the same symbol and dates as simulated rows. The loader skips dates that already exist for a symbol, so existing SIMULATED rows would silently remain. Until price versioning exists (roadmap item 9), use a different symbol or a fresh database.

  3. Re-run EOD. Results derived from observed prices are only trustworthy once the origin label propagates to P&L and risk (roadmap item 13).

Other asset classes#

Simulation is implemented for FX only. Crypto would need higher volatility, fat-tailed shocks and 24/7 calendars; bonds would need a yield-curve model with pull-to-par; equities would need dividends and corporate actions. None of that is built; see Multi-asset design.