Simulated data: why and how
===========================

v1 rolls out with a **simulated** FX pair. The market prices (and therefore every
P&L and risk figure derived from them) are generated, not observed. This document
explains why, how the data is produced, how it is labelled, and what it can and
cannot be used for.

Code: ``src/trade_engine/quant/simulate_fx_gbm.py``. Data: ``data/raw/EURUSD.csv``
(canonical v1 artefact) and ``data/generated/`` (regenerated output).

Why simulate
------------

.. list-table::
   :header-rows: 1

   * - Reason
     - Detail
   * - No data licence or feed dependency
     - v1 has no vendor connection; simulated data lets the whole front-to-back flow (load, book, P&L, risk, reports) run end to end with no external dependency.
   * - Deterministic and reproducible
     - A fixed seed yields the same history every time, so tests, demos and bug reports are repeatable.
   * - Known ground truth
     - The generating parameters are known (drift, volatility), so calculations can be sanity-checked against theory. For example, 1-day 95% VaR of a simulated return series should be close to :math:`1.645 \times \sigma/\sqrt{252}`.
   * - Safe sandbox
     - P&L, limits and risk logic can be exercised with no real positions or real money implied.
   * - Plentiful history
     - Ten years of daily bars on demand, enough for historical VaR to have a meaningful window.

Simulation is a stand-in for market data during v1, **not a claim about the
market**.

What is simulated
-----------------

Only daily OHLC prices for one FX pair. Trades are entered by hand (or from the
seed script); the sample trade prices in ``services/seeding.py`` were chosen near
the simulated series and are equally hypothetical.

The model
---------

Spot follows geometric Brownian motion with the interest-rate differential as
drift, the standard Garman-Kohlhagen setting for an FX underlying:

.. math:: dS_t = (r_d - r_f)\,S_t\,dt + \sigma\,S_t\,dW_t

The daily close path uses the exact solution, with :math:`Z_t \sim N(0,1)`:

.. math:: S_{t+1} = S_t \exp\!\left((r_d - r_f - \tfrac{1}{2}\sigma^2)\,\Delta t + \sigma\sqrt{\Delta t}\,Z_t\right)

Why GBM: it keeps the rate lognormal (never negative), ties the drift to covered
interest parity (no arbitrage), and needs one volatility parameter, which is
appropriate for a limited v1 dataset rather than a stochastic-volatility or jump
model.

.. list-table::
   :header-rows: 1

   * - Parameter
     - Default
     - Meaning
   * - ``s0``
     - 1.0850
     - Starting spot (the first bar's open)
   * - ``r_domestic``
     - 0.045
     - USD risk-free rate, annualised
   * - ``r_foreign``
     - 0.030
     - EUR risk-free rate, annualised
   * - ``sigma``
     - 0.07
     - Annualised volatility
   * - ``years``
     - 10
     - History length; :math:`\Delta t = 1/252`, so 2,520 bars
   * - ``start_date``
     - 2016-01-04
     - First bar
   * - ``seed``
     - 42
     - Random seed

How OHLC is built
-----------------

Only closes come from the model. The rest is layered on top:

- ``open[t] = close[t-1]`` (no overnight gap process); the first open is ``s0``.
- ``high`` and ``low`` widen the ``[open, close]`` range by a half-day-volatility
  fraction: ``high = max(open, close) * (1 + |z| * 0.5 * sigma * sqrt(dt))`` and
  ``low = min(open, close) * (1 - |z'| * 0.5 * sigma * sqrt(dt))``, with
  independent standard normal draws, so high and low always bracket open and
  close and ranges scale with the volatility parameter.
- Dates are business days (Monday to Friday, ``pandas.bdate_range``), with no
  holiday calendar.
- Values are rounded to 5 decimals. Rounding is monotonic, so the OHLC ordering
  survives it.

The close path is drawn from ``default_rng(seed)`` and the OHLC noise from
``default_rng(seed + 1)``, so the two are independent but both reproducible.

Price basis
~~~~~~~~~~~

The simulated series is a single path with no spread, so it represents a ``MID``
price. Bid and ask would be derived as ``mid -/+ spread / 2`` with an explicit,
recorded spread parameter (for example in pips), and labelled ``SIMULATED`` with
that spread in ``source``. A constant spread is a simplification: real spreads
widen around news and at rollover. See :doc:`/database/market` (Bid, ask and mid).

Reproducibility
---------------

``data/raw/EURUSD.csv`` is byte-identical to the output of the generator with its
default arguments (checked by comparing SHA-256 hashes). NumPy does not
guarantee that ``Generator`` streams are identical across releases, so treat the
committed CSV as the canonical v1 artefact and write regenerated data to
``data/generated/`` rather than overwriting it.

Generating data
---------------

.. code:: bash

   python -m trade_engine.quant.simulate_fx_gbm --years 10 --seed 42

Output goes to ``data/generated/<symbol>.csv`` by default (``--output`` to change).
The command prints a reminder that the file is simulated.

Labelling
---------

Simulated data must never be mistaken for observed data. The rules are enforced,
not just documented:

- Every ``prices_eod`` row has ``data_origin`` (``OBSERVED`` or ``SIMULATED``), a closed
  set checked by the database. There is no default.
- The loader requires ``origin`` explicitly (``--origin`` on the CLI, ``origin`` in the
  web API) and refuses to load anything under ``data/generated/`` as ``OBSERVED``.
- ``source`` records the generator or vendor detail, for example ``gbm-seed-42``.
- ``data/raw/EURUSD.csv`` is simulated even though it lives under ``raw/`` (that folder
  holds original *input* files). It must be loaded with ``--origin SIMULATED``.

.. code:: bash

   python scripts/load_market_data.py --symbol EURUSD --csv data/raw/EURUSD.csv --origin SIMULATED --source gbm-seed-42

``GET /market`` shows ``data_origin`` on each bar. P&L and risk
rows do **not** yet carry the label, so a figure computed from simulated prices
is not marked as simulated in those tables (see
:doc:`/database/hardening_roadmap`, item 13).

What simulated data can and cannot be used for
----------------------------------------------

**Suitable**

- Developing and testing data loading, booking, P&L, risk and reporting.
- Checking calculations against known parameters.
- Demonstrating the system.

**Not suitable**

- **Any conclusion about real markets or real risk.** Historical VaR on this
  data is simply a function of ``sigma``: with constant volatility and independent
  normal shocks, the 5th percentile of daily returns is about :math:`-1.645 \times
  0.07/\sqrt{252} \approx -0.73\%`. Real FX has volatility clustering and fat
  tails, which this model lacks.
- **Strategy backtesting or forecasting.** A random walk has no exploitable
  structure by construction, so any apparent edge on it is overfitting or a
  look-ahead bug. Backtests need observed data (or a model deliberately built to
  contain the effect being tested), with walk-forward separation of training and
  evaluation periods.

Model limitations: constant volatility, no gaps, no jumps, no fat tails, no
volatility clustering, no weekends or holidays, no intraday structure, and
high/low generated independently of any real intraday path.

Moving to real data
-------------------

1. Load real prices with ``--origin OBSERVED`` and a ``--source`` naming the vendor.
2. **Do not load them over the same symbol and dates as simulated rows.** The
   loader skips dates that already exist for a symbol, so existing ``SIMULATED``
   rows would silently remain. Until price versioning exists (roadmap item 9),
   use a different symbol or a fresh database.
3. Re-run EOD. Results derived from observed prices are only trustworthy once
   the origin label propagates to P&L and risk (roadmap item 13).

Other asset classes
-------------------

Simulation is implemented for FX only. Crypto would need higher volatility,
fat-tailed shocks and 24/7 calendars; bonds would need a yield-curve model with
pull-to-par; equities would need dividends and corporate actions. None of that is
built; see :doc:`/database/multi_asset`.
