William Caron-Bastarache
Personal Project · Phases A–D complete

Event-Driven Market Probability Engine

A reproducible research pipeline that estimates next-day market move probabilities, validates that estimate under conditions that do not flatter it, and reports the honest conclusion: the model has no tradeable edge.

Summary

A layered PostgreSQL warehouse and a walk-forward-validated probability model, built to answer one question honestly: is there a tradeable edge in next-day ETF moves? The measured answer is no, and reporting that is the point. Most of the engineering exists to make the verdict trustworthy rather than to make it flattering.

Architecture, data engineering, modeling, evaluation
PostgreSQL, Python (asyncio, pandas, psycopg, scikit-learn, matplotlib), SQL window functions, Make, pytest
15 ETFs · 32,295 rows · 2018–2026 · 36 tests · 6 database-level invariants
Phases A–D complete. Calibration next, then event ingestion.

The result

Five expanding walk-forward folds with a 60-day embargo, then a backtest with transaction costs, benchmarked against simply holding the market. No strategy beats the benchmark.

Cumulative return by strategy, each fold backtested separately, against an always-long benchmark

Sharpe: 1.48 holding the market, 1.22 for the volatility filter, 0.66 for the directional rule.

What it demonstrates

Point-in-time correctness

Features may only read data at or before t, using trailing-only window frames; forward-looking values live in a separate table. Assertions recompute each stored value from its definition and fail the build when the two disagree.

Honest validation

Walk-forward expanding folds, never a random split, with a 60-day purge/embargo matched to the longest feature lookback so no test row's rolling window reaches into training data.

Reporting a negative result

Every strategy scored against a hold-the-market benchmark and a naive majority-class baseline. The finding is that the model has no tradeable edge, and it is published rather than buried.

Testing for silent failure

Temporal leakage raises no error and changes no row count; it just quietly inflates the metrics. 36 pytest cases plus database-level invariants, each verified to fail by deliberately breaking it.

Concurrent ingestion

Prices are fetched concurrently with asyncio and a bounded semaphore. That introduced two silent failure modes, both now guarded by cross-asset validations rather than per-row checks.

Reproducible by construction

Stage order is declared as Make dependencies and every stage recomputes its own outputs, so rebuilding from scratch is the normal path. Every figure is generated from the warehouse, never exported by hand.

The finding worth reporting

A real signal that still loses money

Next-day direction has no signal: ROC-AUC ~0.51, straddling chance, exactly as market efficiency predicts for price-only features on liquid ETFs. Large moves are genuinely predictable at ~0.57, in 14 of 15 instruments.

And it still does not convert. The label is unsigned: it knows a big move is coming but not which way, so stepping aside forgoes as much upside as downside. Being right about volatility is not the same as having an edge, and only a backtest could have shown that.

Calibration: ranking is not probability

Reliability diagram: predicted probability against realised frequency, for direction and large moves

The model sorts days correctly but exaggerates the gap: it claims its most turbulent tenth is 5.1x riskier than its calmest, where only 2.6x occurred. A threshold rule needs the ranking to be right; a sizing rule needs the value to be right. This model has the first and not the second.

Architecture

Layered warehouse

Data flows one way through three schemas: raw as ingested, staging cleaned and re-keyed onto surrogate ids, analytics for features, labels, predictions and backtest results. A validation gate between raw and staging fails the build on bad input before it can reach a feature.

Next: events, from an agent system

The event layer is not built yet, deliberately. An event feature is only worth something if you can measure what it adds, and that needs an honest baseline to measure against. This is that baseline. Events are intended to come from a separate agent system extracting structured, timestamped facts from FOMC statements and macro releases.

Contact Me

Interested in this project? Write me.