William Caron-Bastarache
Personal Project · Planning

Multi-Source Financial Intelligence

A planned agentic system that reads official documents and the informal commentary around them, produces a research report and a point-in-time feature panel, feeds the Event-Driven Market Probability Engine, and is judged by a pre-registered backtest before going live.

The idea

Two kinds of source speak about the same event. Official documents state policy; journalists, forums and interviews say what people made of it. Read both, and the gap between them becomes measurable.

The system turns both into structured facts, builds a daily feature panel and a readable report from them, and hands the panel to a probability model that decides whether any of it is worth acting on.

The premise

"An agent that researches a topic and writes a report" is the 2026 equivalent of a React todo app. What makes this one worth building is that its output has to survive a pre-registered backtest, so there is a number at the end that cannot be argued with.

Planning. Architecture and build order set.
Dagster, LangGraph, PostgreSQL + pgvector, Claude and open models behind one interface

Planned architecture

OFFICIAL FOMC statements · minutes · speeches SEP + dot plot · Beige Book BLS · BEA · fed funds futures INFORMAL Fed-beat journalism · interviews r/economics · r/investing · StockTwits newsletters · econ commentary L0 · Source adapters every source arrives in the same shape L1 · Normalisation story-cluster dedup · entity resolution effective trading date · source tiering L2 · Agent tier (LangGraph) coordinator → numeric · language · change · claim structured outputs · confidence routing · review queue L3 · Fact store entity × topic × time × source L4 · Features point-in-time daily panel L5 · Report HTML → PDF, every figure traceable L6 · EDMP integration events CSV + narrative feature columns L7 · Backtest pre-registered lift study

A data pipeline with agents inside it, not an agent system with data attached. Dagster owns the graph; LangGraph owns only the extraction subgraph, which is the one place agency earns its keep.

The layers, in one line each

LayerWhat it does
L0 · AdaptersA filing, an article, a Reddit post and a transcribed interview all arrive in the same envelope, anchored on publication time and tagged with a source tier from regulatory down to social.
L1 · NormalisationCollapse near-identical coverage into story clusters, resolve entities and speakers, and map every timestamp to the session it can actually be used on.
L2 · Agent tierA coordinator routes each document to a numeric, language, change or claim agent. Structured output only, with low-confidence fields escalated and then queued for review rather than written as a guess.
L3 · Fact storeOne row per extracted stance: entity, topic, time, source tier, confidence, and the quote it came from.
L4 · FeaturesA trailing daily panel: stance, volume, disagreement across sources, topic shift, novelty, and the gap between officials and the crowd.
L5 · ReportA deterministic template over the fact store. Prose is generated, but the report cannot state a number the warehouse does not hold.
L6 · EDMP integrationTwo CSVs, one for events and one for narrative features. This system emits files rather than writing to the engine's database.
L7 · BacktestPrice features alone against price plus narrative, walk-forward, pre-registered before it runs.

What the report shows

A stance grid, each topic tracked across recent meetings, so the interesting cell is the one that changed. Summarising a single statement is a commodity; noticing that the Committee dropped "patient" needs a memory per entity.

FOMCJanMarMayJun
hawkish↑→↓↓
inflation↑↑→↓
labour→→↑↑
guidance→↓↓→

The feature worth building it for

The gap between what policymakers say and what the crowd believes: official and wire coverage on one side, social on the other, same topic, same day. Nothing in the existing engine can express it, and it is only meaningful if the deduplication underneath is correct.

Alongside it: trailing stance, cluster volume, disagreement across sources, topic shift, novelty, and the futures-implied policy surprise. Longest window is 30 days, inside the engine's 60-day embargo.

How it gets judged

Three gates, in order, so a failure is attributable to one layer rather than to the whole thing.

1 · Is the extraction right?

Numbers scored against the authoritative print. Language scored against a hand-labelled sample, plus agreement across repeated runs and between two different models.

2 · Are the features honest?

Shuffle the features within a date and accuracy must collapse. Shift them forward one day and it should improve; if the unshifted version already matches, there is leakage.

3 · Is it worth anything?

Feature set, label and test committed first, then run once. Walk-forward, with per-fold results reported rather than the mean alone.

A null result is the expected outcome, and the backtest is not the end

An FOMC statement is repriced within seconds, so a next-day label asks the features to predict something the market absorbed the previous afternoon. With eight meetings a year and around ten candidate features, something will look significant on chance alone, which is why the test is committed before it runs.

That is also why the plan does not stop there. The same panel is meant to run live, scoring each document as it is published rather than waiting for the next close, on a horizon short enough that the information is still in play. The offline study exists to decide what is worth putting live.

Planned stack

LayerChoiceWhy
OrchestrationDagsterDate partitions and backfills, which point-in-time reprocessing needs and a Makefile cannot express.
AgentsLangGraphThe extraction subgraph only, wrapped as one step of the pipeline.
WarehousePostgres + pgvectorEmbeddings live beside the facts rather than in a second system.
ModelsClaude + GLM / QwenOne interface, chosen per task, so cost and agreement can be compared instead of guessed at.
SpeechWhisperXInterviews and press conferences. Speaker separation is what tells the Chair from the reporter.
Modellingscikit-learn, then LightGBMMatches the engine's baseline; trees for the interaction terms.
ReportJinja2 → WeasyPrintDeterministic template, traceable numbers.

Build order

Phase 0 · Foundations

Pipeline skeleton, shared document envelope, adapter base class, deduplication, entity resolution, the trading-date function, the model interface and the review queue. Built generically, but every piece is exercised by Phase 1.

Phase 1 · Macro, end to end

Macro only, on the 15 ETFs the engine already covers. Fed, BLS, futures, one news API, Reddit and one transcription path, then the fact store, the feature panel, one report per FOMC cycle, and the three gates in order. Not done until there is a rendered report and a backtest number.

Phase 2 · Live

Run the same pipeline forward in real time, scoring documents as they are published. After that the corporate vertical reuses everything above and adds only new adapters and topic seeds.

Contact Me

Interested in this project? Write me.