Multi-Source Financial Intelligence
A planned agentic system that reads official documents and the informal commentary around them, produces a research report and a point-in-time feature panel, feeds the Event-Driven Market Probability Engine, and is judged by a pre-registered backtest before going live.
The idea
Two kinds of source speak about the same event. Official documents state policy; journalists, forums and interviews say what people made of it. Read both, and the gap between them becomes measurable.
The system turns both into structured facts, builds a daily feature panel and a readable report from them, and hands the panel to a probability model that decides whether any of it is worth acting on.
The premise
"An agent that researches a topic and writes a report" is the 2026 equivalent of a React todo app. What makes this one worth building is that its output has to survive a pre-registered backtest, so there is a number at the end that cannot be argued with.
Planned architecture
A data pipeline with agents inside it, not an agent system with data attached. Dagster owns the graph; LangGraph owns only the extraction subgraph, which is the one place agency earns its keep.
The layers, in one line each
| Layer | What it does |
|---|---|
| L0 · Adapters | A filing, an article, a Reddit post and a transcribed interview all arrive in the same envelope, anchored on publication time and tagged with a source tier from regulatory down to social. |
| L1 · Normalisation | Collapse near-identical coverage into story clusters, resolve entities and speakers, and map every timestamp to the session it can actually be used on. |
| L2 · Agent tier | A coordinator routes each document to a numeric, language, change or claim agent. Structured output only, with low-confidence fields escalated and then queued for review rather than written as a guess. |
| L3 · Fact store | One row per extracted stance: entity, topic, time, source tier, confidence, and the quote it came from. |
| L4 · Features | A trailing daily panel: stance, volume, disagreement across sources, topic shift, novelty, and the gap between officials and the crowd. |
| L5 · Report | A deterministic template over the fact store. Prose is generated, but the report cannot state a number the warehouse does not hold. |
| L6 · EDMP integration | Two CSVs, one for events and one for narrative features. This system emits files rather than writing to the engine's database. |
| L7 · Backtest | Price features alone against price plus narrative, walk-forward, pre-registered before it runs. |
What the report shows
A stance grid, each topic tracked across recent meetings, so the interesting cell is the one that changed. Summarising a single statement is a commodity; noticing that the Committee dropped "patient" needs a memory per entity.
| FOMC | Jan | Mar | May | Jun |
|---|---|---|---|---|
| hawkish | ↑ | → | ↓ | ↓ |
| inflation | ↑ | ↑ | → | ↓ |
| labour | → | → | ↑ | ↑ |
| guidance | → | ↓ | ↓ | → |
The feature worth building it for
The gap between what policymakers say and what the crowd believes: official and wire coverage on one side, social on the other, same topic, same day. Nothing in the existing engine can express it, and it is only meaningful if the deduplication underneath is correct.
Alongside it: trailing stance, cluster volume, disagreement across sources, topic shift, novelty, and the futures-implied policy surprise. Longest window is 30 days, inside the engine's 60-day embargo.
How it gets judged
Three gates, in order, so a failure is attributable to one layer rather than to the whole thing.
1 · Is the extraction right?
Numbers scored against the authoritative print. Language scored against a hand-labelled sample, plus agreement across repeated runs and between two different models.
2 · Are the features honest?
Shuffle the features within a date and accuracy must collapse. Shift them forward one day and it should improve; if the unshifted version already matches, there is leakage.
3 · Is it worth anything?
Feature set, label and test committed first, then run once. Walk-forward, with per-fold results reported rather than the mean alone.
A null result is the expected outcome, and the backtest is not the end
An FOMC statement is repriced within seconds, so a next-day label asks the features to predict something the market absorbed the previous afternoon. With eight meetings a year and around ten candidate features, something will look significant on chance alone, which is why the test is committed before it runs.
That is also why the plan does not stop there. The same panel is meant to run live, scoring each document as it is published rather than waiting for the next close, on a horizon short enough that the information is still in play. The offline study exists to decide what is worth putting live.
Planned stack
| Layer | Choice | Why |
|---|---|---|
| Orchestration | Dagster | Date partitions and backfills, which point-in-time reprocessing needs and a Makefile cannot express. |
| Agents | LangGraph | The extraction subgraph only, wrapped as one step of the pipeline. |
| Warehouse | Postgres + pgvector | Embeddings live beside the facts rather than in a second system. |
| Models | Claude + GLM / Qwen | One interface, chosen per task, so cost and agreement can be compared instead of guessed at. |
| Speech | WhisperX | Interviews and press conferences. Speaker separation is what tells the Chair from the reporter. |
| Modelling | scikit-learn, then LightGBM | Matches the engine's baseline; trees for the interaction terms. |
| Report | Jinja2 → WeasyPrint | Deterministic template, traceable numbers. |
Build order
Phase 0 · Foundations
Pipeline skeleton, shared document envelope, adapter base class, deduplication, entity resolution, the trading-date function, the model interface and the review queue. Built generically, but every piece is exercised by Phase 1.
Phase 1 · Macro, end to end
Macro only, on the 15 ETFs the engine already covers. Fed, BLS, futures, one news API, Reddit and one transcription path, then the fact store, the feature panel, one report per FOMC cycle, and the three gates in order. Not done until there is a rendered report and a backtest number.
Phase 2 · Live
Run the same pipeline forward in real time, scoring documents as they are published. After that the corporate vertical reuses everything above and adds only new adapters and topic seeds.