Adam Burich Open to work

← Portfolio

2026 · Ongoing · Paper-traded research

Can reasoning catch a news story's second-order effect before the market does?

Solo · Since July 2026 · Runs nightly · Forward-only, never backtested

When a policy story breaks, the obvious names get repriced almost instantly. The knock-on effects — whose costs quietly go up, whose competitor just got a gift, whose supply chain tightened — seem to take longer for the market to work out. I don't know yet whether that gap is real and catchable. This is a pre-registered experiment built to find out without fooling myself.

Where the edge would have to live

The framing is a speed barbell. High-frequency traders own the first-order repricing in milliseconds; institutions take days to weeks to reposition. If there's room for judgment at all, it's in the middle — hours to weeks, one step removed from the headline. So every prediction has to name its payer: who is on the other side of the trade, and why they're stuck — forced to act, slow to react, or trading on emotion. Often I can't name one, and the idea gets passed on.

One example: a scary chip-export headline had people selling crypto miners. The thought going the other way was that fewer imported rigs makes the network less competitive, which helps the machines already running. That became the first prediction on the ledger — and whether it was right is exactly the kind of question one data point can't answer.

01Why forward-only

The reasoner is an LLM. On any past event it may already know how things turned out, so a backtest would measure recall, not judgment. The only honest evidence comes from predictions written on news the model has never seen resolved — which means the experiment can only run forward, in real time, and has to be slow. Prediction timestamps against the model's knowledge cutoff make that auditable.

02The nightly loop

  1. Capture the newsA standard-library script pulls policy headlines from free RSS in three tiers — editorial sections, press wires behind a strict filter, and targeted news searches standing in for the paid terminals — then clusters them into stories ranked by how many outlets carry them. A typical night: 707 raw items, 350 policy leads, 287 distinct stories. Each night's snapshot is committed to git as the record of exactly what the reasoner saw.
  2. Reason, price-blindClaude reads only the snapshot. It may research company and policy facts, but may never look at price reactions before writing a prediction. A preflight records beta, typical daily range, and the next earnings date — without showing prices.
  3. Challenge itselfBefore logging a prediction or a pass, a mandatory self-challenge: is my blocker ignorance or impossibility? What's the strongest opposite case? Is this still fresh?
  4. Review what's openEach open thesis is rated intact, strengthened, weakened, or broken against the new news, and each position is called hold, exit, or roll.
  5. Execute and scoreAn evening routine records decisions and exits at the close; orders meant for the open are queued and filled by a morning routine at 9:32 ET, into a paper-broker ledger. Matured predictions are scored weekly.

03What a prediction looks like

Every prediction is one append-only JSON line, written before the outcome is knowable and never edited or deleted:

id, ts, event_id, headline
thesis          the second-order chain, 1–3 sentences
payer           forced | slow | emotional — and who, specifically
ticker, direction, horizon_days, confidence (1–3)
size_pct        risk-normalized by the name's typical daily range
beta_est, est_range_pct, news_age_hours, earnings_date
action          "BUY X at the open …; no stop; SELL at the close of the Nth session"

04Rules fixed before the first prediction

The hard part isn't the ideas, it's not fooling myself. It's easy to remember the guesses that worked and quietly forget the rest, so the rules were written down first:

05Amendments, in the open

Rules change as the experiment teaches me things, but every change is versioned, dated, and never retrofitted onto predictions that are already open. The protocol is on version 1.7. Among the amendments: a mandatory exit when news breaks a thesis, with every early exit also scored as if it had been held to the original horizon, so the exit option itself gets measured; a pre-set profit ceiling scaled to each name's range; and holding idle cash in the index with a separate, untouched index account opened on day one as the control, so the comparison isolates the picks.

When two amendments reached the automated routines ten days late, the lag was disclosed and reconciled in the ledger rather than quietly backdated.

06Where it stands

As of September 2026 there are eleven predictions on the ledger, against a bar of thirty. That's not enough to say anything, and the write-up deliberately doesn't. The one early lesson worth naming is about the exit option: the first discretionary early exit, scored against its original horizon, cost more than it saved — which is exactly why every early exit is scored both ways.

07The groundwork: an intraday data pipeline

Before the news experiment, the same repo started as a study of intraday market structure: six months of Nasdaq TotalView-ITCH one-minute bars from Databento — about 24,000 symbols — normalized into date-partitioned Parquet and queried with Polars, with a locked holdout half that wasn't touched until the end. Across fourteen experiments the conclusion was that intraday direction is close to a coin flip, while magnitude and structure are predictable. That result is part of why the news work targets reasoning over days rather than prediction over minutes, and the range-based position sizing comes directly from it.

Stack

Python · Claude (nightly and morning cloud routines) · RSS ingestion (standard library) · Yahoo daily prices · paper-broker · Polars · Parquet · Databento · scikit-learn

A personal research experiment and paper-trading exercise. Not investment advice, not a service, and no real orders are placed.