Flint: design of a continuously learning intraday market model
1. Problem and principles
The task is short-horizon, cross-sectional return forecasting under a regime of continual change: the distribution of one-hour returns is non-stationary, the set of names worth modeling changes week to week, and the data arrives as a stream. Classical practice fits a model to a fixed history and evaluates it on a held-out period. Flint takes the opposite stance, closer to how a discretionary trader operates: a single model that is always on, scores itself on every forecast an hour after making it, and is judged primarily on that live, out-of-sample record.
Six principles follow from that stance and recur throughout the design.
- Distributions, not points. A one-hour return is dominated by noise. A forecast of five quantiles and two direction probabilities carries its own uncertainty, which the policy can act on; a point forecast cannot.
- Online, with the clock as the only split. Every label is used for training only after the forecast it scores has been recorded. There is no separate validation period; the validation is the live record plus a held-out slice of the replay.
- One network for the whole cross-section. Names move together. A network that sees all of them at once can condition each forecast on the others and on the pooled market state.
- Defined risk only. No position the system takes, on paper or in principle, may carry unlimited loss.
- Everything observable. Every feed, feature, forecast, decision and training step is traced to a console the operator can read.
- One machine, nothing leaves it. Market data comes in; keys, positions and the narrative brief's language model stay local.
2. Data
Bars and features
Sources live in a registry with a priority order, and each symbol is owned by the best source that is up and supports it; ownership fails over without operator action. Bars are built only from real trades. Quote and heartbeat ticks update the last price and the bid-ask spread but never form a candle, because a feed that echoes a stale quote while the market is closed would otherwise manufacture flat bars and the model would learn from fiction. Every modeled name needs a price for a bar row to emit, so a name without history is seeded flat at zero rather than allowed to stall the universe.
Each name contributes 25 features per bar, standardized by a running mean and variance kept per name: nine from the bars themselves (return, range, log volume, trade imbalance, spread, log trade count, close position within the bar, momentum, realized volatility), four from attention and tone (news sentiment and attention, retail-forum sentiment and attention), a fear-and-greed index, derivatives positioning (funding, open-interest change, long/short ratio), the net tilt of tracked fund books and an investor-council bias, two fundamental factors, market breadth and VIX, and at-the-money implied volatility and skew from option chains. The label for a window ending at bar t is the log return in basis points from that bar's close to the close at t + 12, capped at 2,000 bps, with a mask that removes any name that did not trade at both ends.
| Group | Feature | Definition |
|---|---|---|
| Bars | ret | log(close / previous close), bps |
range | log(high / low), bps | |
logvol | log(1 + volume) | |
imbalance | (buy volume − sell volume) / volume from tick classification | |
spread | bid-ask spread in bps from the last quote | |
logtrades | log(1 + trade count) | |
closepos | (close − low) / (high − low) | |
mom | sum of the last 12 bar returns | |
rvol | standard deviation of the last 12 bar returns (close-to-close; see Section 7) | |
| Attention | news_sent, news_attn | tone from a local FinBERT (P(positive) − P(negative) per headline, the name's own headlines with the market tone as one headline of prior) and attention as the name's share of the pass's most-covered name, from company news and market desks |
wsb_sent, wsb_attn | the same from retail-forum chatter | |
| Positioning | fng | fear-and-greed index, market-wide |
funding, oi_chg, longshort | derivatives positioning where a venue reports it | |
guru_net, ethos_bias | net tilt of the tracked 13F books; the investor council's bias for the name | |
| Fundamentals | f_value, f_quality | value and quality factor scores from company fundamentals |
| Market | breadth, vix | advancers share on the radar; the VIX level |
| Options | opt_iv, opt_skew | at-the-money implied volatility about 30 days out, and the put-call skew, from the live chain |
The replay buffer
Labeled windows are stored in a ring buffer of shape (windows × names × 128 bars × 25 features) in 32-bit floats. Three decisions about this buffer turned out to matter more than any modeling choice.
Keyed by bar time. Every restart re-walks the last five days of history to warm the feature state. Before windows were keyed, each restart appended another copy of those days: a 526-window buffer was found to hold only 180 distinct label vectors, so the most recent days were being trained on three times over. A window is now stored once per bar timestamp, and a re-walk adds nothing.
Budgeted in bytes. The buffer's capacity is a memory budget, 30 GB on the reference machine, from which the window count follows the universe: about 14,000 windows at 232 names and 7,000 at 400. Growing the universe shrinks the history held rather than the memory left for everything else.
Deep history paged in while the market is closed. A loop pages five-minute candles back in ten-day chunks, one every 150 seconds to respect the provider's rate limit, through an isolated bar and feature pipeline, until 200 days or a full buffer. Paging stops within a chunk of full rather than wrapping the ring onto its newest windows. Five percent of windows, every twentieth by timestamp, are held out of training permanently; the choice is derived from the timestamp so a reload keeps the same slice.
A self-updating universe
The modeled set is the configured core, every holding in the operator's accounts, and up to a fixed number of promoted names, under a cap of 400. Once a day after the close, candidates are drawn from three organic sources and scored on one 0–100 scale by rank within each source: the market radar's movers, weighted by dollar volume and filtered to liquid names above five dollars; tickers named in headlines, counted only when the symbol is known elsewhere; and the tracked funds' 13F positions by weight. The best rank across sources wins, with a small bonus for appearing in more than one.
Promotion fills the room left under the cap from the top of the ranking. Demotion is deliberately slow: a promoted name stays while it keeps ranking inside the top one-and-a-half times the slot count, and leaves only after 14 days without such a hit and at least five days of tenure. The asymmetry exists because a demotion forfeits that name's learned representation and its history in the replay, and a day's excitement is not worth that.
A universe change used to reset the model, which is why no such rule could run automatically. Every parameter in the network is shared across names except a table of per-name embeddings, so a change now migrates: retained names' embedding rows and their optimizer moments are copied, new names receive fresh rows, retained names' replay columns are kept, new names' columns are masked, and the deep backfill re-pages history to fill those columns in. If the saved weights do not fit the current network at all, because the size benchmark landed on a different preset, the network starts fresh but the replay is still kept and migrated; the data outlives the architecture.
3. Model
FlintNet maps a tensor of (names × 128 bars × 25 features) to, per name, five return quantiles at the 10th, 25th, 50th, 75th and 90th percentiles, a probability that the return exceeds +10 bps and a probability that it falls below −10 bps, plus a regime mixture and an attention map that are exposed for inspection. Three architectural choices carry the design.
Dilated causal convolutions over time. Each name's window is read by a stack of one-dimensional convolutions with dilations 1, 2, 4, 8, 16, 32 and 64, which gives a receptive field over the whole window at a cost linear in its length, and causality by construction.
Cross-asset attention. After the temporal stack, a multi-head attention layer runs across names within the bar, so a forecast for one name can condition on the state of the others. A learned embedding per name lets the network tell the names apart; it is the only per-name parameter, which is what makes universe migration cheap.
A regime gate over expert heads. A pooled summary of the market state selects a soft mixture over twelve expert heads, so different regimes use different heads rather than one blended predictor. The gate's weights are published with each forecast. The quantile head is non-crossing by construction: it predicts the median and non-negative increments on either side.
Objective
The loss is the pinball loss over the five quantiles plus a smoothed binary cross-entropy on each direction head, averaged over the names the mask admits:
where ρτ(u) = u(τ − 1[u < 0]) is the quantile check function, θ = 10 bps is the direction deadband, and the direction targets are smoothed by 0.1 so the network cannot be certain. Gaussian noise with standard deviation 0.1 is added to the standardized inputs during training. On CUDA the forward pass runs under bfloat16 autocast with the loss in 32-bit; the training forward is compiled, which doubles throughput on the reference hardware.
Sizing to the machine
Rather than fix a network size, a benchmark at start climbs a ladder of presets and keeps the largest whose training step fits a time budget: the warmup ceiling of 300 seconds over 60 steps, so 5,000 ms per step. Memory is checked separately, before each rung is benchmarked and once more before the start: the rung's measured peak times 1.25, plus the replay budget, plus a 3 GB headroom must fit in what is free. The estimate for the next rung scales the last rung's measured peak by width, since a training peak grows about linearly with model width; a rung whose projected step time is far over budget, or whose need exceeds the machine next to the replay, ends the ladder honestly, while a transient squeeze from another tenant refuses the start rather than caching a downgrade for good.
| Preset | Width | Params | ms/step, 232 names | Peak GB | ms/step, 362 names | Peak GB |
|---|---|---|---|---|---|---|
| XXL | 256 | 4.4M | 960 | 14.4 | 1,507 | 22.4 |
| XXXL | 384 | 11.1M | 1,800 | 25.4 | 2,801 | 39.5 |
| 4XL | 512 | 19.6M | 2,402 | 33.9 | 3,737 | 52.8 |
| 5XL | 768 | 46.3M | 3,992 | 51.0 | not benchmarked: about 110 GB next to the replay | |
4. Learning regimes
Four regimes share one network and one replay, and are never allowed to contend for the GPU at once.
- Warmup. At start, after the checkpoint and replay load and the five-day backfill, 60 optimizer steps over the replay before the engine goes live. A live forecast before this point would be noise.
- Online. Each bar whose forecasts have resolved yields up to eight steps, two per matured label, on batches of 16 drawn 30% from the newest 256 windows and 70% uniformly from the whole buffer. The recency bias is the mechanism of adaptation; the uniform share is the defense against forgetting.
- Idle passes. While the market is closed, up to two passes over the replay per closed session, a budget bounded by data rather than hours so a long night cannot grind the same windows indefinitely. Every 40 steps the held-out slice is scored, and the session's budget ends when that score has not improved over three checks. In-sample loss cannot serve this purpose: it rises whenever the deep backfill adds unseen history and falls as it is fit.
- Deep chunks. Each paged chunk of history is followed by 30 steps so the new windows enter the weights promptly.
5. Calibration
Two online corrections sit between the network and the policy, and both are updated per bar rather than per name, because a bar's several hundred labels are one observation of calibration, not several hundred: updating per name compounds into one enormous step and, with the horizon's delay, drives the correction into oscillation between its clips.
The first is an adaptive conformal scale s on the predicted band widths. The quantiles are widened or narrowed about the median, q̃τ = q0.5 + s(qτ − q0.5), and after each resolved bar the scale moves by
where c is the fraction of that bar's outcomes inside the 10–90 band and γ = 0.05, so the scale settles where one outcome in five falls outside. It starts at 2, wide, and earns its way down. The second is a temperature T on the direction logit, p̃ = σ(T · logit(p)), moved by one gradient step of the calibration cross-entropy per bar with rate 0.01, started sceptical at 0.3. Both corrections, and every per-name record the policy consults, are applied one horizon late in the backtest exactly as they are live.
6. Decision policy and risk
The policy is a composition of gates on the calibrated distribution followed by a sizing rule. It is deliberately not a learned layer: every gate is a live control, auditable, and testable in isolation, and the backtest can replay the identical gates over history. The score is a Sharpe-like ratio of the median to the calibrated interquartile range, z = q̃0.5 / (q̃0.75 − q̃0.25). A name may take a side only if all of the following hold: |z| exceeds a threshold; the tempered direction probability sits far enough from one half and agrees with the sign of z; the expected move exceeds the larger of a flat cost and the live bid-ask spread, and the spread is below a limit; the price is above a floor; the name's own record of resolved calls meets a minimum count and hit rate; and the same side has persisted for two consecutive bars. The book as a whole trades only once the model is trusted, after 48 live labels, and only while its hit-rate average beats a coin flip.
Size is fractional Kelly on the score, hard-capped, then scaled by band coverage relative to the 80% target so an overconfident model trades small:
with k = 0.15. A bearish view is a long put and a volatility view, when both direction probabilities are high and the band is wide, is a long straddle, with premium per position capped at 2% of equity; a bullish view is stock. The paper book runs 100,000 dollars with no leverage, a 15% cap per name, fills marked at the spread, a minimum hold of six bars unless the signal reverses, and flattens at the close. With extended hours enabled, stock may trade in the 4:00–9:30 and 16:00–20:00 sessions; options never do. These gates were added after one evening in which churn and a one-dollar stock cost the book seven percent, and the policy's main job since has been to hold when it has no edge.
7. Pricing models: what Flint uses and what it should borrow
Flint is an empirical forecaster, and it was a fair question whether it uses any of the standard pricing models at all. The audit for this version is in Table 3. The short answer is that it uses one, Black-Scholes, in three places, and that the rest of the canon has something to offer, mostly as features and as corrections to the policy rather than as replacements for the network.
What Flint uses
Black-Scholes-Merton marks every option position in the paper book between quotes, prices the contract selected from the live chain (the at-the-money put about 30 days out for a bearish view; the at-the-money put and call pair for a straddle), and supplies two features from the chain, the at-the-money implied volatility and the put-call skew, which are the market's own forecast of the next month's distribution. The book sizes a put to the delta-equivalent short exposure the policy asked for. Fractional Kelly sizes positions from the score. Factor scores for value and quality, and the fear-and-greed and breadth measures, are joined as features. That is the complete list; realized volatility in the feature set is the plain standard deviation of the last twelve close-to-close bar returns, and the cost of a trade is the quoted spread.
What it should borrow
| Model | Status | What it offers Flint | Where |
|---|---|---|---|
| Black-Scholes-Merton | Used | Beyond marking: the implied volatility scaled to the horizon, σ1h = IV·√(1 / (252 × 6.5)), is a market prior for the band width per name and a direct check on the conformal scale. A name whose calibrated band is far narrower or wider than the implied one is miscalibrated in a way the network cannot see. | features |
| Breeden-Litzenberger | Not used | The second derivative of the call price in strike is the risk-neutral density. Even without the full surface, the 25-delta skew and the wing prices give the market's asymmetry and tail weight for the quantile head to condition on. | features |
| Implied term structure | Not used | A front-expiry implied volatility far above the next expiry's marks a scheduled event, usually earnings. That is a veto gate: do not hold stock or sell a straddle through it; buy a straddle into it only if the model's move beats the implied one. | features, policy |
| Variance risk premium | Not used | IV² − RV² over the last month is a regime variable that predicts index returns at horizons of weeks and conditions the gate; it is two numbers the system already has. | features |
| OHLC volatility estimators (Parkinson, Garman-Klass, Rogers-Satchell, Yang-Zhang) | Not used | Range-based estimators are five to eight times as efficient as close-to-close for the same bars, and Yang-Zhang handles gaps. The current rvol and range can be replaced in place without changing the feature count. | features (in place) |
| HAR-RV, GARCH(1,1) | Not used | A heterogeneous-autoregressive forecast of realized variance from daily, weekly and monthly components gives the horizon a volatility forecast that the conformal scale only approximates after the fact. | features |
| CAPM beta | Used for the target | The hedge ratio for the market-relative target in Section 11 (a ridge loading per modeled index since 0.5.14), and a rolling beta per name as a feature, so the network sees how much of a name's move is the index. | target, features |
| Fama-French factors | Partly | Value and quality are in; the momentum (12-1) and short-term reversal (1-month) factors are cheap daily features with the best-documented premia. | features |
| Greeks on the book (theta, gamma) | In the policy since 0.5.14 | An at-the-money option about 30 days out decays by roughly premium / (2 × 30 × 6.5) per trading hour, a quarter of a percent of premium, which belongs in the cost edge for puts and straddles. A straddle's hourly profit is about ½Γ(ΔS)² − θΔt, so the entry rule should be "the model's expected |move| exceeds the implied one, σimpl·S·√Δt", in place of the present band-width threshold. | policy |
| Square-root impact (Almgren-Chriss, Gatheral) | In the policy since 0.5.14 | Cost = spread / 2 + η·σ·√(Q / ADV). At 100,000 dollars on large caps impact is negligible; for promoted small caps it is the whole edge. The cost gate should include it. | policy |
| Amihud illiquidity, Kyle's lambda | The gate since 0.5.14; not a feature | |return| per dollar of volume is a liquidity gate that works for every source, including those with no quotes, where the spread gate is blind. | policy, features |
| Roll spread estimator | Not used | The effective spread from the negative autocovariance of returns, for names whose source supplies no quotes. | policy |
| Order-flow imbalance (Cont, Kukanov, Stoikov)10 | Partly | The present imbalance feature classifies trades; the quote-size imbalance at the best bid and ask is the stronger, nearly linear predictor of the next move, available where a feed reports sizes. | features |
| Index futures basis, VIX term structure | Not used | The overnight futures move is the pre-open market factor; the futures basis against the ETF and the VX1−VX2 slope are regime variables. Needs a futures quote source. | features |
| Ornstein-Uhlenbeck, pairs | Not used | Cross-sectional residuals against the index mean-revert; a fitted half-life per name is a feature and the natural companion of the market-relative target. | features |
| Merton jump-diffusion, Heston | Not used, by choice | Their lesson, fat tails and clustered volatility, is what the quantile head and the conformal scale learn empirically; the implied surface already encodes the market's version. A parametric model here would constrain what the data should decide. | none |
| Avellaneda-Stoikov | Not applicable | Inventory-based market making; Flint takes liquidity, it does not provide it. | none |
The authors' ranking by expected value per unit of effort: first the implied-volatility horizon band, the variance risk premium and the term-structure event flag, three features that reuse data the system already fetches; second the range-based volatility estimators and a HAR forecast, which upgrade existing features in place; third theta and the gamma breakeven in the options policy and the square-root impact and Amihud gates in the cost edge, which change nothing in the network; fourth beta and the market-relative target; fifth a futures feed.
Using them together
Nothing prevents all of these from entering at once; each is a few numbers per name, and the network's job is precisely to weigh many weak inputs. Two cautions govern the batch. The inputs should be distinct: four range-based volatility estimators say nearly the same thing, so one of them (Yang-Zhang) plus the HAR components earns its place while the rest add noise and cost samples, which matter with an online learner that sees thousands of windows rather than millions. And any change to the feature set resets the network once, so the batch should land as one change, on a weekend, with the backtest run before and after it. The batch the authors would ship is in Table 4, with one addition that costs no feature at all: macro ETFs as members of the universe. Because the network attends across names, treasuries, credit, the dollar, gold, oil and a volatility product become market factors the moment they are modeled, with no feature engineering.
| Family | Features | Source |
|---|---|---|
| Option-implied | σ1h from at-the-money IV; term slope (front minus next expiry IV, the event flag); butterfly convexity; variance risk premium IV² − RV²; dealer gamma exposure, Σ open interest × gamma by strike, signed by side, which separates pinning days from trending days | the live chain, already fetched for IV and skew |
| Realized volatility | Yang-Zhang in place of rvol; Garman-Klass in place of range; HAR components: realized variance over the last day, week and month | the bars |
| Factors | rolling 60-day beta to the index; momentum 12-1; one-month reversal | the bars and the index ETF |
| Microstructure | Amihud illiquidity over 20 days; quote-size order-flow imbalance where the feed reports sizes; Roll's effective spread as the fallback where it does not | quotes and bars |
| Mechanical | an ex-dividend flag for the bar after the ex-date | the fundamentals feed |
| Macro, as names | TLT, HYG, LQD, UUP, GLD, USO, VIXY added to the core universe | no feature change; the attention layer does the rest |
Other forecasters as opinions
A pricing model gives the network numbers about the world. A second class of inputs gives it other forecasters' answers to the same question, computed causally per bar and offered as features, so the network can weigh and correct them rather than replace them. Each answers "what happens over the next hour" from a different theory: a GARCH(1,1) and a HAR forecast of how large the move should be; the option-implied distribution for the horizon, the market's own quantiles and the most important second opinion available; an Ornstein-Uhlenbeck fit on the residual against the index, giving a z-score and a half-life for whether the name is stretched and how fast it reverts; a Kalman trend filter's local drift and uncertainty; and a ridge regression and a gradient-boosted tree trained online on the same features and labels, whose disagreement with the network is itself a feature. The investor council already supplies this kind of opinion for positioning. The combiner above all of them is the ensemble of Section 11 with a meta-labeling layer deciding whether to act, which makes "more opinions" a measurable claim: the backtest scores each opinion alone and the blend. Two rules keep the layer honest: no opinion may see a bar after the one it forecasts, and opinions must disagree to be worth their sample cost.
8. Evaluation
Live scoring
Every forecast is scored when its bar resolves: pinball loss on the raw quantiles, a hit if the median's sign matched the realized sign, coverage of the calibrated band, and paper profit if a position was taken. Three loss figures are shown side by side and mean different things: train pinball is in-sample over the replay; held-out pinball is the 5% slice the trainer never samples; live pinball is out of sample on forecasts made before their labels existed and is the only one that speaks to edge.
Walk-forward backtest
Because live labels accrue only while the market is open, a live record takes weeks. The backtest compresses that into a weekend by replaying the history in the replay with a fresh network of the live preset, day by day in time order. On day D it forecasts every window of the day before it has trained on any label from the day; the live policy's gates and sizing decide; the realized hour-ahead returns score the calls; and only then does the network train on day D, with the online budget plus the idle trainer's passes over everything seen so far. The first ten days are training only; signals wait for the trust rule; calibration and per-name records lag by one horizon. The procedure is therefore out of sample by construction and reproduces the live regimen including its adaptation.
Book accounting is honest about capital. Per bar the signals form a one-hour portfolio with gross weight capped at one; the bar's return is Σi wiyi − c Σi |wi|; a day's return is the mean over its bars, because capital can fund only one hour at a time; the curve is the running sum. The report gives per-day pinball, hit rate, coverage, signals, gross and net return, an equity curve against an always-long equal-weight book, and the best and worst names. A time budget shrinks the per-day training uniformly so a run always finishes within the weekend, and the report records what it used.
8. System design
Flint is one asyncio process: a bar clock, the forecasting path, the four training regimes, and periodic loops for market status, signals, the radar, news, checkpoints, the brief, the portfolio, deep backfill, idle training, the universe refresh and the backtest. A FastAPI server exposes the state as JSON and streams it over one websocket to a dashboard of plain HTML and JavaScript with no build step. The checkpoint is written atomically every five minutes; the replay, tens of gigabytes when full, is written uncompressed as one memory-mapped array per field at most every 30 minutes and at shutdown, swapped into place whole so a crash leaves the previous directory intact.
The reference machine's unified memory is shared by the model, the replay and everything else, and its GPU allocations are invisible to process accounting and outside any control-group limit. Several invariants exist because this was learned the hard way.
- A backtest owns the GPU. While one runs, matured labels are stored and trained on afterwards, the idle trainer skips, deep paging pauses so the replay cannot grow, and a manually requested run waits for paging to finish. The backtest sizes its network to the memory actually free by scaling the live rung's measured peak by width, and stops itself at the next day boundary if free memory falls below a floor.
- The caching allocator is bounded. The allocator keeps blocks per shape, so the training step, the validation pass and per-bar prediction each held their own working set. Validation now runs at small batch and releases the cache after; the idle trainer and the deep loop release it every 40 steps; the allocator runs with expandable segments.
- A benchmark measures only its own peak, above whatever the process already holds, and a universe rebuild stops a running backtest first. A benchmark that ran beside a backtest once read the backtest's working set as its own and refused.
- A failed rebuild never leaves the engine dead. It saves, restores the full list the model was last built for, and resumes the lifecycle on the model it still has.
- A compile failure costs speed, not the engine. The first compiled training step is where the compiler really compiles; a toolchain problem there logs its cause once and the learner trains eagerly thereafter. An engine that died at this point while its server kept answering was the first incident of the day.
- Shutdown must complete. The server's graceful shutdown waits for every websocket handler to return and cancels nothing, so a push-only handler held the termination signal until the service manager killed the process and the final save never ran. Handlers now read their sockets, and the graceful wait is capped.
- Keys apply live and are verified across the board. A key saved on the dashboard is written to the file the configuration reads on the next start and pushed into every consumer that copied it at construction, and one test walks every service through save, status, restart read-back and live apply.
9. Limitations and threats to validity
- Small effects. The intraday regularities the literature documents explain about two percent of variance at best1,2, and daily-frequency deep models have generally failed to beat buy-and-hold after costs9. Edges of a few basis points are the expectation, and the backtest's cost assumption of five basis points a round trip may be generous for the less liquid promoted names.
- Selection in the backtest universe. Names promoted because they moved have their past windows filled by the re-page, so a backtest over the current universe includes names chosen with knowledge of their later behavior. A clean test freezes the universe at each day's date; this is not yet done.
- Feature state at ingest. Paged windows are standardized with the normalization state at the time of paging, not the state that prevailed historically.
- Execution. Paper fills are marked at the quoted spread with no market impact, no partial fills and no queue position; the extended-hours sessions in particular are thinner than the quotes suggest.
- One machine, one process. The design trades fault isolation for simplicity; a bug in any loop can take the whole engine down, which is why the invariants in Section 9 carry tests.
- No causal claims. Nothing here identifies why a forecast works when it does; the attention map and the regime gate are diagnostics, not explanations.
10. Future work
Ranked by expected effect per unit of effort, each item is a hypothesis for the walk-forward backtest.
- A market-relative target (shipped in 0.5.14 as a switch, judged by the next backtest). Over one hour the index dominates every name's move and nothing in a five-minute history predicts the index. Nonlinear models beat linear ones mainly when the target is the market-relative abnormal return4. The label becomes the residual of a ridge regression of the name's return on the modeled index names' returns over the same hour (SPY, QQQ, IWM, SMH and XLF, so a semiconductor name is measured against SMH and a small cap against IWM, not everything against SPY), loadings refit once a session and shrunk toward the broad market; the replay keeps absolute labels, so the target is a run-time switch and one replay serves both. The book holds, per index, the position that cancels the basket's loading: an index put under a long basket, index stock over a bearish one. The weekend backtest runs both targets back to back on the same history and reports them side by side.
- The pricing-model pass of Section 7, in its stated order: implied-volatility horizon band, variance risk premium and the term-structure event flag; range-based volatility and a HAR forecast; theta, the gamma breakeven, square-root impact and the Amihud gate in the policy.
- Calendar and daily-lag features. A name's return in a given half-hour continues at the same half-hour on following days for about forty days, strongest at the open and close1; the first half-hour's index return predicts the last half-hour's2; overnight and intraday components carry opposite, persistent signals3. None are features yet.
- An ensemble of mid-size networks. Tens of millions of parameters on some thousands of windows is heavily over-parameterized, as the overfit guard's early stops confirm. Two 4XL networks fit the machine, and aggregating ensemble members' distributions improves accuracy and calibration8.
- Meta-labeling. A second, small classifier trained on whether acting on each primary forecast paid after costs, sizing by its probability5,6: the evaluable form of a decision layer, trainable online from the outcome records Flint already keeps.
- Calibration to decision loss net of costs7, sizing by the probability that the move clears the spread from the full quantile set rather than the median alone; and exits with a profit target, a stop, a time limit and a daily loss limit.
- A policy sweep over stored forecasts (shipped in 0.5.14). Every forecast a backtest makes is stored, and a vectorized replay of the policy over them reproduces the run's net to the basis point in under a second, so the gates can be swept: a coordinate descent on the first half of the scored days, confirmed on the second, with the result proposed on the dashboard rather than applied.
- SEC filings beyond 13F. Form 4 insider transactions, 13D and 13G stakes and 8-K item codes as direct features; free-text filings summarized by a local 7–8B language model into a fixed schema of direction, materiality, horizon and topic flags, one cached call per filing; and notable filings as a fourth candidate source for the universe.
References
- S. L. Heston, R. A. Korajczyk and R. Sadka, "Intraday Patterns in the Cross-section of Stock Returns," Journal of Finance 65(4), 2010. PDF
- L. Gao, Y. Han, S. Z. Li and G. Zhou, "Market Intraday Momentum," Journal of Financial Economics, 2018. Summary
- D. Lou, C. Polk and S. Skouras, "A Tug of War: Overnight versus Intraday Expected Returns," Journal of Financial Economics, 2019. PDF
- "Design Choices in Machine Learning for Stock Returns," review of target and model choices. Article
- M. López de Prado, Advances in Financial Machine Learning, Wiley, 2018: triple-barrier labeling and meta-labeling. Overview
- Hudson & Thames, "Does Meta-Labeling Add to Signal Efficacy?" Article
- "Utility-Weighted Forecasting and Calibration for Risk-Adjusted Decisions under Trading Frictions," arXiv:2601.07852, 2026. Abstract
- Aggregation of deep-ensemble predictive distributions, Machine Learning, 2025. Record
- Transformer-based deep learning for stock return forecasting, with transaction costs. PDF
- R. Cont, A. Kukanov and S. Stoikov, "The Price Impact of Order Book Events," Journal of Financial Econometrics 12(1), 2014. arXiv
Appendix: principal settings
| Setting | Default | Role |
|---|---|---|
bar_seconds / horizon / window | 300 / 12 / 128 | five-minute bars, one hour ahead, 128 bars of context |
quantiles / direction_threshold_bps | 10, 25, 50, 75, 90 / 10 | the distributional head and its direction deadband |
lr / weight_decay / dropout / input_noise | 5e-4 / 1e-2 / 0.15 / 0.1 | optimizer and regularization |
batch_size / recent_frac / recent_n | 16 / 0.3 / 256 | recency-biased sampling |
steps_per_label / warmup_steps / max_warmup_seconds | 2 / 60 / 300 | online budget; the size ladder's time budget |
idle_train_epochs / overfit_patience (live) | 2 / 3 | closed-session passes and the held-out stop |
deep_backfill_days (live) / replay_gb / replay_save_minutes | 200 / 30 / 30 | history paged, buffer budget, replay write cadence |
max_universe / interesting_count (live) / interesting_grace_days / interesting_min_tenure_days | 400 / 150 / 14 / 5 | the universe and its promotion and demotion rule |
min_labels / min_hit_rate (live) / skill_min_n (live) | 48 / 0.5 / 8 | trust, and the per-name record |
score_threshold / prob_margin / move_floor_bps (live) | 0.35 / 0.06 / 8 | the entry gates |
max_spread_bps / min_price / confirm_bars / stock_min_hold_bars (live) | 25 / 5 / 2 / 6 | liquidity and persistence gates |
kelly_fraction / max_size / option_max_frac (live) | 0.15 / 1.0 / 0.02 | sizing |
band_gamma / temper_lr | 0.05 / 0.01 | the two calibration rates |
backtest_hours / backtest_night_epochs / backtest_cost_bps / backtest_warm_days (live) | 30 / 2 / 5 / 10 | the walk-forward run |
memory_floor_gb (live) | 10 | the floor every GPU tenant keeps free |