Whitepaper · version 0.5.13 · October 2026

Flint: design of a continuously learning intraday market model

Abstract Flint is a single-process, self-hosted system that forecasts the distribution of each stock's return over the next hour on five-minute bars for several hundred names at once, trains online as those forecasts resolve, and acts on the forecasts through a defined-risk paper book. This paper records the design decisions and the reasons behind them. We argue for distributional outputs with online conformal calibration over point forecasts; for a single cross-sectional network with a regime gate rather than one model per name; for a replay buffer keyed by bar time that survives restarts and universe changes; for a universe that maintains itself through a promotion and demotion rule; for a policy composed of auditable gates rather than a learned decision layer; and for a walk-forward backtest that is out of sample by construction and runs on the weekend GPU. We also document the engineering invariants that a shared-memory GPU imposes on such a system, several of which were learned by taking the machine down. The question of whether the forecasts carry economic edge after costs is left open and is precisely what the backtest exists to answer.

1. Problem and principles

The task is short-horizon, cross-sectional return forecasting under a regime of continual change: the distribution of one-hour returns is non-stationary, the set of names worth modeling changes week to week, and the data arrives as a stream. Classical practice fits a model to a fixed history and evaluates it on a held-out period. Flint takes the opposite stance, closer to how a discretionary trader operates: a single model that is always on, scores itself on every forecast an hour after making it, and is judged primarily on that live, out-of-sample record.

Six principles follow from that stance and recur throughout the design.

  1. Distributions, not points. A one-hour return is dominated by noise. A forecast of five quantiles and two direction probabilities carries its own uncertainty, which the policy can act on; a point forecast cannot.
  2. Online, with the clock as the only split. Every label is used for training only after the forecast it scores has been recorded. There is no separate validation period; the validation is the live record plus a held-out slice of the replay.
  3. One network for the whole cross-section. Names move together. A network that sees all of them at once can condition each forecast on the others and on the pooled market state.
  4. Defined risk only. No position the system takes, on paper or in principle, may carry unlimited loss.
  5. Everything observable. Every feed, feature, forecast, decision and training step is traced to a console the operator can read.
  6. One machine, nothing leaves it. Market data comes in; keys, positions and the narrative brief's language model stay local.
Feedspriority failoverper name Bars5 min, OHLCVreal trades Features25 per namez-scored FlintNetquantiles, up/downgate, attention128-bar window Calibrationband scale stemperature T Policy, then the bookgates, fractional Kellystock, puts, straddles Online learner2 steps per labelnight passes + guard weights Deep backfill, closedcandles 200 days back10-day chunk per 150 s Replay bufferkeyed by bar time5% held out, mmap label at t + 12, bps
Figure 1. The bar loop. Hot boxes run on the GPU. A forecast made at bar t becomes a training label at bar t + 12, so every window is scored before it is learned from.

2. Data

Bars and features

Sources live in a registry with a priority order, and each symbol is owned by the best source that is up and supports it; ownership fails over without operator action. Bars are built only from real trades. Quote and heartbeat ticks update the last price and the bid-ask spread but never form a candle, because a feed that echoes a stale quote while the market is closed would otherwise manufacture flat bars and the model would learn from fiction. Every modeled name needs a price for a bar row to emit, so a name without history is seeded flat at zero rather than allowed to stall the universe.

Each name contributes 25 features per bar, standardized by a running mean and variance kept per name: nine from the bars themselves (return, range, log volume, trade imbalance, spread, log trade count, close position within the bar, momentum, realized volatility), four from attention and tone (news sentiment and attention, retail-forum sentiment and attention), a fear-and-greed index, derivatives positioning (funding, open-interest change, long/short ratio), the net tilt of tracked fund books and an investor-council bias, two fundamental factors, market breadth and VIX, and at-the-money implied volatility and skew from option chains. The label for a window ending at bar t is the log return in basis points from that bar's close to the close at t + 12, capped at 2,000 bps, with a mask that removes any name that did not trade at both ends.

Table 0. The 25 features per name per bar and where each comes from. The first nine are computed from the bars; the rest are joined from the signal loops and held constant between their refreshes.
GroupFeatureDefinition
Barsretlog(close / previous close), bps
rangelog(high / low), bps
logvollog(1 + volume)
imbalance(buy volume − sell volume) / volume from tick classification
spreadbid-ask spread in bps from the last quote
logtradeslog(1 + trade count)
closepos(close − low) / (high − low)
momsum of the last 12 bar returns
rvolstandard deviation of the last 12 bar returns (close-to-close; see Section 7)
Attentionnews_sent, news_attntone from a local FinBERT (P(positive) − P(negative) per headline, the name's own headlines with the market tone as one headline of prior) and attention as the name's share of the pass's most-covered name, from company news and market desks
wsb_sent, wsb_attnthe same from retail-forum chatter
Positioningfngfear-and-greed index, market-wide
funding, oi_chg, longshortderivatives positioning where a venue reports it
guru_net, ethos_biasnet tilt of the tracked 13F books; the investor council's bias for the name
Fundamentalsf_value, f_qualityvalue and quality factor scores from company fundamentals
Marketbreadth, vixadvancers share on the radar; the VIX level
Optionsopt_iv, opt_skewat-the-money implied volatility about 30 days out, and the put-call skew, from the live chain

The replay buffer

Labeled windows are stored in a ring buffer of shape (windows × names × 128 bars × 25 features) in 32-bit floats. Three decisions about this buffer turned out to matter more than any modeling choice.

Keyed by bar time. Every restart re-walks the last five days of history to warm the feature state. Before windows were keyed, each restart appended another copy of those days: a 526-window buffer was found to hold only 180 distinct label vectors, so the most recent days were being trained on three times over. A window is now stored once per bar timestamp, and a re-walk adds nothing.

Budgeted in bytes. The buffer's capacity is a memory budget, 30 GB on the reference machine, from which the window count follows the universe: about 14,000 windows at 232 names and 7,000 at 400. Growing the universe shrinks the history held rather than the memory left for everything else.

Deep history paged in while the market is closed. A loop pages five-minute candles back in ten-day chunks, one every 150 seconds to respect the provider's rate limit, through an isolated bar and feature pipeline, until 200 days or a full buffer. Paging stops within a chunk of full rather than wrapping the ring onto its newest windows. Five percent of windows, every twentieth by timestamp, are held out of training permanently; the choice is derived from the timestamp so a reload keeps the same slice.

A window: names × 128 bars × 25 features, plus label and mask slot held out free keyed by bar timestamp: a re-walk of the same bars adds nothing A universe change, by column beforeABC afterBCD A dropped; B and C keep their rows and columns D masked until the re-page fills it Memory budget, not window count windows = budget / (names × 128 × 25 × 4 bytes) 30 GB ≈ 14,000 windows at 232 names, 7,000 at 400 paging stops near full; a smaller buffer on load keeps the newest every 20th bar is held out for good: the honest overfitting signal
Figure 2. The replay buffer and what a universe change does to it.

A self-updating universe

The modeled set is the configured core, every holding in the operator's accounts, and up to a fixed number of promoted names, under a cap of 400. Once a day after the close, candidates are drawn from three organic sources and scored on one 0–100 scale by rank within each source: the market radar's movers, weighted by dollar volume and filtered to liquid names above five dollars; tickers named in headlines, counted only when the symbol is known elsewhere; and the tracked funds' 13F positions by weight. The best rank across sources wins, with a small bonus for appearing in more than one.

Promotion fills the room left under the cap from the top of the ranking. Demotion is deliberately slow: a promoted name stays while it keeps ranking inside the top one-and-a-half times the slot count, and leaves only after 14 days without such a hit and at least five days of tenure. The asymmetry exists because a demotion forfeits that name's learned representation and its history in the replay, and a day's excitement is not worth that.

Candidates, scored 0–100radar movers × $ volumetickers named in headlines13F books of tracked funds Promotedranking fills the roomsince, last_hit = todayweakest leave first Retaineda hit: top 1.5 × countrefreshes last_hitrows and history kept Demoted14 days without a hitand 5 days of tenurerow and history lost top of the ranking keeps ranking goes quiet may re-enter fresh Fixed: the configured core and every holdingnever scored or demoted; a new holding enters within a minute room = min(count, cap − fixed); when it shrinks the weakest go first once a day after the close, once all three sources have reported
Figure 3. The universe's promotion and demotion rule. The asymmetry is deliberate: entry is cheap, exit forfeits learning.

A universe change used to reset the model, which is why no such rule could run automatically. Every parameter in the network is shared across names except a table of per-name embeddings, so a change now migrates: retained names' embedding rows and their optimizer moments are copied, new names receive fresh rows, retained names' replay columns are kept, new names' columns are masked, and the deep backfill re-pages history to fill those columns in. If the saved weights do not fit the current network at all, because the size benchmark landed on a different preset, the network starts fresh but the replay is still kept and migrated; the data outlives the architecture.

3. Model

FlintNet maps a tensor of (names × 128 bars × 25 features) to, per name, five return quantiles at the 10th, 25th, 50th, 75th and 90th percentiles, a probability that the return exceeds +10 bps and a probability that it falls below −10 bps, plus a regime mixture and an attention map that are exposed for inspection. Three architectural choices carry the design.

Dilated causal convolutions over time. Each name's window is read by a stack of one-dimensional convolutions with dilations 1, 2, 4, 8, 16, 32 and 64, which gives a receptive field over the whole window at a cost linear in its length, and causality by construction.

Cross-asset attention. After the temporal stack, a multi-head attention layer runs across names within the bar, so a forecast for one name can condition on the state of the others. A learned embedding per name lets the network tell the names apart; it is the only per-name parameter, which is what makes universe migration cheap.

A regime gate over expert heads. A pooled summary of the market state selects a soft mixture over twelve expert heads, so different regimes use different heads rather than one blended predictor. The gate's weights are published with each forecast. The quantile head is non-crossing by construction: it predicts the median and non-negative increments on either side.

InputN × 128 bars × 25 Dilated causal convsdilations 1 to 64per name, shared weightssees the whole window + name embeddingper-name parameter Cross-asset attentionacross the N namesmulti-headmap published Pooled statemean, spread Regime gatesoftmax, 12 experts 12 expert heads, mixed by the gatemedian + increments: q10 ≤ … ≤ q90logits for P(up), P(down) Losspinball + 2 × smoothed BCEmasked to names that traded gate weights and the attention map are shown on the dashboard
Figure 4. FlintNet. Width is the preset's d_model (512 or 768 on the reference machine); the temporal stack, the attention and the heads are shared across names.

Objective

The loss is the pinball loss over the five quantiles plus a smoothed binary cross-entropy on each direction head, averaged over the names the mask admits:

(1)L = (1/|M|) Σi∈M [ Στ ρτ(yi − qi,τ) + BCE(piup, 1[yi > θ]) + BCE(pidown, 1[yi < −θ]) ]

where ρτ(u) = u(τ − 1[u < 0]) is the quantile check function, θ = 10 bps is the direction deadband, and the direction targets are smoothed by 0.1 so the network cannot be certain. Gaussian noise with standard deviation 0.1 is added to the standardized inputs during training. On CUDA the forward pass runs under bfloat16 autocast with the loss in 32-bit; the training forward is compiled, which doubles throughput on the reference hardware.

Sizing to the machine

Rather than fix a network size, a benchmark at start climbs a ladder of presets and keeps the largest whose training step fits a time budget: the warmup ceiling of 300 seconds over 60 steps, so 5,000 ms per step. Memory is checked separately, before each rung is benchmarked and once more before the start: the rung's measured peak times 1.25, plus the replay budget, plus a 3 GB headroom must fit in what is free. The estimate for the next rung scales the last rung's measured peak by width, since a training peak grows about linearly with model width; a rung whose projected step time is far over budget, or whose need exceeds the machine next to the replay, ends the ladder honestly, while a transient squeeze from another tenant refuses the start rather than caching a downgrade for good.

Table 1. The preset ladder as benchmarked at 232 and 362 names on the reference machine (NVIDIA GB10, 121 GB unified memory), eager mode, batch 16. Step time grows with the universe, so the chosen rung falls from 5XL to 4XL as the universe grows past about 300 names.
PresetWidthParamsms/step, 232 namesPeak GBms/step, 362 namesPeak GB
XXL2564.4M96014.41,50722.4
XXXL38411.1M1,80025.42,80139.5
4XL51219.6M2,40233.93,73752.8
5XL76846.3M3,99251.0not benchmarked: about 110 GB next to the replay

4. Learning regimes

Four regimes share one network and one replay, and are never allowed to contend for the GPU at once.

A trading week, Eastern time warmup open, 09:30 to 16:00 closed: overnight weekend warmup load, migrate if needed, backfill five days, 60 steps a restart resets the live trust count open forecast every bar; two steps per resolved label trusted after 48 live labels, about 14:30 on a fresh start closed backfill chunks, 30 steps each; idle passes until the held-out loss stalls; the universe refresh after the close weekend the walk-forward backtest owns the GPU: live training waits, paging pauses, the net is sized to free memory always calibration and per-name records update one horizon late
Figure 5. When each learning regime runs. Nothing heavy ever shares the GPU with anything else heavy. The timer starts the engine 45 minutes before each weekday open and never restarts a running one.

5. Calibration

Two online corrections sit between the network and the policy, and both are updated per bar rather than per name, because a bar's several hundred labels are one observation of calibration, not several hundred: updating per name compounds into one enormous step and, with the horizon's delay, drives the correction into oscillation between its clips.

The first is an adaptive conformal scale s on the predicted band widths. The quantiles are widened or narrowed about the median, q̃τ = q0.5 + s(qτ − q0.5), and after each resolved bar the scale moves by

(2)s ← clip( s · exp( γ · ((1 − c) − 0.2) ), 0.5, 25 )

where c is the fraction of that bar's outcomes inside the 10–90 band and γ = 0.05, so the scale settles where one outcome in five falls outside. It starts at 2, wide, and earns its way down. The second is a temperature T on the direction logit, p̃ = σ(T · logit(p)), moved by one gradient step of the calibration cross-entropy per bar with rate 0.01, started sceptical at 0.3. Both corrections, and every per-name record the policy consults, are applied one horizon late in the backtest exactly as they are live.

Raw forecast at tquantiles, up/down Calibrated forecastq̃τ = q50 + s·(qτ − q50)p̃ = σ(T · logit p)to the policy Outcome at t + 12y per name, masked One update per bars ← s · exp(γ((1 − c) − 0.2))T ← T − η · mean(grad), eq. 2hit, coverage, pinball recorded updated s and T the next forecast uses the updated s and T: one horizon of lag
Figure 6. The calibration loop. A bar's several hundred labels count as one observation, which keeps the scale from oscillating.

6. Decision policy and risk

The policy is a composition of gates on the calibrated distribution followed by a sizing rule. It is deliberately not a learned layer: every gate is a live control, auditable, and testable in isolation, and the backtest can replay the identical gates over history. The score is a Sharpe-like ratio of the median to the calibrated interquartile range, z = q̃0.5 / (q̃0.75 − q̃0.25). A name may take a side only if all of the following hold: |z| exceeds a threshold; the tempered direction probability sits far enough from one half and agrees with the sign of z; the expected move exceeds the larger of a flat cost and the live bid-ask spread, and the spread is below a limit; the price is above a floor; the name's own record of resolved calls meets a minimum count and hit rate; and the same side has persisted for two consecutive bars. The book as a whole trades only once the model is trusted, after 48 live labels, and only while its hit-rate average beats a coin flip.

Size is fractional Kelly on the score, hard-capped, then scaled by band coverage relative to the 80% target so an overconfident model trades small:

(3)w = wmax · min(1, k · |z|) · max(0.25, min(1, c̄ / 0.8))

with k = 0.15. A bearish view is a long put and a volatility view, when both direction probabilities are high and the band is wide, is a long straddle, with premium per position capped at 2% of equity; a bullish view is stock. The paper book runs 100,000 dollars with no leverage, a 15% cap per name, fills marked at the spread, a minimum hold of six bars unless the signal reverses, and flattens at the close. With extended hours enabled, stock may trade in the 4:00–9:30 and 16:00–20:00 sessions; options never do. These gates were added after one evening in which churn and a one-dollar stock cost the book seven percent, and the policy's main job since has been to hold when it has no edge.

7. Pricing models: what Flint uses and what it should borrow

Flint is an empirical forecaster, and it was a fair question whether it uses any of the standard pricing models at all. The audit for this version is in Table 3. The short answer is that it uses one, Black-Scholes, in three places, and that the rest of the canon has something to offer, mostly as features and as corrections to the policy rather than as replacements for the network.

What Flint uses

Black-Scholes-Merton marks every option position in the paper book between quotes, prices the contract selected from the live chain (the at-the-money put about 30 days out for a bearish view; the at-the-money put and call pair for a straddle), and supplies two features from the chain, the at-the-money implied volatility and the put-call skew, which are the market's own forecast of the next month's distribution. The book sizes a put to the delta-equivalent short exposure the policy asked for. Fractional Kelly sizes positions from the score. Factor scores for value and quality, and the fear-and-greed and breadth measures, are joined as features. That is the complete list; realized volatility in the feature set is the plain standard deviation of the last twelve close-to-close bar returns, and the cost of a trade is the quoted spread.

What it should borrow

Table 3. The standard models, their status in Flint, and the insight each offers. "Features" means a change to the input set, which resets the network once; "policy" means a change to the gates or the book only.
ModelStatusWhat it offers FlintWhere
Black-Scholes-MertonUsedBeyond marking: the implied volatility scaled to the horizon, σ1h = IV·√(1 / (252 × 6.5)), is a market prior for the band width per name and a direct check on the conformal scale. A name whose calibrated band is far narrower or wider than the implied one is miscalibrated in a way the network cannot see.features
Breeden-LitzenbergerNot usedThe second derivative of the call price in strike is the risk-neutral density. Even without the full surface, the 25-delta skew and the wing prices give the market's asymmetry and tail weight for the quantile head to condition on.features
Implied term structureNot usedA front-expiry implied volatility far above the next expiry's marks a scheduled event, usually earnings. That is a veto gate: do not hold stock or sell a straddle through it; buy a straddle into it only if the model's move beats the implied one.features, policy
Variance risk premiumNot usedIV² − RV² over the last month is a regime variable that predicts index returns at horizons of weeks and conditions the gate; it is two numbers the system already has.features
OHLC volatility estimators (Parkinson, Garman-Klass, Rogers-Satchell, Yang-Zhang)Not usedRange-based estimators are five to eight times as efficient as close-to-close for the same bars, and Yang-Zhang handles gaps. The current rvol and range can be replaced in place without changing the feature count.features (in place)
HAR-RV, GARCH(1,1)Not usedA heterogeneous-autoregressive forecast of realized variance from daily, weekly and monthly components gives the horizon a volatility forecast that the conformal scale only approximates after the fact.features
CAPM betaUsed for the targetThe hedge ratio for the market-relative target in Section 11 (a ridge loading per modeled index since 0.5.14), and a rolling beta per name as a feature, so the network sees how much of a name's move is the index.target, features
Fama-French factorsPartlyValue and quality are in; the momentum (12-1) and short-term reversal (1-month) factors are cheap daily features with the best-documented premia.features
Greeks on the book (theta, gamma)In the policy since 0.5.14An at-the-money option about 30 days out decays by roughly premium / (2 × 30 × 6.5) per trading hour, a quarter of a percent of premium, which belongs in the cost edge for puts and straddles. A straddle's hourly profit is about ½Γ(ΔS)² − θΔt, so the entry rule should be "the model's expected |move| exceeds the implied one, σimpl·S·√Δt", in place of the present band-width threshold.policy
Square-root impact (Almgren-Chriss, Gatheral)In the policy since 0.5.14Cost = spread / 2 + η·σ·√(Q / ADV). At 100,000 dollars on large caps impact is negligible; for promoted small caps it is the whole edge. The cost gate should include it.policy
Amihud illiquidity, Kyle's lambdaThe gate since 0.5.14; not a feature|return| per dollar of volume is a liquidity gate that works for every source, including those with no quotes, where the spread gate is blind.policy, features
Roll spread estimatorNot usedThe effective spread from the negative autocovariance of returns, for names whose source supplies no quotes.policy
Order-flow imbalance (Cont, Kukanov, Stoikov)10PartlyThe present imbalance feature classifies trades; the quote-size imbalance at the best bid and ask is the stronger, nearly linear predictor of the next move, available where a feed reports sizes.features
Index futures basis, VIX term structureNot usedThe overnight futures move is the pre-open market factor; the futures basis against the ETF and the VX1−VX2 slope are regime variables. Needs a futures quote source.features
Ornstein-Uhlenbeck, pairsNot usedCross-sectional residuals against the index mean-revert; a fitted half-life per name is a feature and the natural companion of the market-relative target.features
Merton jump-diffusion, HestonNot used, by choiceTheir lesson, fat tails and clustered volatility, is what the quantile head and the conformal scale learn empirically; the implied surface already encodes the market's version. A parametric model here would constrain what the data should decide.none
Avellaneda-StoikovNot applicableInventory-based market making; Flint takes liquidity, it does not provide it.none

The authors' ranking by expected value per unit of effort: first the implied-volatility horizon band, the variance risk premium and the term-structure event flag, three features that reuse data the system already fetches; second the range-based volatility estimators and a HAR forecast, which upgrade existing features in place; third theta and the gamma breakeven in the options policy and the square-root impact and Amihud gates in the cost edge, which change nothing in the network; fourth beta and the market-relative target; fifth a futures feed.

Using them together

Nothing prevents all of these from entering at once; each is a few numbers per name, and the network's job is precisely to weigh many weak inputs. Two cautions govern the batch. The inputs should be distinct: four range-based volatility estimators say nearly the same thing, so one of them (Yang-Zhang) plus the HAR components earns its place while the rest add noise and cost samples, which matter with an online learner that sees thousands of windows rather than millions. And any change to the feature set resets the network once, so the batch should land as one change, on a weekend, with the backtest run before and after it. The batch the authors would ship is in Table 4, with one addition that costs no feature at all: macro ETFs as members of the universe. Because the network attends across names, treasuries, credit, the dollar, gold, oil and a volatility product become market factors the moment they are modeled, with no feature engineering.

Table 4. A feature batch from the pricing canon, computable from data the system already fetches. Thirteen new features plus two replaced in place; one model reset.
FamilyFeaturesSource
Option-impliedσ1h from at-the-money IV; term slope (front minus next expiry IV, the event flag); butterfly convexity; variance risk premium IV² − RV²; dealer gamma exposure, Σ open interest × gamma by strike, signed by side, which separates pinning days from trending daysthe live chain, already fetched for IV and skew
Realized volatilityYang-Zhang in place of rvol; Garman-Klass in place of range; HAR components: realized variance over the last day, week and monththe bars
Factorsrolling 60-day beta to the index; momentum 12-1; one-month reversalthe bars and the index ETF
MicrostructureAmihud illiquidity over 20 days; quote-size order-flow imbalance where the feed reports sizes; Roll's effective spread as the fallback where it does notquotes and bars
Mechanicalan ex-dividend flag for the bar after the ex-datethe fundamentals feed
Macro, as namesTLT, HYG, LQD, UUP, GLD, USO, VIXY added to the core universeno feature change; the attention layer does the rest

Other forecasters as opinions

A pricing model gives the network numbers about the world. A second class of inputs gives it other forecasters' answers to the same question, computed causally per bar and offered as features, so the network can weigh and correct them rather than replace them. Each answers "what happens over the next hour" from a different theory: a GARCH(1,1) and a HAR forecast of how large the move should be; the option-implied distribution for the horizon, the market's own quantiles and the most important second opinion available; an Ornstein-Uhlenbeck fit on the residual against the index, giving a z-score and a half-life for whether the name is stretched and how fast it reverts; a Kalman trend filter's local drift and uncertainty; and a ridge regression and a gradient-boosted tree trained online on the same features and labels, whose disagreement with the network is itself a feature. The investor council already supplies this kind of opinion for positioning. The combiner above all of them is the ensemble of Section 11 with a meta-labeling layer deciding whether to act, which makes "more opinions" a measurable claim: the backtest scores each opinion alone and the blend. Two rules keep the layer honest: no opinion may see a bar after the one it forecasts, and opinions must disagree to be worth their sample cost.

8. Evaluation

Live scoring

Every forecast is scored when its bar resolves: pinball loss on the raw quantiles, a hit if the median's sign matched the realized sign, coverage of the calibrated band, and paper profit if a position was taken. Three loss figures are shown side by side and mean different things: train pinball is in-sample over the replay; held-out pinball is the 5% slice the trainer never samples; live pinball is out of sample on forecasts made before their labels existed and is the only one that speaks to edge.

Walk-forward backtest

Because live labels accrue only while the market is open, a live record takes weeks. The backtest compresses that into a weekend by replaying the history in the replay with a fresh network of the live preset, day by day in time order. On day D it forecasts every window of the day before it has trained on any label from the day; the live policy's gates and sizing decide; the realized hour-ahead returns score the calls; and only then does the network train on day D, with the online budget plus the idle trainer's passes over everything seen so far. The first ten days are training only; signals wait for the trust rule; calibration and per-name records lag by one horizon. The procedure is therefore out of sample by construction and reproduces the live regimen including its adaptation.

For each trading day D, after the warm-up days 1. Forecast day Dno label of D seenfresh net, live preset 2. Decidelive gates and sizingrecords lag 12 bars 3. Scorepinball, hit, coveragebook return, gross ≤ 1 4. Train on Donline + night passesover all days up to D before training on D then, and only then next day the time budget shrinks step 4 so the run ends within the weekend day return = mean over bars of (Σ w·y − c·Σ|w|) the curve is its running sum against an always-long equal-weight book
Figure 7. The walk-forward procedure, reported per day and per name. Nothing in step 1 can see anything from step 4 of the same day. Live training waits for the run; the net is sized to free memory and the run stops itself under the memory floor.

Book accounting is honest about capital. Per bar the signals form a one-hour portfolio with gross weight capped at one; the bar's return is Σi wiyi − c Σi |wi|; a day's return is the mean over its bars, because capital can fund only one hour at a time; the curve is the running sum. The report gives per-day pinball, hit rate, coverage, signals, gross and net return, an equity curve against an always-long equal-weight book, and the best and worst names. A time budget shrinks the per-day training uniformly so a run always finishes within the weekend, and the report records what it used.

Results. As of this version no complete run exists; two attempts on the reference machine were interrupted by a universe rebuild and by the memory failure described in Section 9, and the first full run is scheduled for the weekend of October 11, 2026. The authors are not willing to report numbers from partial runs. The paper book's live record to date has been negative, which is consistent with the literature's finding that documented intraday effects are small, and is the motivation for the work in Sections 7 and 11.

8. System design

Flint is one asyncio process: a bar clock, the forecasting path, the four training regimes, and periodic loops for market status, signals, the radar, news, checkpoints, the brief, the portfolio, deep backfill, idle training, the universe refresh and the backtest. A FastAPI server exposes the state as JSON and streams it over one websocket to a dashboard of plain HTML and JavaScript with no build step. The checkpoint is written atomically every five minutes; the replay, tens of gigabytes when full, is written uncompressed as one memory-mapped array per field at most every 30 minutes and at shutdown, swapped into place whole so a crash leaves the previous directory intact.

The reference machine's unified memory is shared by the model, the replay and everything else, and its GPU allocations are invisible to process accounting and outside any control-group limit. Several invariants exist because this was learned the hard way.

121 GB at 362 names and 4XL during a weekend backtest replay 30 GBmmap on disk backtest net 53 GB at peaksized to free memorythe live net does not train weights 2, system 7 floor 10 headroom 19 The same 121 GB on October 10, before these rules, with no backtest running: replay 12 90 GB of allocator cache outside any processone working set each for the training step,the validation pass and per-bar prediction 19 free Earlier that night two such peaks ran at once and the box rebooted
Figure 8. The memory map that the rules in this section enforce, and the two states that preceded them.

9. Limitations and threats to validity

10. Future work

Ranked by expected effect per unit of effort, each item is a hypothesis for the walk-forward backtest.

  1. A market-relative target (shipped in 0.5.14 as a switch, judged by the next backtest). Over one hour the index dominates every name's move and nothing in a five-minute history predicts the index. Nonlinear models beat linear ones mainly when the target is the market-relative abnormal return4. The label becomes the residual of a ridge regression of the name's return on the modeled index names' returns over the same hour (SPY, QQQ, IWM, SMH and XLF, so a semiconductor name is measured against SMH and a small cap against IWM, not everything against SPY), loadings refit once a session and shrunk toward the broad market; the replay keeps absolute labels, so the target is a run-time switch and one replay serves both. The book holds, per index, the position that cancels the basket's loading: an index put under a long basket, index stock over a bearish one. The weekend backtest runs both targets back to back on the same history and reports them side by side.
  2. The pricing-model pass of Section 7, in its stated order: implied-volatility horizon band, variance risk premium and the term-structure event flag; range-based volatility and a HAR forecast; theta, the gamma breakeven, square-root impact and the Amihud gate in the policy.
  3. Calendar and daily-lag features. A name's return in a given half-hour continues at the same half-hour on following days for about forty days, strongest at the open and close1; the first half-hour's index return predicts the last half-hour's2; overnight and intraday components carry opposite, persistent signals3. None are features yet.
  4. An ensemble of mid-size networks. Tens of millions of parameters on some thousands of windows is heavily over-parameterized, as the overfit guard's early stops confirm. Two 4XL networks fit the machine, and aggregating ensemble members' distributions improves accuracy and calibration8.
  5. Meta-labeling. A second, small classifier trained on whether acting on each primary forecast paid after costs, sizing by its probability5,6: the evaluable form of a decision layer, trainable online from the outcome records Flint already keeps.
  6. Calibration to decision loss net of costs7, sizing by the probability that the move clears the spread from the full quantile set rather than the median alone; and exits with a profit target, a stop, a time limit and a daily loss limit.
  7. A policy sweep over stored forecasts (shipped in 0.5.14). Every forecast a backtest makes is stored, and a vectorized replay of the policy over them reproduces the run's net to the basis point in under a second, so the gates can be swept: a coordinate descent on the first half of the scored days, confirmed on the second, with the result proposed on the dashboard rather than applied.
  8. SEC filings beyond 13F. Form 4 insider transactions, 13D and 13G stakes and 8-K item codes as direct features; free-text filings summarized by a local 7–8B language model into a fixed schema of direction, materiality, horizon and topic flags, one cached call per filing; and notable filings as a fourth candidate source for the universe.

References

  1. S. L. Heston, R. A. Korajczyk and R. Sadka, "Intraday Patterns in the Cross-section of Stock Returns," Journal of Finance 65(4), 2010. PDF
  2. L. Gao, Y. Han, S. Z. Li and G. Zhou, "Market Intraday Momentum," Journal of Financial Economics, 2018. Summary
  3. D. Lou, C. Polk and S. Skouras, "A Tug of War: Overnight versus Intraday Expected Returns," Journal of Financial Economics, 2019. PDF
  4. "Design Choices in Machine Learning for Stock Returns," review of target and model choices. Article
  5. M. López de Prado, Advances in Financial Machine Learning, Wiley, 2018: triple-barrier labeling and meta-labeling. Overview
  6. Hudson & Thames, "Does Meta-Labeling Add to Signal Efficacy?" Article
  7. "Utility-Weighted Forecasting and Calibration for Risk-Adjusted Decisions under Trading Frictions," arXiv:2601.07852, 2026. Abstract
  8. Aggregation of deep-ensemble predictive distributions, Machine Learning, 2025. Record
  9. Transformer-based deep learning for stock return forecasting, with transaction costs. PDF
  10. R. Cont, A. Kukanov and S. Stoikov, "The Price Impact of Order Book Events," Journal of Financial Econometrics 12(1), 2014. arXiv

Appendix: principal settings

Table 2. Defaults on the reference machine. Items marked live can be changed from the dashboard without a restart.
SettingDefaultRole
bar_seconds / horizon / window300 / 12 / 128five-minute bars, one hour ahead, 128 bars of context
quantiles / direction_threshold_bps10, 25, 50, 75, 90 / 10the distributional head and its direction deadband
lr / weight_decay / dropout / input_noise5e-4 / 1e-2 / 0.15 / 0.1optimizer and regularization
batch_size / recent_frac / recent_n16 / 0.3 / 256recency-biased sampling
steps_per_label / warmup_steps / max_warmup_seconds2 / 60 / 300online budget; the size ladder's time budget
idle_train_epochs / overfit_patience (live)2 / 3closed-session passes and the held-out stop
deep_backfill_days (live) / replay_gb / replay_save_minutes200 / 30 / 30history paged, buffer budget, replay write cadence
max_universe / interesting_count (live) / interesting_grace_days / interesting_min_tenure_days400 / 150 / 14 / 5the universe and its promotion and demotion rule
min_labels / min_hit_rate (live) / skill_min_n (live)48 / 0.5 / 8trust, and the per-name record
score_threshold / prob_margin / move_floor_bps (live)0.35 / 0.06 / 8the entry gates
max_spread_bps / min_price / confirm_bars / stock_min_hold_bars (live)25 / 5 / 2 / 6liquidity and persistence gates
kelly_fraction / max_size / option_max_frac (live)0.15 / 1.0 / 0.02sizing
band_gamma / temper_lr0.05 / 0.01the two calibration rates
backtest_hours / backtest_night_epochs / backtest_cost_bps / backtest_warm_days (live)30 / 2 / 5 / 10the walk-forward run
memory_floor_gb (live)10the floor every GPU tenant keeps free