The Finance Lab Introduces TFL Bloodhound, a Financial Reasoning Model Trained on Market Outcomes Instead of Human Preference
TFL_Bloodhound replaces the human evaluator in RLHF with realized market outcomes, and separates quantitative
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.

![]()
TFL_Bloodhound replaces the human evaluator in RLHF with realized market outcomes, and separates quantitative estimation from language reasoning in a compound architecture. In pilot deployment across energy futures; a public playground is open for testing.
TORONTO, ON / ACCESS Newswire / August 27, 2026 / The Finance Lab today introduced TFL Bloodhound Model 1, a financial reasoning system built on a training signal that is unusual for a language model: the reward comes from the market, not from a human rater.
The design starts from a constraint that most financial LLM work routes around rather than through. Markets are not documents. The information that matters is numerical, temporal, relational and non-stationary – a relationship that carried predictive power six months ago may carry none today, and one that looked like noise may become load-bearing under a new volatility regime. Asking a single language model to both estimate high-dimensional time series and reason about them forces it to do the part it is worst at in order to reach the part it is good at.
Bloodhound was built on the opposite assumption: financial AI should not ask one model to do everything.

Separating estimation from reasoning
Specialist TFL models handle the numerical structure of markets – market regime, trend and directional alignment, volatility state, momentum, cross-asset relationships, historical similarity and forward-return distributions – and emit structured evidence rather than predictions. Each returns a machine-generated representation, so the reasoning layer receives a regime classification rather than raw time-series values to interpret in prose.
The reasoning layer does the part that is actually linguistic: comparing competing signals, surfacing contradictions, identifying what is missing, and deciding which evidence matters under current conditions. The domain training, orchestration, memory integration and reinforcement learning around it are what make it Bloodhound.
When the reasoning layer hits uncertainty, it can call another specialist model, retrieve historical context, or recall an analogous situation from memory. The loop is iterative and evidence-seeking rather than single-pass.
Reinforcement Learning from Market Feedback
Conventional alignment uses RLHF – a human tells the model which answer is better. Bloodhound uses Reinforcement Learning from Market Feedback (RLMF), a training approach developed by The Finance Lab in which the reinforcement signal is the realized outcome.
Each expected market movement is emitted as a testable hypothesis, recorded alongside the evidence available, the information the model chose to request, the specialist models it consulted and the reasoning path it took. When the outcome becomes observable, the reward is computed across several dimensions rather than directional accuracy alone: probability calibration, magnitude error, distributional accuracy, risk-adjusted utility, maximum adverse movement, and the relevance of the evidence retrieved. The reasoning policy is trained with Group Relative Policy Optimization (GRPO) around evidence retrieval and hypothesis evaluation.
The last of those dimensions is the one that does the distinctive work. Credit is assigned at the level of evidence, not the response – which specialist outputs genuinely contributed to a calibrated expectation, and which were along for the ride.
This is a harder reward than RLHF, and worth being explicit about why. Market feedback is sparse, heavily noisy, non-stationary, and prone to rewarding a wrong process that happened to produce a right answer. Bloodhound addresses this through graded rather than binary outcome scoring, evidence-level rather than response-level credit assignment, and strict point-in-time construction of the evidence available at hypothesis time. The claim is not that markets are predictable. The claim is that which evidence mattered is answerable after the fact, and that this is a usable training signal.
MeMo: Memory-as-a-Model
Bloodhound’s second component is MeMo (Memory-as-a-Model), which stores prior situations as multidimensional analytical experiences: instrument, market environment, model evidence, hypothesis, expected movement and realized outcome.
Memory relevance is not permanent. MeMo pairs long-term retention with adaptive relevance, so that structurally similar situations across sectors and asset classes can be retrieved as probabilistic evidence while stale relationships decay out of influence – remembering enough of the market to determine what still matters now.
Scale
Bloodhound’s specialist layer was developed across more than 20,000 financial instruments, 100.8 million historical candles and roughly two billion feature-level observations spanning about two decades. The purpose of that breadth is cross-sectional: it lets the system reason beyond a single instrument’s own history and treat analogous situations elsewhere as evidence rather than as deterministic rules.
Foundation model
Bloodhound’s reasoning layer is a fine-tuned variant of Gemma 4 26B A4B, a sparse Mixture-of-Experts model with a 262,144-token context window, retrained for financial reasoning. The foundation model supplies language, long context and tool use; everything that makes the system financial – the specialist models, the orchestration, MeMo and RLMF – is TFL’s.
Parameter count is the wrong yardstick for the resulting system. A monolithic model has to spend capacity approximating numerical work it will never do precisely; Bloodhound’s reasoning layer spends none, because eight specialist quantitative models run in parallel and hand it finished evidence. The capability that matters is that of the composite – the reasoning model plus everything it can call and everything it can remember – and that is not bounded by what 26 billion parameters could compute alone.
Pilot deployment
Bloodhound Model 1 is in pilot deployment at an energy-focused hedge fund in Texas, applied to natural gas, crude oil and heating oil futures. The fund is not named at its request.
Energy was a deliberate first domain rather than a convenient one. The drivers that matter there – weather, storage, spreads, term structure, the relationship between the front and the curve – are precisely the kind that shift regime instead of holding constant, and the crack and spark relationships that anchor the complex can decouple for months at a time. That makes it a demanding test of the two things Bloodhound claims: that evidence relevance can be learned from outcomes, and that memory of a prior situation can be retrieved without assuming the prior situation’s logic still holds. Findings will be published as TFL releases evaluation results.
Availability
Bloodhound Model 1 can be tested now in a public playground at go.thefinancelab.ai, running on a single default instrument. The playground exposes the full reasoning trace – the evidence the model requests, the specialist models it consults, the hypotheses it forms and discards – rather than a summary verdict.
Full access, covering the complete instrument universe and the TFL Financial AI Harness that hosts the model, is available to customers. Within the harness, model-generated hypotheses are validated by portfolio and deterministic risk components before progressing to any downstream action. A technical brief covering the RLMF objective, reward construction and evaluation methodology is available on request at sales@thefinancelab.ai.
Architecture documentation is published at thefinancelab.ai/models.
“I traded options and energy futures for years, and the failure mode was never that a model was wrong on day one,” said Evandro Barros, CEO and Head of AI Research at The Finance Lab. “It was that a relationship you had relied on for two years quietly stopped working, and you found out from the P&L instead of from the model. RLMF is an attempt to move that feedback into the training loop, where it belongs.”
“Quantitative models should calculate, measure and represent the market. The reasoning model should connect evidence, question its own hypothesis and decide what matters now. Most of the difficulty in this system is not in either half – it is in making the second half accountable to what actually happened.”

About TFL Bloodhound
TFL Bloodhound is an AI-powered financial intelligence and reasoning system developed by The Finance Lab, combining proprietary financial models, adaptive reasoning, multidimensional memory and agentic workflows. Reinforcement Learning from Market Feedback (RLMF) and Memory-as-a-Model (MeMo) are proprietary approaches developed by The Finance Lab. Bloodhound outputs are probabilistic analytical information and should not be interpreted as certainty regarding future market movements or as financial advice.
About The Finance Lab The Finance Lab develops AI financial autopilots that combine autonomous agents and proprietary models, enabling users to train their own systems and build their own AI-powered investment firm.
Contact Information:
Evandro Barros
https://thefinancelab.ai · sales@thefinancelab.ai
SOURCE: The Finance Lab
View the original press release on ACCESS Newswire
Media gallery
