Introducing Forezai · TradingAgents — a committee of LLMs decides paper-trades

📊 Full opportunity report: Introducing Forezai · TradingAgents — a committee of LLMs decides paper-trades on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Forezai · TradingAgents has introduced an autonomous system where a committee of specialized LLMs executes paper trades based on structured debates and decision-making protocols. This development aims to explore AI’s potential in trading without risking real money.

Forezai · TradingAgents has unveiled a new system that allows a committee of large language models (LLMs) to autonomously generate and execute paper trades based on structured, multi-role debates and decision-making processes. This development marks a significant step in AI-driven research for trading strategies, focusing on decision transparency and multi-agent reasoning rather than prediction accuracy.

The system is a fork of an existing open-source multi-agent framework called TradingAgents, originally designed to test whether LLMs, when assigned specialized roles, can produce trading decisions comparable to or better than random chance. The new Forezai version adds operational features: an automated scheduler, paper-trading interfaces with multiple broker modes, and a web dashboard for monitoring. It does not trade with real money by default, emphasizing research and simulation.

Specifically, the framework involves multiple analyst roles—covering market structure, news, fundamentals, and social sentiment—that generate independent reports. These reports are debated by bull and bear agents, with a research manager synthesizing them into a cohesive view. A risk team evaluates upside and downside, culminating in a final buy, hold, or sell decision by a trader agent, which is then aggregated by a portfolio manager agent into a detailed rating and target price. The entire process is designed to articulate reasoning explicitly, avoiding reliance on raw data recall.

Forezai’s operational layer introduces a scheduler that runs daily, mapping the decision outputs into paper orders with filtering and risk controls, including position management and exit rules. It supports multiple modes: local Python, Alpaca paper trading, and a shadow mode for comparison. A web dashboard provides real-time analytics, including equity curves, drawdowns, and performance metrics. The system runs locally, with no data sent to external servers, maintaining a focus on controlled research environments.

Introducing Forezai · TradingAgents — Thorsten Meyer AI
AGENTS
● ANNOUNCEMENT / MAY 2026
THORSTEN MEYER AI · FOREZAI · § 03
FOREZAI · 03
TRADINGAGENTS · LAUNCH
Research Series · Companion to Polybot Week 1-2 · 2026-05-17

Introducing Forezai · TradingAgents.
A committee of LLMs
decides paper-trades.

After two weeks of finding out most parametric strategies don’t work, the obvious next research question: can multi-agent LLM judgment do any better?
A fork of the open-source TradingAgents framework (TauricResearch): thirteen LLM agents in four stages — four parallel analysts · a bull-bear debate with research-manager arbitration · a three-voice risk team · a two-layer trader + portfolio-manager decision. The fork keeps the agent graph intact and adds the operational layer the upstream doesn’t ship: an autonomous loop · a multi-broker abstraction · a local web dashboard · Codex OAuth · MCP plug-ins · 520+ unit tests. The question is narrower than “do LLMs predict the market” — that prior is “no, with high confidence.” The narrower question is: when LLMs are structured into specialised adversarial roles, does the committee produce decisions at least no worse than a coin flip after fees? Honest priors before running: it might fail too. If it appears to work, the most likely explanation is variance.
This is not financial advice. Nothing in this announcement should be used to inform real trading decisions. The software described trades simulated money by default. If you reconfigure it to trade real money, you should expect to lose that money — regardless of how clever any individual agent’s reasoning looks. Algorithmic trading is zero-sum after fees and structurally hostile to part-time retail strategies.
13 agents
Specialised roles in four stages
Analysts · Debate · Risk · Decision
78% / -33%
Polybot prior: fleet win rate
combined with -33% bankroll
520+
Passing unit tests across engine,
services, HTTP routes (starting baseline)
€0 floor
LLM cost on Codex OAuth
(falls back to public API per token)
FOREZAI / TRADINGAGENTS· APACHE 2.0 FORK· UPSTREAM TAURIC RESEARCH· LANGGRAPH· 13 AGENTS / 4 STAGES· 4 PARALLEL ANALYSTS· BULL-BEAR DEBATE· 3-VOICE RISK TEAM· TRADER + PORTFOLIO MANAGER· 5-TIER FINAL RATING· ALPACA PAPER + LOCAL + SHADOW· LIVE ENDPOINTS HARD-REFUSED· FASTAPI + REACT VIA CDN· CODEX OAUTH· MCP PLUG-IN REGISTRY· 520+ UNIT TESTS· POLYBOT WEEK 1: 21 EXPERIMENTS· WEEK 2: -33% BANKROLL· 78% FLEET WIN RATE· HONEST RESEARCH, NOT EDGE· FOREZAI / TRADINGAGENTS· APACHE 2.0 FORK· UPSTREAM TAURIC RESEARCH· LANGGRAPH· 13 AGENTS / 4 STAGES· 4 PARALLEL ANALYSTS· BULL-BEAR DEBATE· 3-VOICE RISK TEAM· TRADER + PORTFOLIO MANAGER· 5-TIER FINAL RATING· ALPACA PAPER + LOCAL + SHADOW· LIVE ENDPOINTS HARD-REFUSED· FASTAPI + REACT VIA CDN· CODEX OAUTH· MCP PLUG-IN REGISTRY· 520+ UNIT TESTS· POLYBOT WEEK 1: 21 EXPERIMENTS· WEEK 2: -33% BANKROLL· 78% FLEET WIN RATE· HONEST RESEARCH, NOT EDGE·
FIG. 01 — THE 13-AGENT COMMITTEE
Thirteen specialised roles · four stages · biases made to argue in public
The architecture forces the system to articulate its reasoning rather than relying on what a single context window happens to recall
Stage 1 · Four analysts in parallel4 agents
Market
Structure, ranges, regime indicators
News + Insider
News flow, filings, insider activity
Fundamentals
Balance sheet, earnings, ratios
Social Sentiment
Social-media tone, retail signal
Stage 2 · Bull-bear debate + research-manager arbitration3 agents
Bull researcher
Argues upside thesis from analyst reports
Bear researcher
Argues downside thesis from same reports
Research manager
Arbitrates · writes single synthesis
Stage 3 · Three-voice risk team3 agents
Aggressive
Looks for upside · accepts variance
Conservative
Looks for downside · protects capital
Neutral
Balances · forces downside articulation
Stage 4 · Two-layer decision2 agents
Trader
Three-tier proposal · buy / hold / sell
Portfolio manager
Five-tier rating + price target + horizon · sees arguments only, never raw data
The portfolio manager only sees the arguments, never the raw data — which forces the committee to make its reasoning explicit rather than relying on a single context window’s recall. The upstream framework ships the agent graph; it does not ship the operational machinery to run that graph on autopilot, observe its results honestly, store them for later inspection, or prevent the operator from accidentally trading real money. That gap is what the Forezai fork fills.
FIG. 02 — THE POLYBOT PRIOR · WHY THIS IS A DIFFERENT BET
Two weeks of paper-trading prediction markets · the trap underneath the headline numbers
25 experiments · 78% fleet-wide win rate · -33% bankroll · most parametric strategies are structurally negative-expectation when measured honestly
The flattering number
78%
Fleet-wide win rate · week 2
“You can win four out of five trades and still go broke, because the one loss is bigger than the four wins put together.” Win rate without P&L context is a mechanical illusion.
The honest number
−33%
Fleet bankroll · week 2 close
The strongest possible demonstration of the trap. A parametric trading strategy that looks compelling in a backtest will almost always fail to survive a fresh sample. Most “edges” are mechanical artefacts.
Week 1: 21 parallel strategy experiments · early winners mostly mechanical illusions · exactly one strategy (a fair-value taker on BTC) showed the mathematical signature of real edge over a few hundred settled trades. Week 2: same fair-value strategy with more data collapsed. A separate mid-week hypothesis (market-making) also failed cleanly. Fleet ended week 2 at roughly negative thirty-three percent of bankroll. The honest research finding wasn’t on the winning side — it was on the losing side. Adding more parameters to Polybot wouldn’t change that. TradingAgents is asking a separable question.
FIG. 03 — WHAT THE FORK ADDS · THE OPERATIONAL LAYER
Six layers the upstream framework doesn’t ship
Same agent graph, intact. The fork makes it a research instrument rather than a tech demo.
01 · Loop
An autonomous loop
Scheduler · watchlist · auto-trader maps ratings to paper orders · allow-list filtering · per-ticker cooldowns · sector caps · cash checks · position manager evaluates open positions every 60s for TP / SL / max-hold. Append-only audit logs.
02 · Brokers
Multi-broker abstraction
Three modes: local Python broker (yfinance fills, JSON-persisted) · Alpaca paper-trading adapter · “shadow” mode running both in parallel with divergence view. Real Alpaca live endpoints are hard-refused at multiple layers.
03 · Dashboard
A local web dashboard
FastAPI backend · React via CDN, no Node toolchain · SVG equity curve · rolling-peak drawdown · win-rate by rating / ticker / model · exit-reason breakdown · LLM cost vs realised P&L joined by run ID. Runs locally; nothing sent to a cloud service.
04 · Codex
Codex OAuth
Runs the engine on a ChatGPT Pro subscription via the Codex backend. LLM cost floor effectively zero if you already have ChatGPT Pro. Token stored encrypted locally. Falls back to the regular OpenAI API if you’d rather pay per token.
05 · Alerts
Multi-channel alerts
Slack · Discord · SMTP email · configurable filter on rating events and order fills · append-only history kept locally. Webhook URLs masked in API responses so a screenshot can’t accidentally leak credentials.
06 · MCP
MCP plug-ins
Registry for adding Anthropic Model Context Protocol servers (Kensho · Aiera · FactSet · Morningstar · LSEG) as analyst tools. Plug-ins advertise category (fundamentals · news · market data · social) · probe endpoint tests credentials.
Honest-by-design touches: every generated report prepends “Research, not advice” and appends a footer with version, commit, provider, models used, run ID, and cost. Closed trades carry the same metadata. 520+ passing unit tests across engine, services, and HTTP routes. The intent: when the system loses money, the journal makes it impossible to pretend it didn’t.
FIG. 04 — HONEST PRIORS · BEFORE RUNNING THIS IN ANGER
Three priors stated before the data starts arriving
The bias of the project: when the data says no, the dashboard says no, the article says no
1
It might fail too. LLMs are not oracles, and a sophisticated framework around language-model outputs does not change the underlying error rate of the model. Sample is still everything. The framework’s outputs are subject to the same statistical noise as any prediction system over small samples.
Highest likelihood
2
If it appears to work, the most likely explanation is variance. The same trap that caught the first article’s candidate edge applies here. A high win rate over fifty trades means much less than it looks. Without out-of-sample confirmation, a flattering early sample tells you almost nothing about whether the system has real edge.
Second-most likely
3
If it appears to work for the right reasons — empirical win rate matches stated confidence, and alpha-versus-benchmark persists across non-overlapping samples — that would be a meaningful research finding. Whether that happens, I don’t know. The point of putting it in the open is that the data will say.
Genuinely open
This is explicitly not a launch announcement for a product anyone should connect a real brokerage account to. The Alpaca live endpoints are hard-refused at multiple layers in the code, and the design choice is deliberate. The right next step is data, not deployment. The bias of the whole project is straightforward: when the data says no, the dashboard says no, the article says no, and no one tries to retroactively rescue the thesis. That’s the contribution.
FIG. 05 — WEEK THREE · WHAT THE METHODOLOGY WILL MEASURE
Four concrete measurements before publishing findings
The hope: write the week-three article from a position of “here’s what the data says”. The fear: another candidate falsified at higher sample. Both outcomes are publishable.
M1 · Sample discipline
Small watchlist for a few weeks before publishing
A handful of tickers across two or three sectors. Long enough to gather sample, narrow enough to keep attention on what’s actually happening per agent. Avoid the noise of a 65-ticker autonomous loop until the smaller version has been read carefully.
M2 · Calibration view
Stated confidence vs. realised win rate
When the system says “75% confident”, do the trades actually win 75% of the time? Same measurement applied to Polybot’s fair-value model. If the model is systematically over-confident, that bias dominates everything downstream.
M3 · Cost accounting
Cost per ticker · per rating · per profitable trade
With Codex OAuth the marginal LLM cost is effectively zero. With the public OpenAI API, each run is hundreds of agent turns. The honest question: does this scale economically if you ever did run it at real cost?
M4 · Non-overlapping windows
Alpha vs benchmark · out-of-sample
Not within-sample alpha — trivially inflatable. Hold out one period entirely, run the system on the next, then check whether the held-out result matches the in-sample stats. If they diverge sharply, the in-sample was curve-fit.
Open under Apache-2.0 with upstream cited from every relevant surface. Not open: the operator’s running results, the specific watchlist, the per-agent prompt customisations, the alert channels, the trade journals — kept local for the same reason Polybot’s per-experiment data is kept local. Publishing exact configurations encourages people to copy them with real money, which is the opposite of what an honest research project should do. Summary findings will be published. Recipes will not.
The bet is on a different mechanism, not a different parameter setting. The point is not to find a money-printing AI. The point is to put honest measurements of these systems into the public record — so the next person looking at the space starts a step further along than the last.
Thorsten Meyer AI · Introducing Forezai · TradingAgents · § 03

Potential Impact on AI-Driven Trading Research

This development is significant because it demonstrates a practical implementation of multi-agent LLM systems making autonomous trading decisions in a simulated environment. While not designed for real trading, it provides a testbed for exploring how structured reasoning, role specialization, and explicit articulation of decision rationale can improve AI decision-making. If successful, it could influence future research on AI explainability and collaborative reasoning in financial contexts.

Moreover, by emphasizing transparency and modularity, Forezai · TradingAgents offers a framework for researchers to experiment with multi-voice AI systems, potentially leading to more robust and interpretable AI trading agents. It also highlights the current limitations of LLMs in prediction and the importance of structured debate and reasoning, rather than raw forecasting, in trading applications.

Chemical Process Simulation and the Aspen HYSYS Software

Chemical Process Simulation and the Aspen HYSYS Software

Used Book in Good Condition

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Evolution of AI in Trading Research

Previous research with parametric trading strategies has shown that many seemingly promising rules fail to survive out-of-sample testing, often due to overfitting or mechanical artifacts. Recent efforts have shifted toward understanding whether less rule-bound AI systems, like multi-agent LLM frameworks, can produce more reliable decisions. The original TradingAgents project was designed to test whether specialized LLM roles, engaging in structured argumentation, could outperform random or naive strategies.

Forezai’s fork builds on this foundation, adding operational features necessary for systematic research. While earlier experiments focused on paper-trading against prediction markets, the new system emphasizes the decision process itself, with a focus on explainability, risk management, and automation. This aligns with broader trends in AI research that seek to combine multi-agent reasoning with practical testing environments.

“This system allows us to test whether a committee of specialized LLMs can produce decisions that are at least as reliable as random chance, with the added benefit of explicit reasoning.”

— Thorsten Meyer, lead researcher

The Intelligent AI Investor: A Beginner’s Guide to Using AI Tools for Informed Investment Decisions, Risk Management, and Wealth Building (Trading & Investing Series Book 7)

The Intelligent AI Investor: A Beginner’s Guide to Using AI Tools for Informed Investment Decisions, Risk Management, and Wealth Building (Trading & Investing Series Book 7)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations and Unanswered Questions About AI Trading Agents

It remains unclear how well the system’s decision-making would translate to live trading environments or whether the structured debate approach significantly improves performance over simpler models. The effectiveness of the multi-voice reasoning in reducing overfitting or bias has yet to be empirically validated in real market conditions. Additionally, the impact of operator intervention and the system’s robustness under volatile market scenarios are still being studied.

Advanced Stock Market Analysis Dashboard: Design and Implementation of a Real-Time Financial Analytics Platform

Advanced Stock Market Analysis Dashboard: Design and Implementation of a Real-Time Financial Analytics Platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Developing AI-Driven Trading Systems

Researchers plan to run extended experiments using Forezai · TradingAgents, analyzing the decision quality, consistency, and explainability of the AI committee. Future work may include integrating real-time market data, refining role definitions, and testing the system in live paper trading with more complex risk controls. The goal is to establish whether structured multi-voice AI can meaningfully contribute to more transparent and reliable trading decision frameworks.

Algorithmic Trading - Algorithmic Trading Strategies - Compendium: Volumes 1 to 20: Trading Systems Research and Development

Algorithmic Trading – Algorithmic Trading Strategies – Compendium: Volumes 1 to 20: Trading Systems Research and Development

Used Book in Good Condition

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Can Forezai · TradingAgents trade with real money?

Currently, the system is configured for paper trading only. While it supports modes that could be adapted for live trading, actual real-money trading requires deliberate operator override and additional risk controls.

How does the system ensure decision transparency?

The system explicitly articulates reasoning through multiple specialized agents debating and synthesizing their arguments, which are logged and accessible for review.

Is this system designed to outperform human traders?

No. The primary goal is research and understanding of multi-agent AI reasoning in trading contexts, not to replace or outperform human traders in live markets.

What are the main limitations of this approach?

Its effectiveness depends on the quality of the LLMs and the structure of the debate; it has not yet demonstrated consistent outperformance or robustness in volatile or real trading environments.

When will the system be available for broader research use?

It is currently active for research and testing; broader deployment or open access will depend on ongoing validation and development efforts.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
You May Also Like

Azerbaijan Surges In Global Coverage

Azerbaijan’s coverage in global media has surged, with GDELT recording 35 mentions in a recent time window, marking a notable increase.

What Spot and Futures Traders Watch Before Major Events

Knowledge of market indicators and news signals helps traders anticipate major event impacts, but understanding what they focus on can reveal even deeper insights.

World Economic Forum 2026: Crypto and CBDCs Take Spotlight in Davos

Discover how digital currencies and CBDCs are shaping Davos 2026, with key insights that could redefine your financial future—continue reading to learn more.

Why Coinbase Listings Still Move Markets

Finding out why Coinbase listings still move markets reveals the power behind their influence and what it means for investors.