Three of the Top Ten Spots on the FutureX Leaderboard
Yun Zhong, reporting from Aofeisi
In July, DigClaw’s Rhizome v1 prediction framework secured #1, #3, and #7 on the FutureX evaluation platform.
The three spots were achieved using three different foundation models, including Kimi-K3 and DeepSeek-V4-Pro. The same framework placed all three models in the Top 7, making DigClaw the only participant to accomplish this.
The evaluation consisted of 59 real-world event prediction questions spanning politics, economics, and technology, with no possibility of training-data leakage.
These results support a hypothesis that DigClaw is currently testing:
Predictive capability can be developed outside the foundation model.
As foundation models continue to improve, systems can benefit from their growing capabilities. Meanwhile, prediction trajectories, settlement feedback, calibration experience, and continuously evolving workflows can continue to accumulate within the system itself.
△DigClaw 2026 Jul. Week 4 performance on the FutureX leaderboard
This was not an accidental competition result. It was DigClaw’s first external validation of its answer to the question: “How should prediction actually be done?”
Prediction Is AI’s Most Underrated Capability
As prediction tasks become increasingly standardized and industrialized, AI is also expected to handle complex prediction problems at scale.
But this presents a fundamental problem: large language models are not naturally good at prediction.
LLMs learn correlation rather than causation. They extract patterns from past corpora, but “learning from the past” and “predicting the future” are fundamentally different things.
This creates three fatal problems:
-
A blind spot regarding causal direction. A model may know that A and B frequently occur together, but not whether A causes B, B causes A, or both are caused by a common factor C.
-
Failure of intervention reasoning. You cannot effectively ask a model, “If the Federal Reserve cuts interest rates, what impact will that have on Southeast Asian technology stocks?” because it has learned historical co-occurrence patterns without a causal graph to support intervention calculations.
-
Lack of calibration. The “70% probability” produced by a model has not undergone calibration in the probabilistic sense. It is essentially a byproduct of the token distribution.
So far, existing approaches all have limitations: human collective intelligence, or prediction markets, requires liquidity, and prices are not reliable for niche questions; LLM pattern matching lacks causal structure; and end-to-end training suffers from outcome-oriented bias—a logically rigorous reasoning process that arrives at the wrong answer is penalized, while a rough judgment that happens to be correct is reinforced.
This is precisely why DigClaw exists.
DigClaw is building prediction infrastructure with causal structure as its backbone, probabilistic computation as its engine, and search intelligence as its data pipeline.
Its core assumption is that prediction should not be handled end to end by a single model.
Search, causal reasoning, and probabilistic inference are three orthogonal capabilities. They should be handled separately by three specialized systems and then combined in a structured way.
Three Top-Ten Finishes Validate Cross-Foundation-Model Transfer of Predictive Capability
FutureX is currently the most challenging real-time prediction leaderboard: every week, it publishes prediction questions about real-world events. When predictions are submitted, the ground-truth answers do not yet exist; settlement takes place afterward.
The dataset is hosted on Hugging Face, and the evaluation framework is open source on GitHub. The results are reproducible and verifiable.
Rhizome used the same prediction framework and operating conditions across three foundation models, generating and submitting answers independently for each, without aggregating results across models.
Therefore, these three rankings represent separate runs of the same system methodology on three foundation models, providing a clear cross-foundation-model comparison.
This does not mean that foundation models are unimportant.
Foundation models still provide general-purpose capabilities such as language understanding, reasoning, and tool use.
What this leaderboard result demonstrates is that the way retrieval is organized, time is handled, probabilities are expressed, evidence is maintained, and long-running operations are controlled can constitute system-level capabilities independent of model weights.
The following is DigClaw’s technical report on running the Rhizome framework on FutureX:
Rhizome Technical Report
Rhizome’s design is built around three engineering judgments: search and reasoning must be decoupled; prediction trajectories must be fully preserved and turned into calibration assets; and probability updates must be aware of causal structure.
△DigClaw prediction brain architecture
Design Decision One: Decoupling Search and Reasoning Across Multiple Models
Rhizome’s core design insight is that search quality and reasoning quality are orthogonal problems and should not be optimized by the same model simultaneously.
Leading DeepResearch Agents such as Perplexity and Gemini Deep Research bind search and reasoning within a single model: search quality is weighed down by the reasoning burden, while reasoning quality is contaminated by search noise.
Rhizome completely decouples the two. The search Agent is responsible only for finding relevant information, while the reasoning layer is responsible only for structured reasoning based on the available evidence.
Prediction tasks do not require information that can simply “accurately answer the user’s question.” They require the system to “discover as many relevant signals as possible.”
The search Agent is optimized for information relevance, not answer accuracy. If a search Agent is trained to “find the answer,” it will tend to locate content that looks like a conclusion—which is precisely what makes it dangerous.
Training uses a reinforcement learning framework, SearchRL, with two tracks iterating in parallel:
-
Track A: RL fine-tuning of open-source models—training 8B/30B-parameter open-source models with search relevance as the reward (see: Search-R1, COLM 2025; ReSeek, ICML 2026).
-
Track B: Harnesses for closed-source models—building external search-constrained frameworks for closed-source models such as Claude and GPT.
Different foundation models exhibit stable differences in prediction tasks: some are better suited to extended retrieval and complex reasoning, some excel at quantitative modeling, and others have advantages in cost and response speed.
Rhizome places the prediction protocol, Agent orchestration, tool calls, and result evaluation outside the foundation model. Organizations can select foundation models based on the value of the task and the scale of operation, making trade-offs among reasoning capability, cost, and response speed.
Design Decision Two: Trajectory Recording and Probability Calibration as Data Assets
For every prediction, Rhizome preserves a complete, versioned trajectory, including the question’s temporal conditions and settlement criteria, the evidence available at the time of prediction, the Agent orchestration and tool-calling process, the final answer and probability, and the corresponding model and system versions.
These records are created before the outcome is known, preserving Rhizome’s genuine judgment at a time when it did not yet know the answer.
After an event is settled, Rhizome places the prediction trajectory and the real-world outcome back into the same record. It reviews what information was available at the time, which counterevidence was missed, and whether the error occurred during retrieval, temporal reasoning, inference, answer formulation, or probability calibration.
A high-confidence error and an error made at roughly 50–50 may both be recorded as “incorrect,” but they reveal different problems.
Rhizome makes multiple independent predictions for the same question. The system does not simply average the results. Instead, it aggregates them in logit space and then adjusts their degree of extremity using the Brier Scores of settled questions.
The more independent the information provided by different trajectories, the more decisive the aggregate result can be. The more overlap there is among the evidence, the more conservative the adjustment.
△DigClaw prediction probability calibration method
Building on this, Rhizome uses Platt scaling to identify persistent overconfidence or excessive conservatism from settled questions and calibrate subsequent probabilities.
The calibration standard is intuitive: over the long term, events assigned a 60% probability should occur roughly 60% of the time; events assigned an 80% probability should occur roughly 80% of the time.
Each run is associated with the system version and key configurations in effect at that time, allowing changes in probability and anomalous results to be traced back to specific causes.
Foundation models can be upgraded or replaced, while the definitions of questions, evidence records, belief changes, settlement results, and calibration experience can remain continuous.
This type of feedback data cannot be generated in bulk after the outcome is known. Every sample must record a judgment before the future has occurred, and then wait for reality to provide the answer.
Code can be replicated; data assets accumulated over time cannot be accelerated into existence.
Design Decision Three: Continuous Updates with Causal-Chain Awareness
Before an event is settled, new data, policies, and market information continue to emerge. The system must determine whether the original probability should be updated.
Rhizome maintains a belief state for unsettled questions. Each piece of evidence records the event occurrence time, the content publication time, and the time the system read it.
When new information is completely duplicative of an existing record, the system skips the update. When new evidence conflicts with an old record, the old evidence is not deleted; the system preserves the process by which its judgment was revised.
Any single probability change of more than 0.15 must point to the specific new evidence that triggered the change.
There is another, more difficult question: Does new evidence represent several independent forces, or signals left by the same causal chain at different points?
An interest-rate hike announcement, changes in interest-rate differentials, and capital flows may occur one after another, but they are not necessarily three independent pieces of information.
If the system counts each of them separately in its probability update, the same cause may be counted repeatedly, pushing the probability toward overconfident extremes.
Rhizome is developing a causal-chain-aware Bayesian updating framework. Before evidence enters the probability update, the system first identifies the causal transmission chain to which it belongs, then adjusts the evidence weight according to relationships within that chain.
This framework consists of four layers:
-
Causal knowledge base: Stores validated causal chains and records the transmission lag, impact decay, and historical reliability of key relationships.
-
Within-chain signal deduplication: Subsequent signals from the same causal chain are not counted again at full weight, preventing the same force from being calculated multiple times.
-
Global posterior cap: Limits the cumulative impact of multiple pieces of evidence pointing in the same direction and adjusts the strength of the constraint according to the prediction horizon.
-
Transmission-lag awareness: After an event at the head of a chain occurs, the system updates progressively according to the progress of causal transmission rather than immediately assigning the full impact to the outcome at the tail of the chain.
△DigClaw causal-chain reasoning diagram
Within DigClaw’s internal prediction system, this framework has been prototyped and initially validated, and has guided several investment decisions.
Experiments show that, compared with direct Bayesian aggregation, within-chain signal deduplication and the global posterior cap reduced the overconfidence rate—the proportion of predictions assigned a probability above 85% that ultimately proved incorrect—from approximately 25% to 12%.
Why an Investment Firm Is Building a Prediction Model
Newborn Ventures was launched by DigClaw as an investment and incubation firm driven by AI-based discovery of Beta trends.
DigClaw’s trend-prediction foundation is the underlying technology framework supporting this AI-native VC.
The essence of investing is prediction.
Determining whether a sector will take off, whether a team will succeed, and whether a technology will become mainstream are all prediction problems.
Traditional investing relies on the experience, intuition, and information advantages of partners. DigClaw believes these judgments can be systematized and modeled.
From a broader perspective: investment returns = Beta (event/trend-driven) + Alpha (asset-specific).
Research into Alpha is relatively mature. But Beta—the prediction of macro events and trends—still lacks a genuinely effective AI solution. This is precisely the problem DigClaw aims to solve.
Once prediction becomes a callable, integrable, and calibratable parameter rather than an intuitive judgment, it can be embedded into far more decision-making scenarios than those currently visible.
-
For public companies, this means making strategic forecasts and issuing risk warnings before changes in the industrial chain and policy direction take shape;
-
For investment institutions, this means identifying value before consensus forms;
-
For government-guided funds, this means using systematic methods to assess industry trends and policy outcomes.
The same predictive capability is evaluated publicly on FutureX, drives investment decisions internally, and is applied to industry trend analysis, strategic risk assessment, and other scenarios.
For DigClaw and Newborn Ventures, reaching first place on FutureX is a starting point, not the finish line.
About DigClaw
DigClaw is an AI technology company focused on predictive intelligence. Its core mission is to build prediction infrastructure that is calibratable, auditable, and integrable.
Its flagship prediction framework, Rhizome, uses a three-layer architecture of causal reasoning, probability calibration, and search intelligence to achieve system-level capabilities independent of the foundation model. Foundation models can be replaced, while predictive assets continue to accumulate.
DigClaw’s prediction system has received external validation on public evaluation platforms such as FutureX. The same capabilities are being applied to investment decisions, industry trend analysis, strategic risk assessment, and other real-world scenarios.
About Newborn Ventures
Newborn Ventures is an AI-native early-stage venture capital and incubation firm.
Its core belief is that the essence of investing is prediction. Determining whether a sector will take off, whether a team will succeed, and whether a technology will become mainstream are all prediction problems that can be systematized and modeled.
Through its proprietary Deep Research Agent and causal prediction system, Newborn Ventures tracks innovative signals in AI reasoning worldwide and identifies structural opportunities that have not yet been priced in by the market.