Ant International recently officially released version 2.0 of its self-developed time-series AI forecasting foundation model, “FalconTST” (Time-Series Transformer). FalconTST 2.0 achieved state-of-the-art (SOTA) results on authoritative global benchmarks, reducing mean absolute scaled error (MASE) to 0.666 and outperforming time-series forecasting models from several leading global technology companies. The model was purpose-built for foreign-exchange risk management in cross-border payments and will be expanded to additional industry scenarios, including e-commerce supply-chain demand forecasting and airline operations management.
Barclays, Citi, Deutsche Bank, Standard Chartered, and other major financial institutions have already applied FalconTST 2.0 to cash-flow forecasting and foreign-exchange management, improving their ability to forecast liquidity-risk exposure. FalconTST 2.0’s forecasting accuracy has remained consistently above 93%, and the model is capable of expanding into industries beyond finance, including logistics, aviation, and e-commerce.
Since 2025, Amazon Chronos-2, Google TimesFM 2.5, Salesforce Moirai 2.0, and IBM FlowState have successively brought capabilities such as multivariate forecasting, long-context processing, probabilistic forecasting, and cross-sampling-rate adaptation to the forefront. At the same time, the low signal-to-noise ratio, regime shifts, fat-tail risks, and cross-market linkages characteristic of financial markets make it difficult to directly convert advantages on general-purpose leaderboards into stable financial forecasting gains.
The evolution of Ant International’s Falcon TST is representative: in 2025, the model first completed business validation in foreign-exchange exposure and treasury management; in 2026, Falcon-2.0 improved univariate forecasting efficiency, while Falcon-X complemented it with heterogeneous multivariate relationship modeling. FalconTST 2.0 represents a specialized attempt to combine “general-purpose time-series foundation capabilities with financial-process validation.” Its continued iteration and steady progress demonstrate Ant International’s pioneering role in AI and finance.
I. Financial Time Series Are Putting the Real Capabilities of General-Purpose Foundation Models to the Test
General-purpose time-series foundation models already demonstrate clear cross-dataset transfer capabilities. The GIFT-Eval benchmark released by Salesforce AI Research in November 2024 covers 28 datasets, more than 144,000 time series, and 177 million test data points, together with approximately 230 billion pretraining data points that do not overlap with the test set. The benchmark has helped move the industry from “training one model for each dataset” toward “pretraining once and performing zero-shot forecasting across tasks.” Models such as Google TimesFM, Amazon Chronos, and Salesforce Moirai subsequently evolved rapidly, expanding their core objectives from single-task forecasting accuracy to cross-frequency modeling, cross-variable forecasting, and uncertainty estimation.
Financial markets pose a more stringent test for general-purpose models. Financial returns typically exhibit low signal-to-noise ratios, fat tails, volatility clustering, and regime shifts. Intraday prices, daily closes, yield curves, and macroeconomic data also suffer from inconsistent frequencies, holiday misalignment, and missing values. A large-scale empirical finance working paper published by researchers from the University of Manchester, University College London, and other institutions in November 2025 found that directly using off-the-shelf pretrained TSFMs for zero-shot forecasting and conventional fine-tuning produced relatively weak overall results. After additional pretraining on financial data, both forecasting performance and economic returns improved significantly. This finding indicates that general-purpose pretraining provides valuable priors for temporal patterns, but the financial domain remains highly dependent on the alignment between the training distribution and the target task.
The value of financial TSFMs should be distinguished from their ability to “generate stable trading returns.” In June 2026, an independent study compared models including TimeGPT, TimesFM 2.5, Moirai 2.0, Chronos, and Chronos-2 across five highly liquid U.S. stocks. The results showed that TSFMs achieved the best results in 8 of 10 tasks, but their overall gains over the random-walk baseline were modest, with only a small number of tasks passing significance tests. In addition, a locally supervised iTransformer still outperformed them on two META tasks. Financial institutions are therefore better served by treating TSFMs as transferable forecasting foundations and productivity tools for development, followed by secondary validation through domain data, rolling backtests, and risk constraints.
II. Mainstream TSFMs Have Shifted Toward Multivariate Forecasting, Long Contexts, and Probabilistic Prediction
Chronos-2 advances Amazon’s technical roadmap from univariate forecasting to general-purpose in-context learning. On October 20, 2025, Amazon released Chronos-2, a 120-million-parameter model. It natively supports univariate, multivariate, and covariate-conditioned forecasting, and uses Group Attention to share contextual information among related series.
It is important to distinguish between the two versions. The original Chronos did use the approach of “scaling and quantizing values into discrete tokens, then training with language-model cross-entropy.” Chronos-2 has introduced new group-attention and general-purpose in-context-learning mechanisms, so it cannot be evaluated simply by applying the criticism that “value bucketing limits tail accuracy.” Amazon’s reported “win rate of over 90%” refers to a head-to-head comparison between Chronos-2 and Chronos-Bolt, not to its win rate on the overall GIFT-Eval leaderboard. Likewise, Chronos-Bolt’s “up to 250× faster inference and 20× lower memory usage” is a comparison with the original Chronos at the same scale.
TimesFM 2.5 trades a smaller parameter count for longer context and continuous quantile forecasting. On September 15, 2025, Google Research released TimesFM 2.5, reducing the model size from 500 million parameters in version 2.0 to 200 million parameters, while extending the context length from 2,048 to as many as 16,384 tokens. It also added an optional 30-million-parameter continuous quantile forecasting head, supporting quantile forecasts for up to 1,000 steps.
Google subsequently restored support for XReg covariates in October 2025, added an Agent skills interface in March 2026, and introduced LoRA-based fine-tuning examples in April. Therefore, it is inaccurate to claim that “version 2.5 added quantiles, LoRA, and Agent capabilities all at once in 2026”; the model release and subsequent ecosystem enhancements occurred at different times.
Moirai 2.0 demonstrates that time-series foundation models do not need to keep accumulating parameters. Salesforce released Moirai 2.0 on August 8, 2025, replacing the early Moirai’s Masked Encoder with a Decoder-Only Transformer. It also changed the training objectives to quantile loss and multi-token prediction, while adding missing-value information embeddings, random patch masking, and data-quality filtering.
Salesforce’s small version has approximately 11.4 million parameters—about 96% fewer than the earlier Moirai Large—while delivering a 44% increase in inference speed. It also achieved leading MASE on the version of GIFT-Eval available at the time among models without test-data leakage. Its significance lies in showing that TSFM competition is shifting from “bigger” to “better-suited data, objective functions, and training strategies.”
FlowState uses state-space models to address Transformers’ limited adaptability to changing sampling frequencies. IBM Research presented the ICML 2026 paper FlowState on July 6, 2026. The model uses a state-space-model encoder and a functional-basis decoder to perform continuous-time modeling, dynamically adjusting its time scale as the input sampling rate and forecast horizon change. IBM’s public GIFT-Eval submissions include a 9.1-million-parameter version, substantially smaller than mainstream Transformer-based TSFMs.
This approach places greater emphasis on cross-sampling-rate adaptation and computational efficiency. The focus of the current public paper is not complex heterogeneous multivariate relationships, so its performance in financial scenarios still needs to be evaluated together with its cross-asset and cross-factor modeling capabilities.
III. Falcon TST’s Differentiation Is Already Supported by Validation in Financial Scenarios
Falcon TST first established validation in financial businesses before gradually developing into a publicly available model family. In May 2025, Ant International began collaborating with Barclays on TST-based foreign-exchange forecasting. In July, it conducted a pilot with Citi for foreign-exchange risk management for aviation clients. In August, Standard Chartered integrated TST into its SCALE liquidity engine. Falcon-1.0 was not officially released on Hugging Face until October.
This shows that Falcon’s development path differs from the typical TSFM approach of “first optimizing public benchmarks, then searching for industry applications.” Real-world treasury management and foreign-exchange exposure forecasting appeared earlier in its public development history.
Falcon-1.0’s core design uses a hierarchical mixture of experts to process multiscale temporal patterns. Ant International’s official repository defines Falcon-1.0 as a hierarchical mixture-of-experts model that handles different time scales through patch-level expert specialization and sample-level hierarchical routing.
The current Falcon website presents a product profile of 300 billion time points and 2.5 billion parameters, while joint announcements with multiple banks in 2025 generally described the model as having “nearly 2 billion parameters.” The two figures likely correspond to different model stages or statistical definitions. Because the public materials do not provide a complete comparison, this article does not treat parameter scale as core evidence of Falcon’s advantages. More reliable public evidence lies in its business forecasting capabilities across multiple time scales—including hourly, daily, and weekly cash flows and foreign-exchange exposure.
IV. Falcon-2.0 and Falcon-X Strengthen Univariate Efficiency and Multivariate Relationship Modeling, Respectively
Falcon-2.0 shifts the model’s focus toward univariate forecasting efficiency and probabilistic output. In July 2026, Ant International opened the Falcon-2.0 API. The company officially describes it as an Encoder-Only univariate TSFM trained with the ORBIT framework.
Compared with the original hierarchical mixture-of-experts architecture, Encoder-Only is better suited to completing forecasts directly through a single forward pass, reducing the latency and cumulative error associated with step-by-step autoregressive generation. The public API can directly output 21 quantiles, from 0.01, 0.05, and 0.10 through 0.95 and 0.99, and provides input_mask to explicitly mark missing values. This makes it more suitable for handling financial data affected by holidays, trading suspensions, or misaligned trading days.
Falcon-X addresses the heterogeneous multivariate modeling capability most critical to financial applications. The Ant International team published the Falcon-X paper on May 26, 2026. Its largest publicly disclosed experimental version has 591 million parameters. Falcon-X converts variables with different physical meanings into a unified latent-space representation, uses differential attention to identify both positive and negative correlations, then models relationships among variables and reconstructs the trajectory of each variable.
The paper reports a MASE of 0.687 on GIFT-Eval and evaluates multivariate forecasting on fev-bench. For financial applications, its more important significance is that the model directly begins to capture relationships among heterogeneous variables such as exchange rates, interest rates, volatility, and commodities, rather than relying solely on the historical information of individual series.
Multivariate inputs generate gains only when the relationships among variables are genuinely informative. In May 2026, researchers from Santa Clara University used Chronos-2 to conduct rolling forecasts from 2000 to 2025 for “leading U.S. technology stocks” and U.S. Treasury yields. The study found that joint input of related variables generally outperformed univariate input, but forecasting accuracy declined when stock and interest-rate panels were mixed directly.
This result is equally instructive for Falcon-X: a shared latent space can broaden the information base, but in financial production environments, variable selection, information timing, and rolling backtests must still be strictly controlled to prevent irrelevant factors from introducing noise into the model.
Quantile forecasting can be incorporated into risk-management processes, but it cannot be directly equated with regulatory VaR or ES. Both Falcon-2.0 and Falcon-X provide multiple quantile outputs, which can be used to construct pessimistic, neutral, and optimistic scenarios, as well as important inputs for value-at-risk and expected-shortfall models.
If the forecasting target itself is an asset return or portfolio profit and loss, lower quantiles can be used to estimate tail losses. However, in a real risk-control framework, coverage tests, conditional-coverage tests, stress scenarios, and model-risk management are still required to determine whether the quantiles are stably calibrated. The fact that an API can output 1% or 5% quantiles alone does not demonstrate that the model satisfies VaR or ES requirements.
V. Partnerships with Financial Institutions Are Beginning to Embed Forecasting Capabilities in Real Treasury Processes
The Barclays partnership validated the feasibility of integrating Falcon into a bank’s foreign-exchange hedging platform. On May 7, 2025, Ant International announced that Barclays had integrated the TST model into its BARX NetFX foreign-exchange hedging platform to improve the accuracy of Ant International’s foreign-exchange exposure forecasts.
Public materials state that TST can forecast cash flows and foreign-exchange exposure on an hourly, daily, and weekly basis, with forecasting accuracy exceeding 90% in Ant International’s own scenarios. The partnership’s core value lies in reducing unnecessary hedging and bank risk premiums through more accurate exposure forecasts, after which risk management is executed through the bank’s existing Guaranteed FX process.
The Citi pilot applies Falcon to online-ticketing foreign-exchange risk management for aviation clients. On July 18, 2025, Citi and Ant International announced a joint pilot combining Falcon TST with Citi Fixed FX Rates for airline clients selling tickets online across multiple currencies.
Citi’s official announcement confirmed that the first airline client had reduced hedging costs in actual transactions, citing Ant International data that hedging costs for pilot customers fell by approximately 30%. The partnership demonstrates that time-series foundation models are beginning to enter banking products as “forecasting as a service,” while final pricing and trading continue to be handled by Citi’s existing foreign-exchange systems.
The Standard Chartered partnership integrates Falcon into a 24-hour foreign-exchange liquidity-management process. On August 25, 2025, Standard Chartered and Ant International announced that Falcon TST would be integrated into the bank’s SCALE aggregated liquidity engine to enable real-time, 24-hour foreign-exchange exposure forecasting.
Standard Chartered stated that following the integration, the accuracy of Ant International’s foreign-exchange exposure forecasts exceeded 90%. It also disclosed that Falcon at the time covered more than 60% of Ant International’s transactions involving foreign-exchange conversion; related foreign-exchange costs fell by as much as 60%, while liquidity-management costs fell by as much as 50%. These cost-reduction figures reflect the business metrics disclosed by the two parties and are better understood as indicators of commercial deployment rather than directly comparable results on public academic benchmarks.
Capital A provides some of the clearest third-party enterprise-customer performance data currently available. On October 7, 2025, AirAsia, a Capital A company, disclosed in an official press release that after Falcon was used to forecast multicurrency cash flows and foreign-exchange exposure, forecast accuracy reached 90% on an hourly, daily, and weekly basis, while foreign-exchange hedging costs fell by as much as 40%.
Capital A stated that the project was Falcon’s first commercial deployment and disclosed that the aviation-industry version incorporated travel-industry data. The case shows that the practical value of industry specialization comes primarily from the combination of the model, industry data, bank transaction infrastructure, and corporate treasury rules—not from a single foundation model replacing the existing Treasury system.
VI. Leadership on General Benchmarks Cannot Be Directly Equated with Leadership in Financial Applications
GIFT-Eval remains an important window into the general capabilities of TSFMs, but leaderboard results must be interpreted in light of timing and data overlap. Salesforce created GIFT-Eval to standardize zero-shot testing, reduce data leakage, and compare models across different frequencies and numbers of variables. Since 2026, new versions of Falcon-X, Falcon-2.0, FlowState, and various Agent models have continued to be added, resulting in frequent changes to the rankings.
Some subsequent papers have also specifically distinguished between Chronos-2 results that may overlap with GIFT training data and stricter versions without leakage. For financial institutions, repeatedly citing “number one globally” at a particular point in time can easily be misleading. More meaningful questions are whether a model can remain stable across an institution’s own assets, frequencies, market regimes, and backtesting windows.
No single technical approach has established an absolute advantage on financial tasks at this stage. Chronos-2 excels at using related variables and covariates as context; TimesFM 2.5 combines long-context capabilities with a mature ecosystem to provide strong plug-and-play performance; Moirai 2.0 demonstrates that small models can remain competitive on general-purpose benchmarks; FlowState emphasizes cross-sampling-rate adaptation; and Falcon-2.0 and Falcon-X respectively focus on efficient univariate forecasting and heterogeneous multivariate relationship modeling.
Financial tasks differ in their requirements for extreme market conditions, temporal alignment, exogenous variables, quantile calibration, and deployment latency. Model selection should therefore be driven backward from specific business objectives.
Table 1. Public Technical Approaches and Financial-Application Observations for Mainstream Time-Series Foundation Models
| Model | Public version and scale | Core technical approach | Multivariate and exogenous information | Financial-application observations |
|---|---|---|---|---|
| Falcon TST 2.0 | Largest experimental version of Falcon-2.0: 585M; largest publicly disclosed experimental version of Falcon-X: 591M | Falcon-2.0 uses Encoder-Only + ORBIT; Falcon-X uses a shared latent prototype space | Falcon-X natively supports heterogeneous multivariate forecasting; both provide quantile-forecasting APIs | Already validated in foreign-exchange and Treasury scenarios; the complete Falcon-2.0 paper has yet to be released |
| Amazon Chronos-2 | 120M | Encoder-based Group Attention and in-context learning | Natively supports univariate, multivariate, and covariate-conditioned forecasting | Independent financial research shows that joint input of related variables is effective; irrelevant cross-market variables may introduce noise |
| Google TimesFM 2.5 | 200M; optional 30M quantile head | Decoder-Only, supporting contexts of up to 16k | Natively focused primarily on univariate forecasting, with covariates available through XReg | Strong average performance on financial benchmarks, although local supervised models still outperform it on some assets |
| Salesforce Moirai 2.0 | Small version: approximately 11.4M | Decoder-Only, using quantile loss and multi-token prediction | The Moirai family targets general-purpose multivariate forecasting; version 2.0 strengthens missing-value and robustness handling | Strong general-purpose efficiency; its publicly disclosed training corpus is not finance-specific |
| IBM FlowState | Public GIFT-Eval submissions include a 9.1M version | State-space-model encoder + functional-basis decoder | Focuses on cross-sampling-rate adaptation and variable forecast horizons | Suitable for cross-frequency tasks; complex heterogeneous multivariate relationships require additional validation |
Data sources: Official papers, technical blogs, and public materials for each model, as well as GIFT-Eval
Note: GIFT-Eval is a dynamic leaderboard, and some submissions are marked as having training-data overlap. To avoid presenting rankings from a particular stage as permanent conclusions, this table does not list a uniform “current ranking.”
VII. Competition in Financial TSFMs Will Increasingly Depend on Data Adaptation, Risk Calibration, and Engineering Deployment
The next stage of competition among financial TSFMs will shift from “whether they can perform zero-shot forecasting” to “whether they can be used reliably in real risk processes.” General-purpose models have already shown that large-scale cross-domain pretraining can significantly lower the barriers to time-series modeling. Empirical financial research has also repeatedly shown that domain data and task adaptation remain critical.
For scenarios such as treasury management, liquidity forecasting, and foreign-exchange exposure, the core evaluation criteria should gradually shift toward rolling-backtest stability, failure boundaries under extreme market conditions, quantile calibration, data-leakage control, inference costs, and system-integration capabilities. General-purpose leaderboards remain valuable, but they can answer only whether a model “can forecast,” not by themselves whether it “can enter a risk-control system.”
Falcon TST 2.0 has already established a relatively clear path toward financial applications. The main focus in 2025 was embedding the TST model into the real treasury-management and foreign-exchange operations of institutions including Barclays, Citi, Standard Chartered, and Capital A. In 2026, the focus shifted to using Falcon-2.0 and Falcon-X to address efficient univariate forecasting and heterogeneous multivariate forecasting, respectively.
Public evidence indicates that this path has a strong foundation for financial productization. Further information is still needed—including the complete Falcon-2.0 paper, the composition of its pretraining data, reproducible evaluations, and stress tests under extreme market conditions—to verify the sustainability of its claimed advantages.
Time-series foundation models are forming a market structure in which general-purpose foundations and industry-specific models coexist. General-purpose TSFMs such as TimesFM and Chronos are suitable for low-cost experimentation, cold starts, and rapid cross-industry transfer. Finance-specific models are better suited to production environments with abundant data, clearly defined metrics, and high risk sensitivity.
The significance of Falcon TST 2.0 lies in pushing the focus of competition further toward real financial processes. A model must not only understand temporal patterns, variable relationships, and uncertainty; it must also be absorbed into the existing risk-control, hedging, and liquidity-management systems of banks and enterprises. What financial institutions ultimately purchase is forecasting capability that can be validated, calibrated, governed, and continuously generate business value. Model size is merely one technical parameter among many.
VIII. Main Sources
-
Salesforce AI Research, GIFT-Eval: A Benchmark for General Time Series Forecasting Model Evaluation, 2024-11-12.
-
Amazon Science / GitHub, Chronos-2: From Univariate to Universal Forecasting / Chronos repository, 2025-10-20.
-
Google Research / GitHub, TimesFM 2.5 release and 2026 updates, 2025-09-15—2026-04-09.