Skip to content
Main Site News Console

Amap Releases ABot-Recon, Its First Streaming 3D Reconstruction Model Without Long-Range Dependencies: Reconstructs 10,000-Frame 3D Scenes from 12 Frames

· 量子位
国内AI

On August 28, Amap, a subsidiary of Alibaba Group, officially released ABot-Recon, the first streaming 3D reconstruction model capable of processing tens of thousands of frames without long-range dependencies. Without requiring long-term memory, the model can use only 12 consecutive frames of local views to reconstruct 3D scenes in real time and with high stability across tens of thousands of frames—enabling reconstruction as the camera moves. On authoritative benchmarks including KITTI, Oxford Spires, and VBR, ABot-Recon achieved SOTA performance with the lowest errors.

As technologies such as autonomous driving and embodied intelligence move into open environments, physical AI needs to do more than simply “see.” It must continuously understand its surroundings while moving: Where am I? What is around me? What should I do next? This is especially important in private-domain environments such as shopping malls, campuses, and warehouses, where positioning signals may be weak and conditions can change rapidly. In these scenarios, systems cannot rely on one-time mapping or reconstruction performed afterward.

At present, mainstream solutions use streaming 3D reconstruction to rebuild three-dimensional scenes in real time. However, this technology faces a range of challenges, including error accumulation and excessive GPU memory consumption. To maintain global consistency over long distances, conventional streaming 3D reconstruction methods typically establish memory anchors during reconstruction. The model must rely on long-range memory to continuously store, retrieve, and integrate historical information. As sequences become longer, inference slows down, accuracy declines, and GPU memory usage increases.

ABot-Recon takes an entirely new technical approach—“local pose prediction + residual optimization”—to address the challenges that long sequences pose to accuracy, speed, and GPU memory usage. Unlike comparable methods, ABot-Recon abandons memory alignment based on long-range memory anchors. Instead, the model uses only the latest 12 consecutive frames as a local context window. For each step, it predicts the local point cloud in the current camera coordinate system and the relative pose between adjacent frames, keeping computational complexity constant while ensuring a smooth trajectory. It then progressively assembles the complete global trajectory and 3D scene through online composition. Because the prediction targets are always local and independent of sequence length, the model’s memory usage and per-frame computational cost no longer increase as the video becomes longer.

To address error accumulation in local prediction, ABot-Recon introduces error-correction and error-constraining mechanisms on both the prediction and training sides, enabling real-time calibration of trajectory errors. With its innovative model design, ABot-Recon outperformed existing methods on multiple public long-sequence benchmarks.

In terms of accuracy, on the Oxford Spires long-sequence benchmark, ABot-Recon reduced the average trajectory error by 40.6% compared with LingBot-Map, a representative method in the field. Its relative pose rotation error, RPE-R, was as low as 0.12°, reaching the best level among comparable streaming methods and representing an approximately 40% reduction from the previous SOTA.

In terms of speed, ABot-Recon achieved real-time reconstruction at 24.45 FPS on the KITTI-02 test, with an inference speed 1.24 times that of the representative industry method. Among all methods with an average trajectory error below 20 meters, ABot-Recon delivered the highest throughput and the lowest GPU memory usage.

In terms of GPU memory, ABot-Recon’s peak GPU memory usage was only approximately 6.71 GB—about one-third that of comparable models. This means that a consumer-grade GTX 1080 Ti is sufficient to run the model, substantially lowering the barrier to using streaming 3D reconstruction.

Notably, ABot-Recon requires only monocular RGB video as input. Without any additional depth sensors or known camera parameters, it can perform real-time 3D reconstruction at the scale of tens of thousands of frames. This allows it to support private-domain mapping and base-map updates in the surveying and mapping industry, while also enabling broad applications in embodied intelligence, autonomous driving, 3D content production, and other scenarios.

Together, these improvements support a counterintuitive conclusion: helping a model go farther does not necessarily require it to remember more. A model that observes only 12 frames has instead found a more accurate, faster, and more efficient path. ABot-Recon has now open-sourced its inference and evaluation code as well as model weights on GitHub, and launched an experience space on the ModelScope community, making it easy for developers to try the model directly.

Appendix:

Project homepage: https://amap-cvlab.github.io/ABot-Recon-html

Technical report: https://github.com/amap-cvlab/ABot-Recon/blob/main/ABot-Recon-Tech-Report.pdf

ModelScope: https://modelscope.cn/studios/amap_cvlab/ABot-Recon/

This article was provided by Amap and reprinted by QbitAI with authorization. All views expressed are those of the original author.