TAU-HOME.COM
LOADING

LingBot-Map: Open Source Foundation Model for 20 FPS Real-Time Streaming 3D Reconstruction

Robbyant open-sourced LingBot-Map, an ECCV 2026 Best Paper candidate. It combines GCT and Paged KV Cache for 20 FPS streaming 3D reconstruction over 10k frames.

tau · October 9, 2026

#LingBot-Map #3DReconstruction #ComputerVision #ECCV2026 #OpenSource #FlashInfer #PyTorch

LingBot-Map: Open Source Foundation Model for 20 FPS Real-Time Streaming 3D Reconstruction

The Robbyant research team has open-sourced the official source code and pretrained weights for LingBot-Map, a feed-forward foundation model designed for real-time streaming 3D reconstruction, under the Apache-2.0 license. Nominated as an ECCV 2026 Best Paper Award candidate, the project enables simultaneous camera trajectory tracking and high-density 3D geometry reconstruction on continuous long-term video streams.

LingBot-Map real-time 3D reconstruction architecture diagram and point cloud visualization

Image source: Robbyant Team / GitHub

Conventional feed-forward 3D reconstruction pipelines frequently struggle with cumulative geometric drift and memory bottlenecks as frame counts grow into the thousands. LingBot-Map tackles these bottlenecks structurally through a dedicated geometric memory architecture and Paged KV cache optimization.

Geometric Context Transformer (GCT) Architecture and Drift Mitigation

At the core of LingBot-Map is the Geometric Context Transformer (GCT), which consolidates anchor context, pose reference windows, and trajectory memory into a single streaming framework.

  • Long-Range Drift Suppression: By jointly referencing salient historical geometric cues and camera trajectories during streaming input, the model mitigates coordinate distortion and accumulated drift across extended sequences.
  • Dense 3D Geometry Recovery: It generates fine-grained point clouds and 3D surface geometry via pure feed-forward neural inference without relying on iterative global bundle adjustment.

20 FPS Real-Time Streaming with Paged KV Cache

To handle long-horizon sequences in real time, LingBot-Map adapts the Paged KV cache mechanism—widely proven in large language model serving—to geometric attention operations.

  • Inference Speed: Achieves a stable real-time streaming throughput of approximately 20 FPS at 518×378 input resolution.
  • Long-Sequence Scalability: Maintains consistent memory consumption across long video sequences exceeding 10,000 frames without out-of-memory (OOM) failures.
  • Acceleration Backends: Recommends the FlashInfer (flashinfer-python) backend for optimal throughput, with fallback support for PyTorch native SDPA (Scaled Dot-Product Attention).

Model Checkpoints and Tooling Ecosystem

The project distributes pretrained model weights across HuggingFace and ModelScope for various deployment scenarios:

  • Available Checkpoints: General-purpose deployment checkpoint lingbot-map and pretraining stage-1 checkpoint lingbot-map-stage1 for fine-tuning.
  • Interactive Web Viewer: Includes a built-in Viser-based real-time 3D point cloud visualizer.
  • Sky Region Masking: Incorporates a dedicated ONNX sky-masking model to filter out outdoor atmospheric artifacts during 3D reconstruction.

Requirements and Operational Considerations

When deploying LingBot-Map for production workflows or extended sequences, keep the following environment and configuration caveats in mind:

  • Recommended Environment: PyTorch 2.8.0 and CUDA 12.8 are recommended. FlashInfer should be installed separately for maximum throughput, and an NVIDIA Kaolin build is required if using the offline batch renderer.
  • Pose Collapse Guard & Windowed Mode: Default settings do not enable state resetting, meaning camera pose tracking may degrade if the motion trajectory substantially exceeds the training distribution bounds. For long video inputs over 3,000 frames, enabling windowed mode (--mode windowed) or tuning the keyframe interval (--keyframe_interval) is recommended.

Sources