Independent research
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
This paper introduces Memory Decoder at Scale, scaling parametric long-term memory models up to 6.9B parameters and pretraining them on 300B tokens. To handle the computational bottleneck of constructing kNN distributions over 207B tokens, the authors develop a distributed Faiss pipeline using embedding compression, index sharding, and parallel search,…
Rubin Wei, Jiaqi Cao, Jiarui Wang, Junming Zhang, et al.- Published
- Jul 2026
- Citations
- 0
- Code
- 13 stars
