Paper 2605.14906
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
01 In brief
Summary
MEMLENS is a new benchmark for evaluating multimodal long-term memory in large vision-language models (LVLMs) and memory-augmented agents.
It comprises 789 questions across five memory abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge update, and answer refusal) at four context lengths (32K–256K tokens).
An image-ablation study confirmed that solving MEMLENS requires visual evidence, as removing evidence images dropped frontier LVLMs below 2% accuracy on the 80.4% of questions with image evidence.
Evaluation of 27 LVLMs and 7 memory-augmented agents revealed complementary failure modes: long-context LVLMs achieve high short-context accuracy but degrade as conversations grow, while memory agents are length-stable but lose visual fidelity under storage-time compression.
Multi-session reasoning caps most systems below 30% accuracy, and neither approach alone solves the task.
The results motivate hybrid architectures combining long-context attention with structured multimodal retrieval.
The benchmark is released with code and data for reproducibility.
02 From the paper
Abstract
Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capability: long-context LVLMs and memory-augmented agents. However, no existing benchmark conducts a systematic comparison of the two on questions that genuinely require multimodal evidence. To close this gap, we introduce MEMLENS, a comprehensive benchmark for memory in multimodal multi-session conversations, comprising 789 questions across five memory abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge update, and answer refusal) at four standard context lengths (32K-256K tokens) under a cross-modal token-counting scheme. An image-ablation study confirms that solving MEMLENS requires visual evidence: removing evidence images drops two frontier LVLMs below 2% accuracy on the 80.4% of questions whose evidence includes images. Evaluating 27 LVLMs and 7 memory-augmented agents, we find that long-context LVLMs achieve high short-context accuracy through direct visual grounding but degrade as conversations grow, whereas memory agents are length-stable but lose visual fidelity under storage-time compression. Multi-session reasoning caps most systems below 30%, and neither approach alone solves the task. These results motivate hybrid architectures that combine long-context attention with structured multimodal retrieval. Our code is available at https://github.com/xrenaf/MEMLENS.