The year/Independent research

Paper 2602.22638

MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios

Published
Feb 2026
Research lab
Independent
Citations
6
GitHub
157 stars

01 In brief

Summary

MobilityBench is a scalable benchmark for evaluating LLM-based route-planning agents in real-world mobility scenarios, built from 100,000 anonymized queries from Amap across 22 countries and over 350 cities.

It covers 11 task scenarios in four families: Basic Information Retrieval, Route-Dependent Information Retrieval, Basic Route Planning, and Preference-Constrained Route Planning.

To ensure reproducibility, it uses a deterministic API-replay sandbox that caches responses, eliminating non-determinism from live services.

The evaluation protocol assesses outcome validity (Delivery Rate, Final Pass Rate) plus instruction understanding, planning, tool use, and efficiency.

Experiments with 11 LLMs (GPT, Claude, Gemini, DeepSeek, Qwen) under ReAct and Plan-and-Execute frameworks show that models perform well on basic tasks but struggle with Preference-Constrained Route Planning.

ReAct achieves higher final pass rates but consumes ~35% more input tokens than Plan-and-Execute.

Scaling models improves performance, and enabling thinking modes boosts success rates at higher computational cost.

The benchmark and toolkit are publicly released.

02 From the paper

Abstract

Route-planning agents powered by large language models (LLMs) have emerged as a promising paradigm for supporting everyday human mobility through natural language interaction and tool-mediated decision making. However, systematic evaluation in real-world mobility settings is hindered by diverse routing demands, non-deterministic mapping services, and limited reproducibility. In this study, we introduce MobilityBench, a scalable benchmark for evaluating LLM-based route-planning agents in real-world mobility scenarios. MobilityBench is constructed from large-scale, anonymized real user queries collected from Amap and covers a broad spectrum of route-planning intents across multiple cities worldwide. To enable reproducible, end-to-end evaluation, we design a deterministic API-replay sandbox that eliminates environmental variance from live services. We further propose a multi-dimensional evaluation protocol centered on outcome validity, complemented by assessments of instruction understanding, planning, tool use, and efficiency. Using MobilityBench, we evaluate multiple LLM-based route-planning agents across diverse real-world mobility scenarios and provide an in-depth analysis of their behaviors and performance. Our findings reveal that current models perform competently on Basic information retrieval and Route Planning tasks, yet struggle considerably with Preference-Constrained Route Planning, underscoring significant room for improvement in personalized mobility applications. We publicly release the benchmark data, evaluation toolkit, and documentation at https://github.com/AMAP-ML/MobilityBench.