Artemis Panagopoulou, Student Researcher, and Mohit Goyal, Senior Software Engineer, Google We propose a system for synthetic data generation to train AI systems to visually follow any route on any map, finally teaching language models to navigate our world. Quick links MapTrace HuggingFace Dataset (2M question answer pairs Share Copy link × Look at a map of a shopping mall or a theme park. Within seconds, your brain processes the visual information, identifies your location, and traces the optimal path to your destination. You instinctively understand which lines are walls and which are walkways. This fundamental skill — fine-grained spatial reasoning — is second nature. For all their incredible advances, multimodal large language models (MLLMs) often struggle with this particular task. While MLLMs can identify a picture of a zoo and list the animals you might find there, they may have a difficult time tracing a valid path from the entrance to the reptile house. They might draw a line straight through an enclosure or a gift shop, failing to respect the basic constraints of the environment. …