Cross-View Geolocalization at VIS: Three Recent Papers
1 August 2026
When a hurricane, wildfire, or flood strikes, the first useful images rarely come from a satellite. They come from people on the ground who upload photos and posts within minutes, showing collapsed roofs, flooded streets, and blocked roads. What these images almost never include is a location. Only around one percent of tweets have ever carried a geotag, and this percentage decreased further after precise geotagging was discontinued in 2019. Without coordinates, a photograph cannot be placed on a map, and something that cannot be placed on a map cannot guide a rescue team.
One way around this is cross-view geo-localization: Take a ground-level photo and compare it to a database of geo-referenced satellite imagery to determine the location of the matching tile. The difficulty is that the two pictures have almost nothing in common visually. One is taken straight down from orbit at a uniform scale, while the other is taken at eye level, at an angle, and points at whatever the person deems alarming. Three papers from VIS and its collaborators, published over the past year, each address a different part of this problem.

They add a step in between. Rather than making a direct comparison, the first paper introduces a third viewpoint between the two: panoramic street-view imagery. This imagery is ground-level, like the photo, but geo-referenced, like the satellite tile. The team assembled MultiIan, a dataset centered on Hurricane Ian in Florida and Cuba. They gathered several thousand images from crowdsourcing platforms and social media. For posts without geotags, a large language model was used to extract the location from the text. A caption naming a street corner or landmark was often sufficient, and manual review was used for ambiguous cases. Two matching strategies were developed and compared and the work was published in the International Journal of Geographical Information Science by Wenping Yin, Fabian Deuser, Ziqi Liu, Jiaqi Wei, Xuanshu Luo, Martin Werner, Hao Li, and Yong Xue.
Adding a second step and more disasters. As we already showed is that one intermediate view helps, now the natural follow-up question arises whether a second one helps more and if any of this survives a change in disaster type. SAGINGeo, published in the IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, inserts drone-altitude imagery alongside street view. This allows satellite, aerial, and ground perspectives to be learned together in one shared representation rather than being compared pairwise. Building this required a dataset that did not exist. SAGINDisaster covers hurricanes, tornadoes, floods, and wildfires across the continental United States and Hawaii. It combines all four viewpoints at more than two thousand locations, totaling 118,000 images. The drone views were generated from 3D map data because flying real missions over that many past disasters is impossible. The paper was written by Wenping Yin, Fabian Deuser, Jiancheng Jiang, Ziqi Liu, Xuanshu Luo, Zhedong Zheng, Martin Werner, Hao Li, and Yong Xue.
Knowing when not to trust the answer. Both of the above return a single coordinate and nothing else. In a rescue context, this is problematic: a confident wrong answer could send a team to the wrong street, and there is no way to distinguish it from a correct answer. The third paper, published in the ISPRS Journal of Photogrammetry and Remote Sensing, modifies the model's output. Instead of one point, ProbGLC produces a probability distribution across the globe together with a score describing the image's initial ambiguity. Some photos are simply not locatable — for example, a wall of debris, a flooded field, or an area with no signage or skyline — and the model can now acknowledge this instead of making an incorrect guess. Images that the model is confident about can be sent straight to a crisis map, while the rest can be flagged for a human analyst or a drone to take a clearer picture. The authors are Hao Li, Fabian Deuser, Wenping Yin, Steffen Knoblauch, Wufan Zhao, Filip Biljecki, Yong Xue, and Wei Huang.

Taken together, the three papers trace a single line of work: first, make the match possible at all, second make it hold across different kinds of disasters and third make the model honest about its own uncertainty. The last of these may matter most in practice. Getting a photo within two hundred kilometers is not helpful for anyone waiting for assistance, and knowing that you have not yet arrived is as important as the estimate itself.
This research was conducted in collaboration with China University of Mining and Technology, Technical University of Munich, National University of Singapore, Heidelberg University, Wuhan University, University of Macau, HKUST (Guangzhou), Tongji University, University College London, and Shandong Territorial and Spatial Planning Institute.
Triple-objective cross-view geolocalization of disaster-related VGI: the case of Hurricane Ian
Wenping Yin, Fabian Deuser, Ziqi Liu, Jiaqi Wei, Xuanshu Luo, Martin Werner, Hao Li, Yong Xue
SAGINGeo: A Space-Aerial-Ground Integrated Framework for VGI Geolocalization in Multi-Disaster Scenarios
Wenping Yin, Fabian Deuser, J. Jiang, Ziqi Liu, Xuanshu Luo, Zhedong Zheng, Martin Werner, Hao Li, Yong Xue
Towards generative location awareness for disaster response: A probabilistic cross-view geolocalization approach
Hao Li, Fabian Deuser, Wenping Yin, Steffen Knoblauch, Wufan Zhao, Filip Biljecki, Yong Xue, Wei Huang