Brian Hung
Projects
Geoguessr Arena
Evaluation of various LLMs on street view photos.
Playing a round
One run per model
ParallelRunner
GWS15k sample
128 of 15k panoramas, by seed
getStreetViewContentPart
the panorama, cached in R2
Single shot
one guess at temperature 0
N-shot
n guesses, geodesic medoid
Tool calling
getStreetViewPhoto, answer
Score the guess
distance in km, correctCountry
analyze-cross-model
ranks runs, 95% CI
gpt-5, claude-sonnet-4, gemini-2.5-pro and grok-4 each play the same 128 panoramas.
The leaderboard ranks by correct country first, then by mean distance.
Grading a screenshot answer
matchMatchesLabel
no
no
no
no
match
The agent's pick
a place_id from its searches
Audited labels
Google ids on 361 of 369
Same place_id
the exact Google listing
Within 150 m
of the label's coordinates
Same street address
70% of tokens, same number
Name tokens overlap
half of them, within 25 km
Correct
grade 1
Miss
grade 0
For a photo with no real place, the right answer is none, or a guess below 85 confidence.