Hi! I am Kaizhen Tan, a Ph.D. student at New York University, supervised by Prof. Chenghe Guan and Prof. Zhan Guo. I received my master’s degree in Artificial Intelligence at Carnegie Mellon University and bachelor’s degree in Information Systems from Tongji University.

My research sits at the intersection of Urban Science and Embodied AI. Driven by the vision of harmonizing artificial intelligence with urban ecosystems, I aim to address the knowledge-to-action gap in digital cities: while urban digital systems are increasingly capable of monitoring conditions, modeling urban dynamics, and anticipating risks, they still struggle to support timely, place-based action. My work seeks to build spatially intelligent and socially aware urban AI systems that make cities more adaptive, inclusive, and governable.

My research integrates:

  • Paradigms: Robotic Urbanization, Agentic Urban Digital Twins, Multimodal Social Sensing, Spatial Intelligence
  • Methodologies: Representation Learning, Geospatial & Spatiotemporal Data Analysis, Agent-Based Simulation
  • Technical Foundations: LLMs, VLMs, AI Agents, World Models

Specifically, my research agenda is organized around four key topics:

🤖 1. Robotic Urbanization & Governance

How can dense cities integrate embodied intelligence while preserving safety, accessibility, and pedestrian experience?

  • Urban Readiness for Robots: Measure whether sidewalks, crossings, curbs, buildings, and public facilities can support safe robot operation.
  • Human-Robot Coexistence: Study conflicts, comfort, right-of-way, and interaction norms between robots, pedestrians, cyclists, and vulnerable groups.
  • Accessibility-Aware Deployment: Design routing and operation strategies that avoid reducing mobility for disabled people, older adults, and children.
  • Curbside and Low-Altitude Governance: Develop spatial rules for delivery robots and drones, including lanes, parking, corridors, privacy, noise, and safety constraints.
  • Public Acceptance and Accountability: Model public perception, responsibility boundaries, and governance mechanisms for city-scale deployment.

🏙️ 2. Agentic Urban Digital Twins

How can urban digital twins evolve from static city models into continuously updated systems for sensing, reasoning, and policy support?

  • Urban Foundation Representations: Fuse remote sensing, street-view imagery, trajectories, POI, IoT, text, and 3D data into unified urban representations.
  • Continuous Urban Sensing: Use robots, drones, mobile devices, and wearables as emerging data sources to update urban conditions over time.
  • 3D City Understanding: Support geo-localization, semantic mapping, and spatial querying across point clouds, meshes, and 3D Gaussians.
  • Urban Agents: Build LLM and VLM agents for map reasoning, spatial RAG, policy QA, public service assistance, and planning workflows.
  • Policy Sandbox: Enable what-if simulation, risk assessment, and implementation checks for urban management and public policy.

🎨 3. Multimodal Social Sensing

How can multimodal human-centered data reveal urban experience, social needs, and governance priorities?

  • AI-Enhanced Geospatial Analysis: Link urban form, environment, mobility, and public services with human behavior and social outcomes.
  • Pedestrian Experience and Accessibility: Detect walking barriers, sidewalk quality, perceived safety, and mobility challenges in everyday urban environments.
  • Urban Perception and Visual Aesthetics: Quantify streetscape quality, neighborhood imagery, and place identity to support design and regeneration decisions.
  • Socio-Cultural Signals: Extract place-based narratives from text, images, and online platforms to understand local identity and public concerns.
  • Participatory Governance: Translate social sensing results into explainable tools for planners, communities, and decision-makers.

🚀 4. Spatial Intelligence & World Models

How can spatial intelligence provide reliable reasoning, memory, and simulation capabilities for urban AI systems?

  • Embodied Spatial Representations: Unify geometry, semantics, physics, affordance, and action for robots, agents, and urban digital twins.
  • Urban World Models: Learn predictive models of how urban spaces change and how agents interact with physical and social environments.
  • Spatial Reasoning with VLMs: Improve map understanding, 3D reasoning, scene interpretation, and location-aware decision-making.
  • Lifelong Updating and Memory: Develop mechanisms for continuous learning, forgetting control, uncertainty tracking, and safe model updates.
  • Interpretable and Robust Decision Support: Make spatial AI systems transparent enough for planning, governance, and real-world deployment.

🔥 News

  • 2026.07: 🎉 Our paper CREG was accepted by the 9th Chinese Conference on Pattern Recognition and Computer Vision (PRCV 2026).
  • 2026.03: 🎓 I am pleased to share that I will begin my PhD at New York University in Fall 2026 under the supervision of Prof. Chenghe Guan and Prof. Zhan Guo.
  • 2026.01: 🎉 The abstract co-authored with Prof. Fan Zhang has been accepted for the XXV ISPRS Congress 2026. See you in Toronto!
  • 2025.12: 🎉 Our paper, led by my senior labmate Dr. Weihua Huan and co-authored with Prof. Wei Huang at Tongji University, was accepted by GIScience & Remote Sensing; honored to contribute as second author and big congratulations to Dr. Huan!
  • 2025.10: 🔭 Joined Prof. Yu Liu and Prof. Fan Zhang’s team at Peking University as a remote research assistant.
  • 2025.08: 🎉 Delivered an oral presentation at Hong Kong Polytechnic University after our paper was accepted to the Global Smart Cities Summit cum The 4th International Conference on Urban Informatics (GSCS & ICUI 2025).
  • 2024.04: 🔭 Began my academic journey at Prof. Wei Huang’s lab in the College of Surveying and Geo-Informatics, Tongji University.

📖 Education

New York University 2026.09 – 2031.05
Ph.D. student in Urban Science · New York / Shanghai
Carnegie Mellon University 2025.08 – 2026.08
M.S. in Artificial Intelligence · Pittsburgh
Tongji University 2021.09 – 2025.06
B.Mgt. in Information Systems · Shanghai

💼 Experience

🔭 Research

Research Assistant · NYU Shanghai, China
Designing embodied-intelligence-friendly urban spaces, and advancing collaborative governance through digital twins and urban agents that model complex city dynamics.
Research Assistant · Peking University, China
Automated urban element measurement from street view using VGGT and semantic segmentation, recovering metric scale via ground-plane fitting and camera-height calibration for large-scale data generation.
Research Intern · Singapore
Modeled air traffic controller communication tasks to predict workload from radiotelephony and trajectory data, integrating multi-source air traffic data with a CNN-Transformer model.
Research Assistant · Tongji University, China
Built a multimodal pipeline over social media reviews and photos to analyze tourist perception of historic quarters, fine-tuning segmentation models and applying sentiment analysis for multi-dimension satisfaction scores.
Research Assistant · Tongji University, China
Modeled traffic congestion propagation with spatiotemporal graphs and multi-scale community search, linking propagation patterns to built environment factors via causal analysis and POI-based features.

💻 Industry

AI Product Manager Intern · Shanghai, China
Ran end-to-end market and competitive analysis for AI products (Migo, Auto-Research) against 50+ vertical tools, built an operational KPI dashboard for retention and behavior funnels, and established evaluation frameworks for Migo's InternLM model.
Data Analyst · Shanghai, China

📝 Publications

Peer-Reviewed

UrbanVGGT
UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
Kaizhen Tan, Fan Zhang
XXV ISPRS Congress, 2026.
A measurement pipeline that estimates metrically scaled sidewalk width from a single street-view image through VGGT-based 3D reconstruction, semantic segmentation, and adaptive ground-plane fitting, reaching 0.25 m MAE on a Washington D.C. benchmark.
STALS
A Spatiotemporal Adaptive Local Search Method for Tracking Congestion Propagation in Dynamic Networks
Weihua Huan, Kaizhen Tan, Xintao Liu, Shoujun Jia, Shijun Lu, Jing Zhang, Wei Huang
GIScience & Remote Sensing, 2025.
A spatiotemporal adaptive local search (STALS) method combining dynamic graph learning and spatial analytics to trace how large-scale urban traffic congestion propagates through a road network.
Tourist perception framework
Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
Kaizhen Tan, Yufan Wu, Yuxuan Liu, Haoran Zeng
Global Smart Cities Summit cum The 4th International Conference on Urban Informatics (GSCS & ICUI), 2025.
A multimodal framework for reading tourist perception in historic Shanghai quarters, combining image segmentation, color theme analysis, and sentiment mining into four-dimensional satisfaction scores.
ATCO command lifecycle
Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
Kaizhen Tan
7th Asia Conference on Machine Learning and Computing, 2025.
A CNN-Transformer framework linking air traffic controller voice commands to aircraft trajectories, modeling command lifecycle and workload dynamics to support command generation and scheduling in terminal airspace.
CREG
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
Kaizhen Tan, Yang Feng, Heqing Du
9th Chinese Conference on Pattern Recognition and Computer Vision (PRCV), 2026.
A training-free diagnostic that maps vision-language model attributions onto a compass graph, measuring whether the evidence a model attends to aligns with the queried spatial direction. Higher accuracy does not imply better directional structure.
CacheWeaver
CacheWeaver: Cache-Aware Evidence Ordering for Efficient Grounded RAG Inference
Kaizhen Tan, Rong Gu, Mingyuan Li
GroundLM @ EMNLP 2026.
A prompt-layer reordering method that rearranges retrieved evidence so overlapping passages share prefix-cache hits, cutting median time-to-first-token by 20–33% across three vLLM configurations without touching the serving engine or the evidence set.

Preprints & Under Review

Real street views compared with six text-to-image generators across six cities
GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation
Kaizhen Tan, et al.
Submitted to NeurIPS 2026.
A benchmark of 7,117 curated Mapillary images over 109 named road segments in 25 cities, asking whether text-to-image models render this street or merely a plausible one. Street and neighborhood names help; GPS coordinates alone do not.
Terminal pose of a failed sidewalk robot evaluated against clear-width and accessibility-feature criteria
Failure Has a Footprint: Rethinking Sidewalk Robot Failure in Public Space
Kaizhen Tan, et al.
Submitted to HRI 2027.
Defines the terminal footprint as the ground area a failed delivery robot occupies until retrieval and evaluates it against published clear-width thresholds and accessibility features. In 1,000 simulated failures on measured New York footways, stopping in place leaves 55.7% of failures below the ADA clear width, and an access-aware siting policy reduces the share to 18.8%. Across 17,136 km of New York footway, 28.0% of the network has no width-compliant pose under PROWAG.
Evidence needed to compare model-predicted and observed human responses to robot actions in public HRI releases
From Robot Actions to Human Responses: What Can Public HRI Data Support?
Kaizhen Tan, et al.
Submitted to HRI 2027.
Introduces Design × Response, an evaluation method that compares the model-predicted response change between robot actions with the observed human-response change for the same action pair, response, and population. Applied to seven public HRI releases, it shows that temporal specificity, action comparability, and population alignment provide independent evidence about whether a release can support model–empirical agreement.
PopNavShift pipeline from persona records to pedestrian motion profiles, and matched replay of one encounter under three controllers
PopNavShift: Stress-Testing Social Navigation Strategies under Population Shift
Kaizhen Tan, et al.
Submitted to ICRA 2027.
Replays identical physical episodes under different synthetic pedestrian-profile distributions to test whether robot navigation strategies keep their relative ranking. Across eight population conditions and 7,488 matched runs, changing pedestrian time pressure reverses 22.4% of mean pedestrian-delay orderings, compared with 8.6% of robot travel-time orderings.
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
Kaizhen Tan, et al.
Submitted to ICLR 2027.
Asks which physical parameters a latent world model actually recovers in its predictive representation, and which stay structurally unidentifiable however well the model predicts.
Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution
Kaizhen Tan, et al.
Submitted to ICLR 2027.
Defines physical meaning through responses to registered interventions and introduces a certificate that compares independently trained sensor mappings on held-out entities. On Cluster Haptic the same-surface audio–acceleration response distance is 4.5 times smaller than the wrong-surface distance for all 19 test surfaces, and in a controlled elastoplastic system a step-factorized executor generalizes to an unseen action order while a program-level executor is 38.5 times worse than an entity-blind predictor.
Token pruning calibration
When Does Visual Token Pruning Improve Calibration? The Role of Evidence Coverage in MLLMs
Kaizhen Tan, et al.
Submitted to AAAI 2027.
For multimodal LLM confidence, the selection rule matters more than the token budget: coverage-based pruning cuts expected calibration error on POPE with LLaVA-1.5 while holding accuracy, and the kept-set coverage tracks accuracy but not confidence.
Study design comparing street-view imagery with existing urban data across paired source conditions
Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing
Kaizhen Tan, et al.
Submitted to Computers, Environment and Urban Systems.
Compares image-based predictions with existing urban data across seven attributes, five public resources, and three VLMs, using paired source conditions, image replacements, and conflicting records. Existing data matched or exceeded image-only models for road damage, curb ramps, and house price, while images were more informative for building type, building function, and low-rise floor count, with the floor-count advantage rising 5.7 percentage points per doubling of distance to the nearest labelled building.
Study area with the eight Qingdao Metro services and 172 stations
Institutional Workflows and the Observed Geography of Metro Lost-Property Records: Evidence from Qingdao
Kaizhen Tan
Submitted to Journal of Transport Geography.
Examines 54,658 lost-property postings from the Qingdao Metro between July 2024 and July 2026 to ask where administrative records locate the events they describe. After accounting for station-area characteristics, terminal-only stations post at 8.76 times the rate of ordinary stations and terminal-plus-transfer stations at 21.15 times, consistent with recovery and attribution workflows shaping the observed geography.

In Preparation

RoboROW
RoboROW: A Right-of-Way Simulation and Policy Toolkit for Urban Service Robots
Kaizhen Tan, et al.
Manuscript in preparation.
A multi-agent co-simulation framework that translates robot right-of-way policies into spatial access masks, priority rules, motion parameters, and independent safety constraints, coupling JuPedSim pedestrians, SUMO traffic, and a ROS 2, Nav2, and Gazebo stack. The formal design compares nine spatial-allocation policies and 14 priority policies across three urban environments, seven sidewalk widths, and five pedestrian levels of service, with a 19-factor Sobol design estimating a joint deployment-feasibility boundary.
Framework for reviewing robotic urbanization in urban ground transportation systems: agents, spaces, and mechanisms
Robotic Urbanization and the Transformation of Urban Ground Transportation Systems: An Updated Review
Kaizhen Tan, et al.
Manuscript in preparation.
A PRISMA-informed systematic mapping review of Scopus and Web of Science literature from 2017 to June 2026, with a final synthesis corpus of 347 full texts. Road automation accounts for 88.5% of primary agent codes and simulation or optimization studies for 42.4% of methods, and the Agent-Space-Mechanism-Outcome framework organises the evidence around the urban interfaces where robots move, stop, load, wait, yield, and interact with people.
One Oakland Street View standpoint photographed over fifteen years, with model perception scores at each visit
You Cannot Photograph the Same Street Twice: Reliability Limits in Vision–Language Measurement of Urban Change
Kaizhen Tan, et al.
Manuscript in preparation.
Uses 4,648 consecutive-epoch Google Street View pairs from 435 standpoints in five US cities to measure how much a VLM perception score changes when the street itself does not. Re-photographing the same street moves a score by 0.80 points on average, equal to 66.5% of the difference between two streets in the same city, while aggregation over hundreds of pairs still recovers a coherent redevelopment signal.
Stimuli, trial structure, and the structure of urban scene appraisal ratings
Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment
Kaizhen Tan, et al.
Manuscript in preparation.
Tests predictive accuracy and neural alignment separately using EEG from 63 adults who rated 56 Berlin street scenes, across seventeen feature spaces. The best representation, DINOv2 ViT-B, reaches 29.6% of the noise-ceiling lower bound and a Gabor energy descriptor is indistinguishable from it, while the same embeddings predict held-out appraisal ratings at up to r = 0.87; the two measures do not track each other across models.
Rescaling the supplied world reference, response slopes across eight vision-language models, and training on the scale symmetry
Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning
Kaizhen Tan, et al.
Manuscript in preparation.
When every world-space quantity in a prompt is rescaled by a common factor, the correct metric answer changes by that factor, but eight vision-language models move only part of the way over four orders of magnitude. EquiSD uses this exact scaling relation as label-free supervision and raises a 3B model’s median response slope from 0.66 to 0.94 with one query per training video.

Technical Reports

sym
BlindNav: YOLO+LLM for Real-Time Navigation Assistance for Blind Users
Kaizhen Tan, Yufan Wang, Yixiao Li, Hanzhe Hong, Nicole Lyu
Technical report, 2025.
BlindNav is a real-time, camera-based navigation assistant that uses YOLO for street-scene detection and a local LLM to turn those signals into concise voice guidance for blind and low-vision pedestrians.

💬 Presentations

  • 2026.07 - XXV ISPRS Congress 2026
    UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
    Toronto, Canada
  • 2025.08 - Global Smart Cities Summit cum The 4th International Conference on Urban Informatics (GSCS & ICUI 2025)
    Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
    Hong Kong Polytechnic University (PolyU), Hong Kong SAR, China
  • 2025.07 - 7th Asia Conference on Machine Learning and Computing (ACMLC 2025)
    Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
    Hong Kong SAR, China

📫 Contact

  • Email(NYU): kt3275@nyu.edu
  • Email(CMU): kaizhent@cmu.edu

Please feel free to reach out if any of these research directions resonate with you. I'd be happy to chat!

🌍 Visitor Map