Hi! I am Kaizhen Tan, a Ph.D. student at New York University, supervised by Prof. Chenghe Guan and Prof. Zhan Guo. I received my master’s degree in Artificial Intelligence at Carnegie Mellon University and bachelor’s degree in Information Systems from Tongji University.

My research sits at the intersection of Urban Science and Embodied AI. Driven by the vision of harmonizing artificial intelligence with urban ecosystems, I aim to address the knowledge-to-action gap in digital cities: while urban digital systems are increasingly capable of monitoring conditions, modeling urban dynamics, and anticipating risks, they still struggle to support timely, place-based action. My work seeks to build spatially intelligent and socially aware urban AI systems that make cities more adaptive, inclusive, and governable.

My research integrates:

  • Paradigms: Robotic Urbanization, Agentic Urban Digital Twins, Multimodal Social Sensing, Spatial Intelligence
  • Methodologies: Representation Learning, Geospatial & Spatiotemporal Data Analysis, Agent-Based Simulation
  • Technical Foundations: LLMs, VLMs, AI Agents, World Models

Specifically, my research agenda is organized around four key topics:

🤖 1. Robotic Urbanization & Governance

How can dense cities integrate embodied intelligence while preserving safety, accessibility, and pedestrian experience?

  • Urban Readiness for Robots: Measure whether sidewalks, crossings, curbs, buildings, and public facilities can support safe robot operation.
  • Human-Robot Coexistence: Study conflicts, comfort, right-of-way, and interaction norms between robots, pedestrians, cyclists, and vulnerable groups.
  • Accessibility-Aware Deployment: Design routing and operation strategies that avoid reducing mobility for disabled people, older adults, and children.
  • Curbside and Low-Altitude Governance: Develop spatial rules for delivery robots and drones, including lanes, parking, corridors, privacy, noise, and safety constraints.
  • Public Acceptance and Accountability: Model public perception, responsibility boundaries, and governance mechanisms for city-scale deployment.

🏙️ 2. Agentic Urban Digital Twins

How can urban digital twins evolve from static city models into continuously updated systems for sensing, reasoning, and policy support?

  • Urban Foundation Representations: Fuse remote sensing, street-view imagery, trajectories, POI, IoT, text, and 3D data into unified urban representations.
  • Continuous Urban Sensing: Use robots, drones, mobile devices, and wearables as emerging data sources to update urban conditions over time.
  • 3D City Understanding: Support geo-localization, semantic mapping, and spatial querying across point clouds, meshes, and 3D Gaussians.
  • Urban Agents: Build LLM and VLM agents for map reasoning, spatial RAG, policy QA, public service assistance, and planning workflows.
  • Policy Sandbox: Enable what-if simulation, risk assessment, and implementation checks for urban management and public policy.

🎨 3. Multimodal Social Sensing

How can multimodal human-centered data reveal urban experience, social needs, and governance priorities?

  • AI-Enhanced Geospatial Analysis: Link urban form, environment, mobility, and public services with human behavior and social outcomes.
  • Pedestrian Experience and Accessibility: Detect walking barriers, sidewalk quality, perceived safety, and mobility challenges in everyday urban environments.
  • Urban Perception and Visual Aesthetics: Quantify streetscape quality, neighborhood imagery, and place identity to support design and regeneration decisions.
  • Socio-Cultural Signals: Extract place-based narratives from text, images, and online platforms to understand local identity and public concerns.
  • Participatory Governance: Translate social sensing results into explainable tools for planners, communities, and decision-makers.

🚀 4. Spatial Intelligence & World Models

How can spatial intelligence provide reliable reasoning, memory, and simulation capabilities for urban AI systems?

  • Embodied Spatial Representations: Unify geometry, semantics, physics, affordance, and action for robots, agents, and urban digital twins.
  • Urban World Models: Learn predictive models of how urban spaces change and how agents interact with physical and social environments.
  • Spatial Reasoning with VLMs: Improve map understanding, 3D reasoning, scene interpretation, and location-aware decision-making.
  • Lifelong Updating and Memory: Develop mechanisms for continuous learning, forgetting control, uncertainty tracking, and safe model updates.
  • Interpretable and Robust Decision Support: Make spatial AI systems transparent enough for planning, governance, and real-world deployment.

🔥 News

  • 2026.07: 🎉 Our paper CREG was accepted by the 9th Chinese Conference on Pattern Recognition and Computer Vision (PRCV 2026).
  • 2026.03: 🎓 I am pleased to share that I will begin my PhD at New York University in Fall 2026 under the supervision of Prof. Chenghe Guan and Prof. Zhan Guo.
  • 2026.01: 🎉 The abstract co-authored with Prof. Fan Zhang has been accepted for the XXV ISPRS Congress 2026. See you in Toronto!
  • 2025.12: 🎉 Our paper, led by my senior labmate Dr. Weihua Huan and co-authored with Prof. Wei Huang at Tongji University, was accepted by GIScience & Remote Sensing; honored to contribute as second author and big congratulations to Dr. Huan!
  • 2025.10: 🔭 Joined Prof. Yu Liu and Prof. Fan Zhang’s team at Peking University as a remote research assistant.
  • 2025.08: 🎉 Delivered an oral presentation at Hong Kong Polytechnic University after our paper was accepted to the Global Smart Cities Summit cum The 4th International Conference on Urban Informatics (GSCS & ICUI 2025).
  • 2024.04: 🔭 Began my academic journey at Prof. Wei Huang’s lab in the College of Surveying and Geo-Informatics, Tongji University.

📖 Education

New York University 2026.09 – 2031.05
Ph.D. student in Urban Science · New York / Shanghai
Carnegie Mellon University 2025.08 – 2026.08
M.S. in Artificial Intelligence · Pittsburgh
Tongji University 2021.09 – 2025.06
B.Mgt. in Information Systems · Shanghai

💼 Experience

🔭 Research

Research Assistant · NYU Shanghai, China
Designing embodied-intelligence-friendly urban spaces, and advancing collaborative governance through digital twins and urban agents that model complex city dynamics.
Research Assistant · Peking University, China
Automated urban element measurement from street view using VGGT and semantic segmentation, recovering metric scale via ground-plane fitting and camera-height calibration for large-scale data generation.
Research Intern · Singapore
Modeled air traffic controller communication tasks to predict workload from radiotelephony and trajectory data, integrating multi-source air traffic data with a CNN-Transformer model.
Research Assistant · Tongji University, China
Built a multimodal pipeline over social media reviews and photos to analyze tourist perception of historic quarters, fine-tuning segmentation models and applying sentiment analysis for multi-dimension satisfaction scores.
Research Assistant · Tongji University, China
Modeled traffic congestion propagation with spatiotemporal graphs and multi-scale community search, linking propagation patterns to built environment factors via causal analysis and POI-based features.

💻 Industry

AI Product Manager Intern · Shanghai, China
Ran end-to-end market and competitive analysis for AI products (Migo, Auto-Research) against 50+ vertical tools, built an operational KPI dashboard for retention and behavior funnels, and established evaluation frameworks for Migo's InternLM model.
Data Analyst · Shanghai, China

📝 Publications

Peer-Reviewed

UrbanVGGT
UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
Kaizhen Tan, Fan Zhang
XXV ISPRS Congress, 2026.
A measurement pipeline that estimates metrically scaled sidewalk width from a single street-view image through VGGT-based 3D reconstruction, semantic segmentation, and adaptive ground-plane fitting, reaching 0.25 m MAE on a Washington D.C. benchmark.
CREG
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
Kaizhen Tan, Yang Feng, Heqing Du
PRCV 2026.
A training-free diagnostic that maps vision-language model attributions onto a compass graph, measuring whether the evidence a model attends to aligns with the queried spatial direction. Higher accuracy does not imply better directional structure.
CacheWeaver
CacheWeaver: Cache-Aware Evidence Ordering for Efficient Grounded RAG Inference
Kaizhen Tan, Rong Gu, Mingyuan Li
GroundLM @ EMNLP 2026.
A prompt-layer reordering method that rearranges retrieved evidence so overlapping passages share prefix-cache hits, cutting median time-to-first-token by 20–33% across three vLLM configurations without touching the serving engine or the evidence set.
STALS
A Spatiotemporal Adaptive Local Search Method for Tracking Congestion Propagation in Dynamic Networks
Weihua Huan, Kaizhen Tan, Xintao Liu, Shoujun Jia, Shijun Lu, Jing Zhang, Wei Huang
GIScience & Remote Sensing, 2025.
A spatiotemporal adaptive local search (STALS) method combining dynamic graph learning and spatial analytics to trace how large-scale urban traffic congestion propagates through a road network.
Tourist perception framework
Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
Kaizhen Tan, Yufan Wu, Yuxuan Liu, Haoran Zeng
Global Smart Cities Summit cum The 4th International Conference on Urban Informatics (GSCS & ICUI), 2025.
A multimodal framework for reading tourist perception in historic Shanghai quarters, combining image segmentation, color theme analysis, and sentiment mining into four-dimensional satisfaction scores.
ATCO command lifecycle
Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
Kaizhen Tan
7th Asia Conference on Machine Learning and Computing, 2025.
A CNN-Transformer framework linking air traffic controller voice commands to aircraft trajectories, modeling command lifecycle and workload dynamics to support command generation and scheduling in terminal airspace.

Preprints & Under Review

Real street views compared with six text-to-image generators across six cities
GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation
Kaizhen Tan, Hanzhe Hong, Siru Tao
Submitted to NeurIPS 2026.
A benchmark of 7,117 curated Mapillary images over 109 named road segments in 25 cities, asking whether text-to-image models render this street or merely a plausible one. Street and neighborhood names help; GPS coordinates alone do not.
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong, Yang Feng, Heqing Du
Preprint.
Asks which physical parameters a latent world model actually recovers in its predictive representation, and which stay structurally unidentifiable however well the model predicts.
Token pruning calibration
When Does Visual Token Pruning Improve Calibration? The Role of Evidence Coverage in MLLMs
Kaizhen Tan, Yang Feng, Heqing Du, Hanzhe Hong, Siru Tao, Xin Xu
Preprint.
For multimodal LLM confidence, the selection rule matters more than the token budget: coverage-based pruning cuts expected calibration error on POPE with LLaVA-1.5 while holding accuracy, and the kept-set coverage tracks accuracy but not confidence.
Adapter capacity measurement framework
How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-Tuning
Kaizhen Tan, Heqing Du, Yang Feng
Preprint.
Measures in bits what a LoRA adapter can actually store: roughly two bits per trainable parameter, governed more by where the parameters sit than how many there are. On Qwen2.5, supervised fine-tuning memorizes verbatim while RL with verifiable rewards does not.

In Preparation

Qingdao Metro lost-and-found
What We Lose, What We Gain: Spatio-temporal Patterns of Lost-and-Found Items in Qingdao Metro
Kaizhen Tan, et al.
Manuscript in preparation.
34,333 lost-and-found records across 173 metro stations, read for spatio-temporal loss patterns, holiday effects, and station-level hotspots that transit operators can act on.
RoboROW
RoboROW: Simulating and Governing Right-of-Way for Urban Service Robots in Shared Urban Spaces
Kaizhen Tan, et al.
Manuscript in preparation.
When should a city give sidewalk robots their own lane? An agent-based simulation that screens shared versus dedicated right-of-way policies against pedestrian delay, comfort, and accessibility.
Framework for reviewing robotic urbanization in urban ground transportation systems: agents, spaces, and mechanisms
Robotic Urbanization and the Transformation of Ground Transportation Systems: Impacts, Challenges, and Opportunities
Kaizhen Tan, et al.
Manuscript in preparation.
A synthesis of how embodied machines entering streets, sidewalks, and curbs reshape ground transportation, and what that demands of infrastructure design, regulation, and accountability.

Technical Reports

sym
BlindNav: YOLO+LLM for Real-Time Navigation Assistance for Blind Users
Kaizhen Tan, Yufan Wang, Yixiao Li, Hanzhe Hong, Nicole Lyu
Technical report, 2025.
BlindNav is a real-time, camera-based navigation assistant that uses YOLO for street-scene detection and a local LLM to turn those signals into concise voice guidance for blind and low-vision pedestrians.

💬 Presentations

  • 2026.07 - XXV ISPRS Congress 2026
    UrbanVGGT: Scalable Sidewalk Width Estimation from Street View Images
    Toronto, Canada
  • 2025.08 - Global Smart Cities Summit cum The 4th International Conference on Urban Informatics (GSCS & ICUI 2025)
    Decoding Tourist Perception in Historic Urban Quarters with Multimodal Social Media Data: An AI-Based Framework and Evidence from Shanghai
    Hong Kong Polytechnic University (PolyU), Hong Kong SAR, China
  • 2025.07 - 7th Asia Conference on Machine Learning and Computing (ACMLC 2025)
    Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
    Hong Kong SAR, China

📫 Contact

  • Email(NYU): kt3275@nyu.edu
  • Email(CMU): kaizhent@cmu.edu

Please feel free to reach out if any of these research directions resonate with you. I'd be happy to chat!

🌍 Visitor Map