Visual AI can now read a city the way a planner walks it, only across millions of images at once. That is the argument of a new book from researchers at the MIT Senseable City Lab and Peking University, published by Routledge this month and covered by MIT News on 24 September 2026. The book is called “How AI Sees the City: Urban Visual Intelligence.”
Its authors are optimistic about what the technology can reveal, from traffic emissions to how much greenery people see on their way to work. They are also blunt about the risks: ubiquitous cameras, eroded privacy and algorithms that inherit the biases of the people who train them. “AI is not neutral,” says co-author Fábio Duarte.
This article explains what the book argues, how visual AI turns images into data, what the research has already found, where the dangers lie, and what councils, planners and businesses should ask before they point a camera at a street. For a related study on forecasting traffic, see A New AI Framework Could Help Cities Plan for Future Traffic.
Table of contents
- What the Book Is About
- From Kevin Lynch to Visual AI: A Long Tradition
- How Visual AI Turns Images Into Data
- The Promise: What Visual AI Has Already Found
- Visual AI Works Best Alongside Other City Data
- The Peril: Visual AI and Surveillance
- The Peril: Visual AI Is Not Neutral
- Privacy by Design: What Good Practice Looks Like
- What Visual AI Means for Councils, Planners and Businesses
- Questions to Ask Before Deploying Visual AI in Public Space
- Frequently Asked Questions About Visual AI and Cities
- References
What the Book Is About
“How AI Sees the City” is a short academic book, but its subject is large. It asks what happens when the enormous volume of images cities now produce becomes something computers can measure.
The authors
The book has four authors. Fábio Duarte is a principal research scientist and associate director of the MIT Senseable City Lab. Martina Mazzarello is a research scientist and leads the lab’s global initiatives. Carlo Ratti is a professor of the practice at MIT and the lab’s founder and director. Fan Zhang is an assistant professor at the Institute of Remote Sensing and GIS at Peking University.
The starting point
Routledge’s description notes that cities “have become vast repositories of digital imagery,” from surveillance cameras and smartphone photos to satellite images and street-level captures. The book presents what it calls “a comprehensive framework for applying visual AI in urban contexts,” with case studies from the United States, Stockholm, Amsterdam, Beijing, Dubai and Singapore.
The chapters
The book moves from history to method to ethics. After an introduction set on a walk through Cambridge, Massachusetts, it covers visual approaches to the city from ancient Rome to MIT, the digital image, using AI to interpret images, visual AI in the city, visual AI indoors, visual AI and the individual, surveillance, generative AI and the city, and how words and images work together.
Who it is for
The publisher aims it at scholars and postgraduate students in urban planning and landscape architecture, and at anyone interested in how technology shapes cities. Michael Batty of University College London called it a “fascinating book” that “shows how we are beginning to interpret the world of urban design.”
From Kevin Lynch to Visual AI: A Long Tradition
The authors deliberately place visual AI in a long history of looking at cities. That framing is an argument in itself.
Looking has always shaped planning
The book traces visual representations of cities from Roman marble maps to early photography, which produced influential images of Haussmann’s reshaping of Paris and of crowded tenements on New York’s Lower East Side in the 19th century. Pictures of cities have always changed how people thought about fixing them.
Lynch and Whyte
Two figures anchor the story. Kevin Lynch, a former MIT professor, published “The Image of the City” in 1960, a study of how people mentally map urban space. William H. Whyte, the sociologist behind “The Organization Man,” later became an urbanist who studied public spaces by watching them closely.
Scaling up the observer
“Kevin Lynch at MIT was only using paper and pen,” Duarte says. “We can now scale up what he was doing, with visual AI, while also looking at many different dimensions of cities.” Ratti adds that visual AI allows researchers “to observe cities at a scale and with a level of detail that was previously impossible.”
A tool, not an oracle
By putting visual AI in this tradition, the authors make a point. Powerful as it is, the technology remains a tool serving human purposes. It extends the planner’s eye rather than replacing the planner’s judgement.
How Visual AI Turns Images Into Data
The central idea of the book is simple to state. “We can treat these digital images as data and quantify features of the city,” Duarte says. “With computer vision techniques, each image is a dataset.”
What the software does
In practice, visual AI means software that finds and labels things in pictures. It can count vehicles and sort them by type, measure how much of a view is taken up by trees or sky, detect pedestrians and cyclists, or classify the style of a building. Each photo becomes a set of numbers that can be mapped and compared.
Where the images come from
Cities generate images from many sources, each with a different view and a different privacy risk. The table sets out the main ones discussed in the book and in the Senseable City Lab’s recent studies.
| Image source | What visual AI can measure | Privacy exposure |
|---|---|---|
| Satellite imagery | Tree cover, green space, building footprints | Low: people rarely identifiable |
| Street-level photos | Greenery people actually see, pavements, shopfronts | Medium: faces and plates appear |
| Traffic cameras | Vehicle counts, vehicle types, signal effects, emissions | Medium: depends on whether plates are read |
| Vehicle dashcams | Movement and congestion across a network | Medium to high: continuous street footage |
| Surveillance CCTV | Crowding, use of public space, safety incidents | High: designed to observe people |
| Listing and interior photos | Interior design styles, housing quality | Medium: private homes on public sites |
Connecting what is seen to how cities work
“The real promise of visual AI is not simply that computers can look at millions of images,” Zhang says. “It is that we can connect what is visible in those images — streets, buildings, greenery, traffic, public space — with larger questions about how cities function and how people experience them.”
The Promise: What Visual AI Has Already Found
The book is not only theory. The Senseable City Lab has published studies that show what visual AI can do with images cities already collect.
Traffic emissions, block by block
In a study of New York City, published in Nature Sustainability and described by MIT News in April, researchers used images from 331 cameras already installed at Manhattan intersections, together with anonymised location records from more than 1.75 million mobile phones. Sorting vehicles into 12 broad categories, the software placed 93% of them in the right category.
Why the detail matters
The team also tested what happens if detailed inputs are replaced with citywide averages. The rougher estimates varied from 49% below to 25% above the detailed results. As the study put it, seemingly small simplifications can introduce large errors. Visual AI earns its keep by removing those simplifications, street by street and hour by hour.
Measuring a real policy
New York introduced congestion pricing south of 60th Street in Manhattan in January 2025. Using their method, the researchers found that traffic volume fell by about 10%, but emissions fell by 16% to 22%. The gap matters: cutting stop-and-go traffic reduces pollution more than cutting the number of cars alone would suggest. A Cornell study had separately found a 22% drop in fine particulate matter inside the zone.
Emissions fell by 1.6 to 2.2 times as much as traffic volume did, using the study’s own ranges.
Greenery people actually see
Satellite images show how many trees a city has. They do not show whether people see them. The book examines how street-level images from phones and other sources can measure how much greenery people glimpse in everyday life, which research links to reported wellbeing. Visual AI can show that two neighbourhoods with similar tree cover feel very different at street level.
Interiors, too
Visual AI reaches indoors, too. Using images from 400,000 Airbnb listings around the world, a Senseable City study found that interior design styles are not becoming globally more alike, contrary to some claims. They still reflect significant geographic differences.
The questions planners can now answer
MIT News lists the kind of questions visual AI can address: why exactly traffic is snarling, what makes particular intersections dangerous, and which parts of plazas or parks attract the most people. These are questions planners have always asked. The difference is that they can now be answered everywhere at once, not one site at a time.
Visual AI Works Best Alongside Other City Data
One lesson from the Senseable City Lab’s work is easy to miss. Its strongest results come from combining images with other sources, not from cameras alone.
Cameras plus phones
The New York emissions study did not rely on traffic cameras by themselves. It combined camera images with anonymised location records from more than 1.75 million phones, which described overall movement across the city. The cameras showed what kinds of vehicles were on the road and how signals affected them. The phone data showed where and when people travelled. Neither source could have produced block-by-block emissions on its own.
Street level plus satellite
The same pattern holds for greenery. Satellite images measure tree cover from above. Street-level images measure what people actually see. Used together, visual AI can show where a city has trees that nobody notices and where a few well-placed trees make a street feel green.
Dashcams as a second network
In related work in Amsterdam, the team used dashboard cameras in vehicles to gather information about movement. “With our model we can make any camera used in cities, from the hundreds of traffic cameras to the thousands of dash cams, a powerful device to estimate traffic emissions in real-time,” Duarte said of that research.
Why combination matters for privacy too
Combining sources can reduce the need for identifying detail in any one of them. A system that knows vehicle types from cameras and movement patterns from anonymised phone data does not need to read licence plates to answer its question. Good design uses each source for what it does best and nothing more.
The Peril: Visual AI and Surveillance
“Across cities, more images means more data, more insight — and more concerns about privacy and fairness,” MIT News summarises. The book’s chapter on surveillance is titled “Eyes on the City.”
Camera density varies enormously
MIT News notes that London, an early adopter of CCTV, has about 210 cameras per square mile. Eight of the world’s ten most camera-heavy cities are in China, and Shanghai has more than 5,000 cameras per square mile. Debate over traffic cameras has also flared in the US this year.
Shanghai’s figure is a floor, so the gap is at least 24 to 1.
The trade-off the authors name
On the safety case for intensive video recording, the authors write that “the benefits must be weighed against the significant erosion of personal freedom and the potential for abuse inherent in a system of constant monitoring.” Visual AI makes that trade-off sharper. A camera that once needed a person to watch it can now be analysed automatically, around the clock.
Analysis changes what a camera is
A traffic camera watched by nobody records very little that matters to privacy. The same camera feeding a visual AI system that tracks movement, recognises vehicles or estimates who is in a crowd is a different thing. The hardware is unchanged. The capability is not.
The Peril: Visual AI Is Not Neutral
The second danger the book identifies is bias. It is quieter than surveillance, but it may be harder to fix.
Training shapes what a model sees
“We need to teach AI to see, and depending on how you teach it, it will see what is embedded in the culture,” Duarte says. “AI is not neutral.” If visual AI systems are trained mainly on majority population groups, they may not evaluate minority groups the same way.
Data that confirms assumptions
MIT News notes that biased systems can produce data that reinforces prior perceptions as much as it reflects reality, about people, neighbourhoods and whole cities. A model that has learned to associate certain street features with danger can mark a neighbourhood as unsafe simply because it looks like other places labelled that way.
People are not neutral either
Mazzarello offers a counterweight. “Our eyes are not neutral, either,” she says. “Every tool has to be guided in the right way, and trained in the best way.” The point is not that human observers were fair and machines are biased. It is that visual AI can scale up whatever bias it inherits, faster and further than any single observer.
Privacy by Design: What Good Practice Looks Like
The book’s authors do not argue that cities should stop using visual AI. They argue it should be used “wisely, critically, and creatively.” The emissions study offers a model.
Recognise types, not people
The New York study recognised types of vehicles, “but without compiling license plate numbers,” according to MIT News. It answered its question, emissions by street and hour, without identifying anyone. That is privacy by design: collect the least identifying data that still answers the question.
Use existing infrastructure carefully
The study used cameras already installed at intersections. That is cheaper and avoids new cameras, but it also means repurposing footage for uses the public may not expect. Clear public notices about secondary uses help maintain trust.
Aggregate early
A system that turns footage into counts on the device, and throws the footage away, carries far less risk than one that stores video for later analysis. Where possible, visual AI should output statistics, not images.
Know the legal frame
In the UK, filming identifiable people is covered by UK GDPR, and the Information Commissioner’s Office publishes guidance on video surveillance. In the EU, the AI Act restricts the use of real-time remote biometric identification in public spaces for law enforcement to narrow exceptions. Anyone planning a visual AI project should take advice early. Our data protection team helps organisations assess these projects before deployment.
What Visual AI Means for Councils, Planners and Businesses
The book is written for academics, but its lessons apply to anyone collecting images in public or semi-public spaces.
For councils and planners
Visual AI offers cheaper, more detailed evidence for decisions on traffic, air quality and public space. The New York work shows how existing cameras can measure the effect of a policy within weeks. Planners exploring immersive tools may also find our piece on AI-Powered VR Helps Planners Design Better Cities useful.
For property owners and retailers
Many businesses already run cameras in car parks, shopping centres and offices. Adding analytics turns security footage into footfall and usage data. The same privacy and bias questions apply, at a smaller scale, and staff and visitors need to be told.
For technology suppliers
Buyers will increasingly ask how a visual AI product handles faces and licence plates, how its models were trained, and how accurate they are for different groups of people. Suppliers that can answer clearly will win contracts. Our data analytics work often starts with exactly these questions.
Questions to Ask Before Deploying Visual AI in Public Space
The book ends with a call to explore the technology “wisely, critically, and creatively.” In practice, that means asking hard questions before a project starts. The checklist below turns the book’s concerns into practical steps.
| Question | Why it matters | Good answer |
|---|---|---|
| What question are we answering? | Stops data being collected just in case | One stated purpose per dataset |
| Do we need to identify anyone? | Identification drives most of the risk | Types and counts, not faces or plates |
| Is footage kept or turned into numbers? | Stored video can be reused for other ends | Aggregate on capture, delete raw footage |
| How was the model trained? | Training data shapes what the model sees | Documented data sources and accuracy by group |
| Have we told the public? | Secondary uses of cameras erode trust | Clear notices and a published purpose |
| Who reviews the results? | Automated findings can confirm bias | A person checks conclusions before action |
Test against the simpler method
The New York study compared its detailed method with citywide averages and found errors ranging from 49% below to 25% above. Every visual AI project should include a similar test. If the simple method gives the same answer, the cameras may not be worth the risk.
Plan for the end of the project
Data collected for a pilot often outlives it. Decide at the start when footage and derived data will be deleted, and who can authorise keeping them.
Frequently Asked Questions About Visual AI and Cities
What is visual AI?
Visual AI is software that analyses images and video to find, label and measure things in them, such as vehicles, trees, people or buildings. In cities it turns photos and camera feeds into data planners can use.
What is “How AI Sees the City” about?
It is a 2026 Routledge book by Fábio Duarte, Martina Mazzarello, Carlo Ratti and Fan Zhang. It sets out a framework for using visual AI to study cities, with case studies from the US, Europe, China, Dubai and Singapore, and examines surveillance, privacy and bias.
What has visual AI revealed about cities?
In New York, it showed that congestion pricing cut traffic by about 10% but emissions by 16% to 22%. Other studies measured how much greenery people see and found that interior design styles still vary by region.
What are the main risks of visual AI in cities?
The book names two: widespread visual surveillance that erodes personal freedom, and bias, where models trained mainly on majority groups evaluate other groups differently.
How many surveillance cameras do cities have?
MIT News reports about 210 cameras per square mile in London and more than 5,000 per square mile in Shanghai. Eight of the world’s ten most camera-heavy cities are in China.
Can visual AI be used without invading privacy?
It can reduce the risk. The New York emissions study recognised vehicle types without recording licence plates. Collecting counts rather than identities, deleting raw footage and telling the public all help.
References
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.