Urban design shapes how much we walk, what we eat, who we meet, the air we breathe and how safe our streets feel. As planners and public health teams start asking chatbots built on a large language model for advice on those decisions, a new study asks whether that advice is ethical, not just technically sensible. The answer, from researchers at the Japan Advanced Institute of Science and Technology (JAIST) and Waseda University, is a qualified yes: the AI avoided obvious harm and treated poorer neighbourhoods fairly, but it often forgot to involve the community or to say that humans should make the final call, especially when money was tight.

The paper, “Ethical assessment of large language model-generated advisory text on designing the built environment for health”, appeared online on 10 August 2026 and will be published in volume 27 of the journal Developments in the Built Environment on 1 October. EurekAlert and Tech Xplore carried the release on 25 September. The authors describe it as the first study worldwide to examine the ethics of AI-generated advice on urban design and health. This article explains what they did, what they found, and what it means for councils, consultancies and health bodies that are already using large language models in planning work.

Why Urban Design for Health Is Turning to AI

urban design for health llm ethical advisers b bicycle rack with three hoops

Large language models are increasingly used to write advice across urban design, planning and public health, the JAIST release notes. They can recommend changes to transport infrastructure, land use, walkability, pedestrian environments and green space, all of which affect health. Because that text can read like expert advice, the question is whether it meets the ethical standards expected of human experts.

The built environment is a health intervention

The link between urban design and health is well established. The Lancet’s 2016 series on urban design, transport and health, led by Billie Giles-Corti, framed city planning as “a global challenge” for population health, because street layouts, density, transport and green space drive physical activity, air quality, road injuries and social contact. Decisions about a cycle lane or a park are, in effect, public health decisions.

Why AI advice raises new questions

A planner who asks a chatbot how to make a low-income neighbourhood healthier gets a fluent, confident answer in seconds. That is useful for brainstorming. It is risky if the answer quietly recommends something harmful, gives richer areas better options, or skips the people who will live with the result. Our earlier coverage of AI-powered VR tools for planners and the promise and peril of visual AI in cities shows how quickly these tools are entering the field.

Where AI Already Shows Up in Urban Design Work

urban design for health llm ethical advisers c street lamp with a curved arm

The study matters because AI is no longer a future question for planning teams. Large language models are already being used for tasks around urban design, often informally, by staff who find them quicker than a literature search.

Drafting and summarising

The most common uses are writing tasks: first drafts of design statements, summaries of consultation responses, plain-English explanations of technical reports and lists of options for a site. Each of these shapes what decision-makers read, even if a human signs it off.

Analysis and imagery

Other tools use computer vision to analyse street imagery, model traffic or generate visualisations of proposed schemes. We have covered AI that helps cities plan for future traffic and the risks of reading cities through computer vision. Text advice is different: it tells people what to do, not just what is there, which is why the ethics of that advice needs its own test.

Informal use is the risk

The riskiest use of AI in urban design is the unrecorded kind. An officer who asks a chatbot for “the best interventions for a deprived estate” and pastes the answer into a briefing has imported the model’s blind spots without anyone knowing. The JAIST study shows what those blind spots look like.

How the Researchers Tested AI Urban Design Advice

urban design for health llm ethical advisers d terrace of three houses facing a lawn

The team was led by Associate Professor Mohammad Javad Koohsari, founder of the Urban Design Science for Health Laboratory at JAIST, and Professor Koichiro Oka of the Faculty of Sport Sciences at Waseda University. The other authors are Becky P.Y. Loo, Jing Zhao, Jiuling Li, Ying Long, Yi Lu and Andrew T. Kaczynski. They evaluated ChatGPT’s responses to urban design prompts across different health pathways, neighbourhood incomes and budgets.

Study elementDetail
ModelChatGPT (described in the paper as “a recent LLM”)
Health pathwaysPhysical activity, dietary intake, social interaction, air pollution, traffic safety and crime, noise
Prompts18 in total: 12 for higher-income and lower-income neighbourhoods, 6 for mixed-income neighbourhoods with a budget constraint
RepetitionsEach prompt run 10 times
Answers analysed180
Ethical criteriaNon-maleficence, distributive justice, collective participation, transparent oversight
CodingTwo co-authors independently coded every answer
MethodContent analysis

Six ways the built environment affects health

The six pathways cover the main routes through which urban design influences health. Physical activity depends on walkable streets and safe cycling. Dietary intake depends partly on access to food. Social interaction depends on public space where people meet. Air pollution, traffic safety, crime and noise are all shaped by how roads, buildings and land uses are arranged.

The table below summarises the kind of urban design levers each pathway usually involves. It is our summary of common planning practice, to show what the AI was being asked to advise on, not a list of the study’s prompts or answers.

Health pathwayTypical urban design leversWho tends to gain or lose
Physical activityConnected street networks, safe cycle routes, nearby destinations, parksAreas with poor walking routes gain most
Dietary intakeAccess to food shops and markets, community growing spaceNeighbourhoods with few healthy food outlets
Social interactionSquares, benches, libraries, community spacesOlder and isolated residents
Air pollutionTraffic reduction, greenery along busy roads, siting of schools and homesHomes and schools next to main roads
Traffic safety and crimeLower speeds, crossings, street lighting, active frontagesChildren, older people, pedestrians
NoiseRoad surfaces, building layout, buffers between homes and trafficDense housing beside transport corridors

Every lever on that list has distributional effects. A new cycle route or a low-traffic scheme helps some residents and inconveniences others, which is why good urban design depends on participation and why the study’s procedural criteria matter.

Rich, poor and mixed neighbourhoods

The design of the prompts is what makes the study useful. By asking the same question for higher-income and lower-income areas, the researchers could check whether the AI offered poorer places weaker proposals. By adding a set of mixed-income prompts with a budget limit, they could see what the model dropped when it had to prioritise, which is the everyday reality of most urban design projects.

The Four Ethical Tests for Urban Design Advice

urban design for health llm ethical advisers e magnetic compass with a needle

The answers were judged against four criteria. Two concern the substance of the advice and two concern the process it recommends.

CriterionQuestion it asksType
Non-maleficenceDoes the advice avoid recommending harmful or unsafe changes?Substantive
Distributive justiceAre lower-income neighbourhoods offered proposals as strong as higher-income ones?Substantive
Collective participationDoes the advice involve the community in decisions?Procedural
Transparent oversightDoes it acknowledge uncertainty and the need for human, professional oversight?Procedural

Substance and process

Non-maleficence and justice are familiar from medical ethics, where “first, do no harm” and fair distribution are core principles. Participation and oversight come from planning and governance. Sherry Arnstein’s 1969 “Ladder of Citizen Participation” made the case that consultation without real power is tokenism, and that idea still underpins how good urban design processes are judged. Advice that proposes a fine park but ignores the residents who will use it fails that test.

Why oversight belongs in the advice

It might seem odd to expect an AI to say “check this with a professional”. But the World Health Organization’s 2021 guidance on the ethics and governance of AI for health sets out six consensus principles, including human autonomy, transparency and accountability. Advice that presents itself as final, without flagging uncertainty or the need for human decision-makers, invites over-reliance.

What the Study Found About AI and Urban Design Ethics

urban design for health llm ethical advisers f drinking fountain on a pedestal

The results split cleanly between the substantive and the procedural criteria.

Share of ChatGPT urban design answers meeting each ethical criterion
Non-maleficence (180 of 180) 100%
Distributive justice (110 of 120 evaluable units) 91.7%
Transparent oversight (136 of 180) 75.6%
Collective participation (109 of 180) 60.6%

No harmful advice

Non-maleficence was satisfied in all 180 answers. The model did not clearly recommend any harmful or unsafe changes to the built environment. For a tool that some fear will hallucinate dangerous suggestions, that is a reassuring baseline.

Poorer areas were treated fairly

Distributive justice was met in 110 of 120 evaluable units, or 91.7%. Lower-income contexts were “rarely given weaker proposals” than higher-income ones. That matters, because a model trained on text about wealthier cities could easily have defaulted to ambitious schemes for rich areas and minimal fixes for poor ones.

Participation and oversight lagged

The procedural criteria were weaker. Collective participation appeared in 109 of 180 answers (60.6%), and transparent oversight in 136 of 180 (75.6%). In roughly two answers in five, the AI proposed changes to a neighbourhood without suggesting the people living there should have a say.

The Budget Effect: When Money Is Tight, Process Disappears

The clearest finding came from the budget-constrained prompts. When the model had to prioritise within a limited budget for a mixed-income neighbourhood, the procedural safeguards fell away.

How a budget constraint changed the advice (share of answers)
Collective participation, no budget constraint (88 of 120) 73.3%
Collective participation, with budget constraint (21 of 60) 35.0%
Transparent oversight, no budget constraint (108 of 120) 90.0%
Transparent oversight, with budget constraint (28 of 60) 46.7%

The counts for the unconstrained prompts follow directly from the totals in the release: 109 participation answers overall minus 21 in the budget set leaves 88 of 120, and 136 oversight answers minus 28 leaves 108 of 120. Participation fell by about half, from 73.3% to 35.0%, and oversight from 90.0% to 46.7%.

Why this matters for real projects

Budget pressure is the normal condition of urban design, not an edge case. The moment the model was asked to choose, it tended to drop the steps that make choices legitimate: asking residents and flagging that experts should decide. That is exactly when those steps matter most, because prioritising means someone loses out.

What Koohsari says

“Our findings show that ethical urban design for health depends not only on what physical changes are proposed, but also on how decisions are made, who is involved, and how uncertainty and human oversight are addressed,” Dr Koohsari said. The paper’s abstract puts the conclusion carefully: current AI outputs “may reproduce some baseline ethical conventions in urban design discourse but are less reliable on procedural concerns.”

Why Process Ethics Is Harder for AI Than Avoiding Harm

The study reports what the model did, not why. But the pattern has a plausible explanation, and it is worth setting out as our interpretation rather than the authors’ finding.

Answering the question asked

A chatbot is built to answer the question in front of it. Asked what physical changes would improve health, it lists physical changes. Harm avoidance and fairness are properties of those changes, so they show up naturally in the answer. Participation and oversight are properties of the decision process around the changes, and a model will only mention them if it treats process as part of the question.

Constraints push out context

A budget constraint makes the question sharper: what should come first? A model optimising for a direct, useful answer tends to rank options and stop. The steps that make urban design decisions legitimate, such as asking residents or noting that officials must decide, look like padding to a system trying to be concise. That may be why they fell away under budget pressure.

What that means for tool builders

If this reading is right, the fix is partly in design. Tools built for urban design and public health work could build participation and oversight prompts into their templates by default, rather than relying on users to ask. The authors’ call to judge future tools on process as well as recommendations points the same way.

Are Large Language Models Ethical Advisers for Urban Design?

The study’s headline question deserves a direct answer. On this evidence, a chatbot is a reasonable ethical first draft, not an ethical adviser.

What the AI does well

It avoids obvious harm and, in this test, it did not short-change poorer neighbourhoods. The researchers say large language models “could serve as an initial input” for urban designers, planners and public health professionals considering changes to streets, public spaces, transport infrastructure, land use and green space.

What it cannot replace

The JAIST release adds that such outputs “should not replace professional judgment or community participation, particularly when limited budgets require prioritization.” An adviser who stops mentioning the community whenever money is short is not one a council should rely on unsupervised.

The conditional optimism

Koohsari’s closing view is hopeful but conditional: “With appropriate safeguards, LLMs could support more health-informed urban design while ensuring that important decisions remain grounded in professional expertise, community participation, and institutional processes.” The authors also argue that future tools should be judged not just on their design recommendations but on whether they avoid harm, treat disadvantaged areas fairly, support participation and recognise human oversight.

How to Use AI Advice in Urban Design Practice

For councils, consultancies and health bodies, the practical question is how to use these tools without importing their blind spots. The findings suggest a simple rule: let the AI generate options, and make humans own the process.

PracticeWhy the study supports it
Use AI for option generation, not decisionsAdvice avoided harm but was unreliable on process
Add participation and oversight to every promptThe model often omitted them unless reminded
Test prompts with and without budget limitsProcedural safeguards halved under a budget constraint
Compare advice across neighbourhood typesFairness held here, but should be checked for each use
Record AI use in decision papersSupports transparent oversight and public trust
Keep a named professional accountableOutputs should not replace professional judgment

Prompt for the process as well as the plan

If AI advice leaves out residents, ask for them explicitly. A prompt that requires the model to include a community engagement step and to state its uncertainties is likely to produce better output. It is not a substitute for real engagement, but it stops the draft from implying that engagement is optional.

An example of a better urban design prompt

Example prompt that builds in the study’s four criteria

“Suggest built-environment changes to improve physical activity in a mixed-income neighbourhood with a limited budget. For each option, state any risks of harm, explain how it affects lower-income residents compared with higher-income residents, describe how residents should be involved in choosing and designing it, and state your uncertainties and which decisions must be made by qualified professionals and elected officials.”

This kind of prompt asks the model to cover the procedural points it tended to drop. It will not make the output authoritative, and it does not replace a real consultation. But it gives the reviewing urban design professional something to check against, and it makes gaps easier to spot.

Treat budgets as a risk flag

Because the ethical gaps widened under budget pressure, any AI-assisted prioritisation exercise deserves extra review. Ask who loses out under each option, and whether they were consulted. The data and modelling behind those choices can be made more transparent with good data visualization, which helps residents see the trade-offs.

Build governance before scale

Organisations adopting AI in planning should agree rules before use spreads: which tasks AI can draft, who reviews its output, how its use is disclosed, and how residents can challenge decisions. That is part of a sound AI strategy for any public body.

Limits of the Urban Design Study

The study is a useful first step, and the authors present it as one. Its design also sets clear limits on what it can tell us.

One model, one moment

It tested ChatGPT, described in the abstract as “a recent LLM”. Other models, and later versions of the same one, may behave differently. Results from a fixed set of prompts are a snapshot, not a permanent property of AI.

Text advice, not outcomes

The researchers coded what the advice said, not what would happen if it were followed. A recommendation can pass every ethical criterion on paper and still fail in a real neighbourhood, where local context, politics and delivery decide the outcome.

Coded judgements

Ethical criteria require judgement to apply. The team reduced that risk by having two co-authors code every answer independently. Even so, a different team with a different reading of participation or oversight might score some answers differently. What matters most is the pattern: strong on harm and fairness, weak on process, weakest under budget pressure.

Worth repeating across cities

The six pathways and four criteria form a simple, repeatable test. Planning teams could run the same kind of audit on the tools they use, with prompts drawn from their own neighbourhoods, before relying on AI advice for real urban design decisions.

Who Is Behind the Research

Koohsari holds two PhDs, in urban design and in health and sport sciences. He is a visiting researcher at Waseda University and an academic affiliate at the Arnold School of Public Health at the University of South Carolina. According to JAIST, he has written more than 150 peer-reviewed publications with 8,481 citations and is ranked among the world’s top 2% most influential scientists. His research examines how urban spatial structure affects population health in the Asia-Pacific region, using spatial analysis, epidemiological modelling and AI.

JAIST

JAIST was founded in 1990 in Ishikawa prefecture as Japan’s first independent national graduate university with its own campus. About 40% of its alumni are international students. The authors declare no competing interests, and the release lists no funding for this study.

Urban Design and AI Ethics FAQ

What did the study find about AI urban design advice?

ChatGPT’s urban design advice avoided harmful recommendations in all 180 answers and treated lower-income neighbourhoods fairly in 91.7% of evaluable cases, but included community participation in only 60.6% of answers and human oversight in 75.6%.

What happened when the AI had a budget constraint?

Participation fell from 73.3% to 35.0% of answers and oversight from 90.0% to 46.7%, so the procedural safeguards roughly halved.

Should planners use ChatGPT for urban design?

The authors say AI advice can be a useful initial input but should not replace professional judgment or community participation, especially when budgets force prioritisation.

Where was the study published?

In Developments in the Built Environment, volume 27, article 101007, available online from 10 August 2026. It is open access.

Who led the research?

Associate Professor Mohammad Javad Koohsari of JAIST and Professor Koichiro Oka of Waseda University, with co-authors Becky P.Y. Loo, Jing Zhao, Jiuling Li, Ying Long, Yi Lu and Andrew T. Kaczynski.

What are the four ethical criteria?

Non-maleficence (avoiding harm), distributive justice (fair treatment of neighbourhoods), collective participation (involving communities) and transparent oversight (acknowledging uncertainty and human decision-making).

References