Skip to main content

signal · indexed from World Bank Blogs

Can AI read the rules for greener buildings? We put it to work in 119 cities

Ritula Anand, Jayashree Srinivasan

Data quality: Source-backedSeen at source 3h agoVerified source
Published
08 Oct 2026

The World Bank tested AI to research building energy laws across 119 cities, combining exact legal citations, independent AI review, and human verification.

Can AI read and cite building energy laws across 119 cities? That is the data collection challenge behind the next edition of Building Green, the World Bank Group’s dataset on building energy codes and their enforcement, due in spring 2027. The dataset pairs two types of data: what the law requires, researched from published legal sources, and what happens in practice, reported by local experts. For the legal half, we put AI to the test: OpenAI’s GPT-5.4 Mini drafted answers to 60 legal questions per city, each tied to an exact excerpt from the law; Anthropic’s Claude independently reviewed the legal basis; and analysts checked every answer. In the first 21 economies reviewed, about four in five answers were correct, at roughly US$27 per city in model costs and an estimated 1,500 researcher hours saved. The stakes are high. Buildings and construction account for around 37 percent of global CO2 emissions, and energy codes are one of the main tools governments use to set efficiency standards for new construction. But a code on the books is not a code in practice. That gap is at the heart of Building Green. Its first edition, published in 2024, found that across 88 economies with a mandatory code, just over a third had a functioning system for qualifying and overseeing the inspectors who check compliance. The next edition covers 119 cities in 110 economies. What AI can and cannot answer Data on regulations and their enforcement are collected through two complementary questionnaires. The regulatory data are matters of record — whether a requirement is mandatory, which law establishes it, what penalties apply — and can be traced to official sources such as gazettes, ministry websites, and standards. This questionnaire covers 60 questions per city, answered by AI, with each answer carrying its legal basis, an exact excerpt, and a URL. The enforcement data exist in no official source: how consistently a requirement is applied, how long a compliance review takes, whether the local workforce and market can deliver what the code requires, and whether incentives work in practice. These answers come from the architects, engineers, and officials who work with the process, through a separate expert questionnaire. Figure 1. Who answers what. Two questionnaires on the same building energy code, split by who can answer How the AI research works For each city, the OpenAI-based workflow first identifies the legal sources in force, prioritizing official gazettes, ministry and municipal websites, and standards bodies. News articles, Wikipedia, law-firm summaries, and similar secondary sources are excluded. The core legal texts are then retrieved and indexed. Each question is checked against that source index first and the web second, keeping the results anchored to the wording of the law rather than whichever summary appears first in a search. Where the evidence is not strong enough, “Insufficient evidence found” is an acceptable answer. An automated consistency check flags unsupported answers and contradictions. Figure 2. From legal source to verified answer Testing it, and checking every answer Early runs exposed a recurring problem: the AI often reached a defensible substantive answer but cited the wrong legal instrument. That mattered because the goal was data that could be checked and reproduced, not simply plausible answers. Three lessons became important. First, a parent law is not always the code: the detailed requirements may sit in a ministerial regulation. Second, an old local rule may remain online after being repealed or superseded. Third, the date of the legal instrument matters more than the date of the webpage hosting it. We turned each into an explicit rule for the workflow. Even so, the accuracy review shows where the work still lies: pinning down the latest amendment of a code, and listing every type of project or exemption it covers. Rules reduce errors; they do not eliminate them. Every answer therefore passes three layers of review (Figure 2): an automated consistency check; an independent review of the legal basis by Anthropic's Claude, in supervised sessions with an analyst, so the drafting model never grades its own work; and a team check against the underlying sources. Only answers that pass all three will be published. Behind the headline figure, accuracy reached 75 percent or higher in 18 of the first 21 economies, and up to 95 percent in the strongest. The AI did best where the law gave a clear answer: nearly nine in ten yes/no and fixed-choice answers were correct. The review also produced something reusable: a benchmark of 7,140 human-graded rows, covering 60 questions across 119 cities. Future automated verifiers can be tested against that set rather than assumed to be reliable. Even with full human checking, the time savings are substantial. A researcher previously needed roughly two to three days per jurisdiction to identify the relevant instruments, read them, and draft 60 answers with legal references; reviewing a city's AI-drafted answers now takes about four to six hours. The technology cost was modest: testing and production from March to August 2026 came to about US$3,200 in OpenAI usage across all 119 cities, or roughly US$27 per city — even though a single city run can process five to ten million words of legal text to fill a 60-row spreadsheet (Figure 3). Figure 3. Anatomy of a city run. An illustrative production run shows what lies behind 60 answers: millions of tokens processed, hundreds of model calls, and modest model cost. Production data, August 2026. Back to the buildings When the next edition is published in spring 2027, each city's legal requirements will sit alongside what local experts report about how they work in practice — and the places where the two disagree will be the most telling. A requirement confirmed in the legal research may sit beside an expert response of "Yes, but inconsistently" on enforcement. That gap is often where the most useful policy questions begin: Are there enough qualified inspectors? Can local professionals meet the requirements? Are compliant materials available? What other researchers can take from this AI let us extend the legal research to 119 cities, but the main lesson is about design rather than models. The split came first: questions that could be answered from a published source went to AI; questions that needed professional experience stayed with local experts. Verification did the rest — an exact citation behind every answer, "insufficient evidence" as an accepted response, review by a second model and then by analysts, and every recurring error turned into a new rule. None of this is unique to building energy codes: any dataset built from public records can be collected, and checked, the same way. This blog includes AI-assisted editing. The final content was reviewed by World Bank Group staff.

Provenance

World Bank Blogs(official channel)Data quality: Source-backed
Publisher
World Bank Blogs
Source system
World Bank Blogs — Official API or feed
Verification
Verified — the publisher's own channel
Published
— stated by the source
Last seen at source
Indexed by GLODA
Last updated
Record version
429a3c0a1001
How GLODA knows this — Source-backed · 1 fieldData quality: Source-backed
How each field is known
FieldHow we knowStateConfidence
Publication dateOfficial API or feedData quality: Source-backed100
Source
World Bank Blogs · Official API or feed · official-docs
Retrieval
seen · indexed
Record
429a3c0a1001 · source id 038813e07f05c8d637c2cf91c1aa1539
Computed
provenance-v1

States describe GLODA's confidence in its own fields, never a judgement on the notice or its parties. Methodology

Open this signal in the GLODA terminal

The notice above is the public record. Signing in adds the analysis: saved views and alerts on searches like this one, workflow and pipeline, exports and full workspace tooling.