LLM · BIM

Automated BIM checks against Indonesian building rules

An LLM agent reads the regulations in Bahasa Indonesia and picks the checks a building model needs. A rule engine does the measuring and decides what passes.

A building information model (BIM) describes every stair, door and wall as an object with dimensions and properties. That makes it tempting to check the design against the building code automatically, before anything is built. Is the riser too tall? Is the door wide enough?

I started investigating this at Nucon. The idea was sound, but the rules were the hard part. Building codes are written as prose, and each requirement had to be found, interpreted and turned into code by hand. For Indonesia that means working from regulations in Bahasa Indonesia.

That has become much, much easier to build. Below is a proof of concept that checks a building model against three Indonesian regulations. The regulations stay in Indonesian: the agent searches them in Indonesian and cites them verbatim.

Pick a viewpoint and press Check this view. Each check is a recording of a real run, played back in about eight seconds. When the report arrives, the failing components turn red. Select a finding to see the measurement drawn on the model, and its source link to read the regulation. Switch the rulebook to compare the national rules with Jakarta’s.

Fig. 01 · A compliance check on a two-storey duplex, replayed from a recorded agent run
Viewpoint
Rulebook

Drag to orbit, scroll to zoom, click a component for its properties. Pick a viewpoint to run a check.

Loading the model…

The model

The building is the Duplex, a two-storey, two-unit residential sample that buildingSMART distributes for testing IFC software. It was exported from Revit in 2011. Its 286 components with geometry each carry property sets: a door knows its width, a stair flight its riser height and tread length.

The three viewpoints are where the recordings were made: the stair in unit A, the Level 1 bathroom door next to it, and a bathroom door upstairs. From any other position you can still orbit the model and click components to read their properties, but only these three views have recorded checks.

The agent picks the checks, the rules decide

A language model is good at deciding what is relevant. It is not a reliable judge of whether 194 is more than 180. So the work is split.

The viewer sends the components in view, a screenshot and the camera position. The backend adds the building around them: the flights and railings of a stair, the wall a door sits in, the rooms on either side. The agent, an OpenAI model with a set of tools, then works through the scene:

  1. It lists the elements and decides which topics apply. The stair view covers stairs, headroom, handrails, doors and corridors.
  2. It searches the regulations for each topic.
  3. It lists the executable rules for the active rulebook and runs the ones that apply.
  4. It writes a report.

Only step 3 produces a verdict. Each rule is a small YAML file: the element types it applies to, a measurement function, and a constraint such as riser_height_max <= 180 mm, with the citation it comes from. The measurement is computed from the IFC model with IfcOpenShell, and the comparison is ordinary code.

The report is constrained too. A finding has to reference a rule result that was actually executed; a failure the agent leaves out is appended anyway; and the numbers in the report are copied from the rule result, never from the model’s text. The agent decides where to look and explains what it saw. It cannot mark its own homework.

Reading the rules in Bahasa Indonesia

The regulation corpus is 483 passages from three documents:

  • Permen PUPR 14/2017, the national ministerial regulation on building accessibility and ease of use.
  • SNI 03-1746-2000, the national standard for means of escape.
  • Pergub DKI Jakarta 72/2021, the Jakarta governor’s regulation on means of egress.

Each passage keeps its original Indonesian text, with a short English gloss for readers like me. Search combines vector embeddings and keyword matching over the Indonesian text. The agent’s queries are in Indonesian: tinggi anak tangga (riser height), ruang bebas tangga (stair headroom), lebar pintu (door width). The source links in the figure show the article it relied on, in the original wording.

The two rulebooks are layers. The national rulebook applies Permen PUPR and the SNI. The Jakarta rulebook applies Pergub 72/2021 first and falls back to the national rules where Jakarta says nothing, so the same stair gets a 180 mm riser limit in one rulebook and 178 mm in the other.

What it found

The Duplex fails on its stair and its doors:

  • The risers are 194 mm. The national limit is 180 mm; Jakarta’s is 178 mm.
  • The treads are 250 mm deep. The national minimum is 300 mm; Jakarta’s is 280 mm.
  • Four doors have 762 mm openings, below the 800 mm minimum in both rulebooks.

Headroom, the national stair width and the corridor widths pass. The Jakarta rulebook adds one more failure: its fire-stair rules require a 1,200 mm stair, and this one is 914 mm. That finding needs a person to decide if it applies at all. Pergub 72/2021 is about means of egress in buildings, and whether a small house’s internal stair counts as a fire stair is a question for someone who knows how the regulation is applied, not for a rule file.

These are not surprising results for a US sample house checked against Indonesian rules. The point is that each one comes with a measurement, a drawn dimension on the model and the article it breaks.

Where the data fights back

The model’s own data can’t be trusted blindly. In the Duplex, the stair flight attributes are in feet even though the file declares metres, a bug in the 2011 Revit exporter. A checker that read those attributes directly would report 194 mm risers as 636 mm.

So every stair measurement is taken twice. The property set gives one value, and the geometry gives another: the height of each tread, found from the mesh. If they agree, the result has high confidence. If they disagree, the geometry wins, the confidence drops and the note says why. A result below the confidence threshold becomes review, not fail. The source panel in the figure lists these notes under “How it was measured”.

Cost and limits

Each recorded run took 54 to 71 seconds and cost between 31 and 52 US cents in model usage. Most of that is input: the scene, the regulation passages and the screenshot come to 60,000 to 100,000 tokens per check. The recordings mean this page costs nothing to view.

This is a proof of concept, not a certified checker. It covers stairs, doors, corridors and headroom, not structure or services, and it checks the part of the building in view rather than the whole model. The regulation values were extracted from the primary documents and still need checking by someone qualified to read them. But the pieces that used to take the longest (finding the requirement, reading it in Indonesian, and connecting it to a measurable property of a model) are now quick to build, and the verdicts still come from code that can be tested.