KAUST · Miami University

Architect-Ant

Editable automatic furnishing of architectural floor plans

Fedor Rodionov1 Aleksandar Cvejić1 Michael Birsak1 John Femiani2 Peter Wonka1
1King Abdullah University of Science and Technology (KAUST) 2Miami University
Paper Code soon AntPlan data 3D comparison
Architect-Ant turns an empty structured floor plan into an editable furniture layout and a furnished 3D scene.
Architect-Ant turns an empty structured floor plan into a furnished 3D scene through an editable, object-level layout.

A fine-tuned vision-language model writes furniture layouts for real floor plans as short, editable code. It learns professional furnishing patterns from AntPlan, a collection of 505 real architectural drawings, and is then refined with reinforcement learning against a room-aware Layout Rule Score. It generates layouts directly, without an iterative agentic critique loop at inference.

505
professional floor plans in AntPlan, 10 room categories
92
annotated furniture classes, 53 objects per plan on average
94%
functional completeness (FUN+), best of eight methods
8.80
mean Layout Rule Score on 109 shared room shells, best overall
95s
estimated per room, with 20 candidates and 3D scene construction (SceneSmith: ~1,555 s)
01 — Interactive comparison

Same room, eight methods

Six SceneSmith room shells for each of four room types, rebuilt in 3D from the benchmark layouts with the paper's assets. Every method furnishes the same walls, doors and windows. Pick a room from a room-type menu and any two methods, then orbit and look from any angle. The two cameras stay linked.

Click a method to load it into the highlighted panel. ← → flip through methods. Walls facing the camera fade automatically. ★ ours

This room, per method

What Architect-Ant wrote for this room


        
02 — Method

Three sources of layout knowledge, combined in training

Good furniture placement draws on examples, explicit constraints and broad world knowledge. Each source is incomplete on its own. Architect-Ant combines all three while it is trained, so that at inference a single model generates constraint-aware layouts directly.

01

Professional layouts

AntPlan provides supervision from real architectural drawings: object choices and arrangements that professionals actually use, including fixtures and service elements.

02

Explicit constraints

The Layout Rule Score turns geometric and functional requirements into a deterministic reward, including containment, collisions, clearance, reachability and room-specific relations.

03

Pretrained knowledge

Gemma-4-31B contributes broad semantic and visual understanding. A LoRA adapter per room type adapts it to floor plans and the layout language.

2.1AntPlan: real professional floor plans

AntPlan contains 505 residential floor-plan drawings with structural, room and furniture annotations across ten room categories and 92 object classes, with 53 objects per plan on average. Walls, doors, windows and railings are extracted with an RT-DETR-X detector trained on CubiCasa5K, and room labels are assigned manually. A furniture detector bootstrapped from a hand-labeled subset proposes objects, which are then reviewed and corrected by hand. The experiments use cleaned bedroom, bathroom, kitchen and living-room subsets with 2,259 training and 243 validation rooms. Rooms from one source plan never appear in both splits. The dataset card is on Hugging Face.

Examples of original architectural floor-plan images from AntPlan.
Original architectural floor-plan images from AntPlan. Each is paired with structural, room and furniture annotations.
An annotated AntPlan example: original drawing, structural and object boxes, and room regions.
One annotated plan. Left: the original drawing. Center: structural and object boxes with semantic labels. Right: annotated room regions.
Dataset# plansStyleRoom cls.Furniture cls.Avg. objects
MSD5,372Synthetic, simple boxes9——
ResPlan17,000Synthetic, simple boxes7——
ZInD2,737Synthetic, from 3D reconstructionfree text——
FloorPlanCAD2,845CAD drawings, tiled—3016.73
SESYD10Synthetic—1217.07
CubiCasa5K5,000Real, professional architectural101027.20
AntPlan505Real, professional architectural109253.01

FloorPlanCAD is distributed as 15,663 tiles cut from 2,845 drawings; SESYD has 10 layouts with 100 variations each; CubiCasa5K furniture classes are coarse.

2.2Layouts as editable code

A layout is a short sequence of objects, each with a class, the top-left corner of its footprint and its extents, in metres in the room's own frame. The format is compact enough to generate directly and simple to edit by hand. It lets containment, overlap and clearance be checked exactly.

Height and orientation are not predicted. When the layout becomes a 3D scene, each footprint is matched to an asset of a compatible class and size. Its height comes from the asset, and its orientation is inferred from walls and neighbouring furniture: chairs face tables, televisions face seating.

FURNITURE
OBJ class=bed x=0.30 y=0.50 w=1.60 h=2.00
OBJ class=nightstand x=1.90 y=0.50 w=0.50 h=0.50
END

Syntax example. The real output for each room above is shown under the viewer.

2.3Supervised adaptation, then GRPO

Gemma-4-31B-it is fine-tuned with LoRA on the language model, with the vision encoder frozen. The inputs are the room geometry as text, a structure-only raster of the room and a furniture request. The target is the final layout alone, with no reasoning traces. Reinforcement learning then starts from this checkpoint. Each rollout contains a free-form reasoning block followed by a layout, and the reward is computed only from the parsed final layout. The model therefore develops its own strategy for satisfying spatial requirements, following the outcome-supervision principle of DeepSeek-R1-Zero.

Data preparation and training pipeline: detection, room extraction, SFT and GRPO on Gemma 4 31B.
Data preparation and training. Detector-assisted annotations form room-level examples for supervised fine-tuning, followed by GRPO with final-layout rewards. Each room category gets its own LoRA adapter.
Rollouts
16 responses per room, 8 rooms per update (128 on-policy rollouts), one policy-gradient step per batch
Advantage
reward centred by the room-group mean, scaled by the pooled batch standard deviation, clipped to [−2, 2]
Dr. GRPO
no per-group variance normalisation; loss normalised by a fixed 1,024-token constant
Schedule
15 updates, learning rate 2×10⁻⁵, KL coefficient 0.001 to the frozen base model, no learned critic

2.4Layout Rule Score

LRS is a deterministic, room-aware score. The evaluator parses the layout and starts it at 10. It then deducts a penalty for every violated rule instance. Thirty-three rule types cover shared structural checks: containment, wall overlap and penetration, door and window obstruction, disallowed furniture overlap and accessibility. They also cover room-specific relations such as appliance mounting, kitchen runs and work zones, focal seating, window blocking, wall alignment and fixture placement.

Thresholds and exemptions are calibrated against AntPlan annotations, such as a rug beneath a bed or a sink set into a cabinet. A typical room needs 50–200 predicate evaluations. A score of 10 means no encoded rule was violated. It is not a guarantee of an ideal layout.

S(𝓛; x) = 10 − Σk pk(𝓛, x)

During GRPO, LRS is the main outcome reward, with auxiliary terms for object-size deviation, gated similarity to the annotated layout and response validity. These terms are used only in training. At inference, the same rule scorer screens and selects among candidates.

2.5Inference and 3D scene construction

For a new room shell, Architect-Ant retrieves the eight most similar same-type training rooms by size and samples four furniture requests from them, each specifying categories, counts and approximate sizes. It generates five layouts per request, 20 candidates in total, and the rule scorer selects one of them. The selected layout remains editable and is converted into a furnished 3D scene.

Inference pipeline: structural input, fine-tuned Gemma, DSL output, rule-based scorer, semantic mask and 3D scene.
Inference. The room geometry, a structure-only image and a retrieved furniture request condition the room adapter, which produces a reasoning trace and an editable layout. Layouts are scored and selected before 3D assets are retrieved and rendered.
03 — Results

Fuller rooms, fewer violations

The benchmark uses 109 released SceneSmith room shells: 31 bedrooms, 32 living rooms, 27 dining rooms, 13 kitchens and 6 bathrooms. Architect-Ant, Holodeck, LayoutGPT and LayoutVLM share the same Gemma-4-31B backbone. SceneSmith contributes its released GPT-5.2 scenes, refined by its own agentic loop.

Scene statisticsMetrics and runtime
MethodBase model ScenesObj./roomObj./10 m²Cov. % COL+ % ↓OOB % ↓Acc. pairs % ↑FUN+ % ↑LRS ↑Time s ↓
ATISSTransformer906.53.1029.524.118.195.2841.5020.2
InstructSceneGraph diff.907.23.1526.313.720.986.963−2.8221.6
DiffuSceneDiffusion908.03.4726.518.120.882.771−3.5574.3
LayoutGPTGemma-41097.84.0021.98.68.597.4842.1067.7
LayoutVLMGemma-41098.24.2228.112.716.094.0866.29272.0
HolodeckGemma-41099.14.5934.70.00.099.5787.84432
SceneSmithGPT-5.21098.94.5023.43.41.399.4918.441,555
Architect-AntGemma-410910.85.4529.10.60.697.2948.8095.4

Room-level means. Bold marks the best result and underline the second best among the geometric and functional metrics. Counts, coverage and runtimes are not ranked. COL+ is the fraction of objects in a disallowed collision after category-aware filtering, OOB the fraction outside the room, and FUN+ the share of required room functions served by contained, reachable objects. SceneSmith's time is an object-count-normalised estimate from its full pipeline, not a measured rerun. Timing scopes and hardware differ between methods.

3.1Blind visual preference

Two vision-language judges compare anonymous Architect-Ant and SceneSmith renders of the same shell, with randomised A/B order. Both prefer Architect-Ant overall. Their preference is strongest on kitchen and dining shells, while SceneSmith wins four of the six bathrooms.

Architect-Ant preferredTieSceneSmith preferred

Judges' votes for the room shown above

3.2What SFT and GRPO each contribute

On the 243 held-out AntPlan rooms, each model generates eight samples per room. Only parseable, non-empty layouts are retained and scored with raw LRS. Supervised fine-tuning gives the largest jump. Reasoning GRPO on top improves both the mean and the best-of-eight score in every room type.

Mean LRSBest of 8
RoomBaseSFT+GRPOBaseSFT+GRPO
Kitchen−11.71−0.753.48−2.375.735.80
Bathroom−0.536.257.004.008.509.00
Living room2.295.416.226.548.778.95
Bedroom−1.515.716.023.898.488.79

Base and SFT generate the layout directly. The +GRPO column adds reasoning-enabled GRPO after SFT. Confidence intervals are reported in the paper.

3.3Knowledge sources in the compared configurations

SourceATISSInstructSceneDiffuSceneLayoutGPTLayoutVLMHolodeckSceneSmithArchitect-Ant
Layout examples
Spatial constraints
Pretrained model knowledge

● used, ○ not used, in the evaluated configuration. LayoutGPT is evaluated zero-shot, without retrieved layouts.

04 — Whole houses

From rooms to complete apartments

Architect-Ant furnishes every room of a floor plan with its room adapters. In a whole-house study, 17 selected Architect-Ant houses are each paired at random with 17 Holodeck and 17 SceneSmith houses. Judges compare anonymous top-down renders for living, circulation, cooking and eating, bathroom and sleeping functions. The three pairs below come from the paper, including one where SceneSmith was preferred.

We selected 17 residential houses from the SceneSmith set and 17 Holodeck houses, and paired each of them with one of 17 Architect-Ant houses. The same Architect-Ant houses are used against both baselines.

Judgevs Holodeckvs SceneSmithTotal
Claude Sonnet 516/17 94.1%14/17 82.4%30/34 88.2%
Kimi K317/17 100%15/17 88.2%32/34 94.1%

Architect-Ant wins out of all judgments.

05 — Citation

BibTeX

@article{rodionov2026architectant,
  title   = {Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans},
  author  = {Rodionov, Fedor and Cveji{\'c}, Aleksandar and Birsak, Michael and
             Femiani, John and Wonka, Peter},
  journal = {arXiv preprint arXiv:2606.10953},
  year    = {2026}
}