AI

How to choose an AI model for construction drawing analysis

Cut to the chase

An AI model can read a floor plan almost perfectly and still miss the cabinets, countertops, or callouts that make the result useful. That is why I would not choose a model for construction drawing analysis from one overall benchmark score.

Coxit’s September 2026 CaseVBench benchmark makes the problem visible. It compares 11 multimodal AI models across about 120 architectural drawing sheets containing about 1,430 hand-labeled objects. The models had to identify cabinets, countertops, floor plans, elevations, and callouts. Each detection was scored using F1 at an intersection-over-union threshold of 0.50.

For a construction workflow, the useful question is not which model sits at the top of the table. It is which model can find the drawing elements the workflow actually depends on, at a cost and speed the job can support.

The hardest drawing element matters more than the overall score

Overall F1 can hide a weak object class. In Coxit’s results, floor-plan scores range from 73% to 99% and elevation scores range from 47% to 94%. Countertops range from 0% to 84%, while callouts range from 0% to 96%. The model ranking changes sharply when the task depends on one of those harder classes.

Selected CaseVBench V2 F1 scores by model and object class, published by Coxit in September 2026
ModelOverall F1CabinetsCountertopsCallouts
GPT-6 Astra92%82%84%96%
Gemini 3.8 Flash70%59%20%68%
Claude Fable 5.170%61%34%64%
GPT-5.6 Sol65%53%24%57%
Qwen3.8-Max62%60%32%47%
GPT-5.6 Terra56%51%23%38%
Gemini 3.5 Flash53%45%16%33%
Gemini 3.1 Pro Preview51%39%10%37%
Claude Opus 540%51%19%15%
Grok 4.630%5%7%0%
Claude Sonnet 521%10%0%4%

GPT-6 Astra leads the published table at 92% overall. The more instructive comparison sits one row lower: Gemini 3.8 Flash and Claude Fable 5.1 both score 70% overall, but Fable scores 34% on countertops against Gemini’s 20%, while Gemini scores 68% on callouts against Fable’s 64%. The same headline score does not describe the same operating behavior.

Cost and processing time change the choice

Median cost and processing time per drawing page published in CaseVBench V2
ModelCost per pageTime per page
Gemini 3.5 Flash$0.02615.5 s
Gemini 3.1 Pro Preview$0.03319.2 s
Qwen3.8-Max$0.055154.9 s
GPT-5.6 Terra$0.05632.0 s
Claude Sonnet 5$0.07260.8 s
Claude Opus 5$0.11237.2 s
GPT-5.6 Sol$0.11445.2 s
Grok 4.6$0.134331.9 s
Gemini 3.8 Flash$0.142100.1 s
Claude Fable 5.1$0.14822.3 s
GPT-6 Astra$0.18628.2 s

At the published median rates, a 500-page pass would cost about $71 with Gemini 3.8 Flash and $74 with Claude Fable 5.1. The serial processing time works out to about 13 hours 54 minutes for Gemini and 3 hours 6 minutes for Fable. That is roughly 4.5 times faster for $3 more across the set, despite the identical overall F1 score.

The cheapest model in the table, Gemini 3.5 Flash, works out to about $13 for 500 pages and about 2 hours 9 minutes of serial processing. GPT-6 Astra works out to about $93 and 3 hours 55 minutes. The $80 difference buys a 39-point increase in overall F1 in this benchmark, plus much stronger countertop and callout scores. Whether that is expensive depends on the cost of the errors and corrections the lower-scoring model leaves behind.

These 500-page figures are calculated from Coxit’s published median per-page values. They assume serial processing with no retries or parallel requests, so they are comparison examples rather than production quotes.

How I would shortlist a model for a drawing workflow

This is the order I would use before putting a model into a construction drawing workflow:

  • Set a minimum score for every object class that can break the workflow.
  • Reject models that miss those thresholds before ranking the overall score.
  • Multiply per-page cost by the real page count and expected revision count.
  • Test latency at the concurrency the product can actually support.
  • Measure human correction time on the company’s own documents.

TL;DR

Coxit’s CaseVBench can narrow the shortlist, but the highest overall F1 score does not automatically make a model suitable for a construction workflow. I would choose against the required drawing elements first, then compare the real job cost, processing time, and correction work on the actual drawing sets.

Source data: Coxit, CaseVBench — AI Object Detection Benchmark for Architectural Drawings, V2 published September 5, 2026.

See more Norml in Google Search

Add Norml as a preferred source
Share
Max Tymoshyn
Written by

Max Tymoshyn

Founder & Architect at Norml Studio

I lead Norml across business, product, design, and technology, setting company direction and shaping the systems behind our work.

LinkedIn ↗