GPT-6 Astra dominates construction drawing analysis in Coxit’s original benchmark
By Max Tymoshyn3 min read
Cut to the chase
Coxit’s latest original research has a clear winner: GPT-6 Astra scored 92% overall F1 for object detection in architectural drawings, 22 points ahead of the next-best models. The benchmark also exposes the trap in choosing from one leaderboard number: a model can read a floor plan almost perfectly and still miss the cabinets, countertops, or callouts that make the result useful.
Coxit’s September 2026 CaseVBench benchmark makes the problem visible. It compares 11 multimodal AI models across about 120 architectural drawing sheets containing about 1,430 hand-labeled objects. The models had to identify cabinets, countertops, floor plans, elevations, and callouts. Each detection was scored using F1 at an intersection-over-union threshold of 0.50.
For a construction workflow, the useful question is not which model sits at the top of the table. It is which model can find the drawing elements the workflow actually depends on, at a cost and speed the job can support.
GPT-6 Astra is 22 points ahead in Coxit’s original benchmark
Overall F1 can hide a weak object class. In Coxit’s results, floor-plan scores range from 73% to 99% and elevation scores range from 47% to 94%. Countertops range from 0% to 84%, while callouts range from 0% to 96%. The model ranking changes sharply when the task depends on one of those harder classes.
| Model | Overall F1 | Cabinets | Countertops | Callouts |
|---|---|---|---|---|
| GPT-6 Astra | 92% | 82% | 84% | 96% |
| Gemini 3.8 Flash | 70% | 59% | 20% | 68% |
| Claude Fable 5.1 | 70% | 61% | 34% | 64% |
| GPT-5.6 Sol | 65% | 53% | 24% | 57% |
| Qwen3.8-Max | 62% | 60% | 32% | 47% |
| GPT-5.6 Terra | 56% | 51% | 23% | 38% |
| Gemini 3.5 Flash | 53% | 45% | 16% | 33% |
| Gemini 3.1 Pro Preview | 51% | 39% | 10% | 37% |
| Claude Opus 5 | 40% | 51% | 19% | 15% |
| Grok 4.6 | 30% | 5% | 7% | 0% |
| Claude Sonnet 5 | 21% | 10% | 0% | 4% |
GPT-6 Astra leads the published table at 92% overall. The more instructive comparison sits one row lower: Gemini 3.8 Flash and Claude Fable 5.1 both score 70% overall, but Fable scores 34% on countertops against Gemini’s 20%, while Gemini scores 68% on callouts against Fable’s 64%. The same headline score does not describe the same operating behavior.
Coxit’s 70% tie breaks on cost and speed
| Model | Cost per page | Time per page |
|---|---|---|
| Gemini 3.5 Flash | $0.026 | 15.5 s |
| Gemini 3.1 Pro Preview | $0.033 | 19.2 s |
| Qwen3.8-Max | $0.055 | 154.9 s |
| GPT-5.6 Terra | $0.056 | 32.0 s |
| Claude Sonnet 5 | $0.072 | 60.8 s |
| Claude Opus 5 | $0.112 | 37.2 s |
| GPT-5.6 Sol | $0.114 | 45.2 s |
| Grok 4.6 | $0.134 | 331.9 s |
| Gemini 3.8 Flash | $0.142 | 100.1 s |
| Claude Fable 5.1 | $0.148 | 22.3 s |
| GPT-6 Astra | $0.186 | 28.2 s |
At the published median rates, a 500-page pass would cost about $71 with Gemini 3.8 Flash and $74 with Claude Fable 5.1. The serial processing time works out to about 13 hours 54 minutes for Gemini and 3 hours 6 minutes for Fable. That is roughly 4.5 times faster for $3 more across the set, despite the identical overall F1 score.
The cheapest model in the table, Gemini 3.5 Flash, works out to about $13 for 500 pages and about 2 hours 9 minutes of serial processing. GPT-6 Astra works out to about $93 and 3 hours 55 minutes. The $80 difference buys a 39-point increase in overall F1 in this benchmark, plus much stronger countertop and callout scores. Whether that is expensive depends on the cost of the errors and corrections the lower-scoring model leaves behind.
These 500-page figures are calculated from Coxit’s published median per-page values. They assume serial processing with no retries or parallel requests, so they are comparison examples rather than production quotes.
How I would shortlist a model for a drawing workflow
This is the order I would use before putting a model into a construction drawing workflow:
- Set a minimum score for every object class that can break the workflow.
- Reject models that miss those thresholds before ranking the overall score.
- Multiply per-page cost by the real page count and expected revision count.
- Test latency at the concurrency the product can actually support.
- Measure human correction time on the company’s own documents.
TL;DR
Coxit’s CaseVBench can narrow the shortlist, but the highest overall F1 score does not automatically make a model suitable for a construction workflow. I would choose against the required drawing elements first, then compare the real job cost, processing time, and correction work on the actual drawing sets.
Original research and source data: Coxit, CaseVBench — AI Object Detection Benchmark for Architectural Drawings, V2 published September 5, 2026.
See more Norml in Google Search
Add Norml as a preferred source