What Coxit’s CaseVBench says about choosing an AI model
By Max Tymoshyn2 min read
Norml’s take: the most useful part of Coxit’s new benchmark is not the winner. It is the proof that model selection has to be tied to a specific task, model version, and operating cost. A general claim that one model is “best” is not useful enough for a production decision.
Coxit’s CaseVBench architectural drawing benchmark compares 11 multimodal AI models on object detection in architectural drawings. Coxit is a Norml Studio client. The benchmark, methodology, and results are Coxit’s. This article is a short Norml overview of the published findings, not an independent validation.
What Coxit measured
The September 2026 V2 release covers about 120 drawing sheets and about 1,430 hand-labeled objects across five classes: windows, doors, kitchen countertops, plumbing fixtures, and callout boxes.
Coxit reports F1 scores at an intersection-over-union threshold of 0.50. The benchmark also places estimated processing time and cost per page beside the accuracy results. That makes the comparison more useful than a leaderboard built around one aggregate score.
What changed between releases
The first CaseVBench release reported that no tested model exceeded 33% F1 on countertops or 50% on callout boxes. In V2, Coxit reports GPT-6 Astra at 84% for countertops and 96% for callout boxes.
That shift is a reminder that an AI benchmark is a dated snapshot. Model names, versions, test conditions, and release dates belong beside the result. Without them, a correct finding can quickly turn into a misleading general claim.
What the results mean for an AI project
The practical lesson is to test the workflow, not the provider’s reputation. A model can score well overall and still fail on the object class that matters most to the product. Class-level results expose that weakness before it becomes an integration problem.
Accuracy is also only one part of the decision. Processing time, per-page cost, consistency, and the amount of human correction required can change which model is sensible in production. CaseVBench publishes several of those tradeoffs together; a product team should add its own documents and review process before choosing a model.
CaseVBench does not decide whether a model is ready for another company’s workflow. It does show a stronger way to narrow the options: define the task, keep the measurement rule consistent, inspect performance by class, and compare the score with the operating cost.
Full source: Coxit, CaseVBench — AI Object Detection Benchmark for Architectural Drawings.
See more Norml in Google Search
Add Norml as a preferred source