AI

What Coxit’s CaseVBench says about choosing an AI model

Norml’s take: the most useful part of Coxit’s new benchmark is not the winner. It is the proof that model selection has to be tied to a specific task, model version, and operating cost. A general claim that one model is “best” is not useful enough for a production decision.

Coxit’s CaseVBench architectural drawing benchmark compares 11 multimodal AI models on object detection in architectural drawings. Coxit is a Norml Studio client. The benchmark, methodology, and results are Coxit’s. This article is a short Norml overview of the published findings, not an independent validation.

What Coxit measured

The September 2026 V2 release covers about 120 drawing sheets and about 1,430 hand-labeled objects across five classes: windows, doors, kitchen countertops, plumbing fixtures, and callout boxes.

Coxit reports F1 scores at an intersection-over-union threshold of 0.50. The benchmark also places estimated processing time and cost per page beside the accuracy results. That makes the comparison more useful than a leaderboard built around one aggregate score.

What changed between releases

The first CaseVBench release reported that no tested model exceeded 33% F1 on countertops or 50% on callout boxes. In V2, Coxit reports GPT-6 Astra at 84% for countertops and 96% for callout boxes.

That shift is a reminder that an AI benchmark is a dated snapshot. Model names, versions, test conditions, and release dates belong beside the result. Without them, a correct finding can quickly turn into a misleading general claim.

What the results mean for an AI project

The practical lesson is to test the workflow, not the provider’s reputation. A model can score well overall and still fail on the object class that matters most to the product. Class-level results expose that weakness before it becomes an integration problem.

Accuracy is also only one part of the decision. Processing time, per-page cost, consistency, and the amount of human correction required can change which model is sensible in production. CaseVBench publishes several of those tradeoffs together; a product team should add its own documents and review process before choosing a model.

CaseVBench does not decide whether a model is ready for another company’s workflow. It does show a stronger way to narrow the options: define the task, keep the measurement rule consistent, inspect performance by class, and compare the score with the operating cost.

Full source: Coxit, CaseVBench — AI Object Detection Benchmark for Architectural Drawings.

See more Norml in Google Search

Add Norml as a preferred source
Share
Max Tymoshyn
Written by

Max Tymoshyn

Founder & Architect at Norml Studio

I lead Norml across business, product, design, and technology, setting company direction and shaping the systems behind our work.

LinkedIn ↗