Text image displaying 'WHITEHOTNOISE AI to CAD Benchmark Review 2026' on a speckled gray background.

This is a very quick and rough analysis to understand how AI benchmarks for CAD are designed, run and scored, and what their results are actually worth in engineering practice.

This is not a thorough academic review, infinite hours have not been spent investigating every benchmark in fine detail, and they are not ranked or really evaluated in any way.

This document is produced to open the conversation about how these benchmarks are built and what they are worth, if they are of any value in really evaluating the ‘AI to CAD’ claims that are rapidly emerging to become White Hot Noise.

We invite submissions to present on the subject at a future CDFAM event.

The full benchmark tables, frontier lab reporting, the papers critiquing these benchmarks, the CD/DC 2026 talks, trends and the PDF are available by email at the end of this summary.

Check the linked primary sources before citing any figure.


As of today, I count 29 dedicated CAD benchmarks from 2024 to September 27, 2026, plus 9 non-academic benchmarks and evaluations found on X, LinkedIn, Reddit, GitHub and vendor sites.

Only 5 dedicated benchmarks have verified peer-reviewed publication; none of the non-academic entries is peer reviewed.

Generated with Claude Opus 5.5. This report was compiled by Claude Opus 5.5 (Anthropic) from web research in September 2026. (yes, I am aware).

Summary of Benchmarks

CategoryCountPeer reviewed (verified)
Dedicated CAD benchmarks29 (27 academic, 2 vendor-run)5: Text2CAD, CADPrompt, CADReview, TriView2CAD, VideoCAD
Adjacent
Drawings, documentation, AEC/BIM, structural engineering, general agents with CAD subsets
121 verified (DesignQA), 1 likely (TechMB)
EDA, listed separately10
Non-academic structured benchmarks40
Non-academic informal evaluations50
  • Releases accelerated sharply in 2026. About a dozen dedicated benchmarks appeared from 2024 to 2025; 16 appeared between April and September 2026, 6 of them in May.
  • Scoring has moved past shape matching to executable tests, editability checks, DFM and FEA, assembly mates, and GUI operation of FreeCAD, AutoCAD, Fusion, SolidWorks, Onshape and Siemens NX but do NOT include design intent or engineering performance metrics.
  • No dedicated meta-analysis of CAD benchmarks was found. Critiques appear in a peer-reviewed survey, in comparison tables inside newer benchmark papers, and in a CD/DC 2026 talk by Metafold; most critics predictably propose their own benchmark as a replacement.
  • Frontier labs now publish CAD scores in model launches. OpenAI reported 95.9% on BenchCAD for GPT-6 Astra on September 3, 2026. The figure is self-reported and not re-graded by the BenchCAD authors.
Peer review status

Five dedicated benchmarks and one adjacent benchmark (DesignQA) have verified peer-reviewed publication. Treat every other score in this report as provisional.

  • Non-academic entries are not peer reviewed. Most are run and scored by the organization that publishes them, and several publishers sell CAD AI products, so…
  • Most academic entries are arXiv preprints, which are also not peer reviewed. Several 2026 entries appear timed for NeurIPS 2026 and may be accepted later.
  • Vendor-reported model scores are self-reported. OpenAI’s and Anthropic’s BenchCAD figures are listed on the BenchCAD leaderboard as vendor self-reported, not re-graded.
Key findings

In the benchmarks, frontier models recover coarse geometry well but fail on precise parameters, advanced operations and engineering behavior. On CADWorld the best computer-use agent reaches 17.5% against 87.0% for a human expert.

  1. CadQuery dominates as the program representation, used by BenchCAD, Text2CAD-Bench, CADExpert and others. FreeCAD Python or GUI is used by Parametric CAD Bench, RealCADBench and CADWorld. CADGenBench is tool-agnostic because it scores STEP output only.
  2. Scoring is shifting from shape to function. Benchmarks from 2024 to early 2026 score IoU and Chamfer distance. Later ones add executable tests (CADTestBench), editability gates (Parametric CAD Bench, HistCAD), design-intent rubrics (MUSE, RealCADBench), FEA and DFM (CADEngBench), and checks on saved native files (CADWorld).
  3. Editing is easier than generation. CADEngBench and CADGenBench both report this. On CADGenBench, top submissions reach about 0.70 to 0.76 on editing and 0.59 to 0.66 on generation.
  4. Harness matters as much as model. gNucleus reports that swapping the agent harness around a fixed model shifts Parametric CAD Bench scores by roughly 10% either way. A community tool that shows the model a render of its output raised one CADGenBench score from 0.360 to 0.457.
  5. Commercial CAD is tested mainly by non-academic benchmarks. CADWorld uses FreeCAD only. AutoCAD-Bench (Markov) and CAD Arena (Normal Research) test AutoCAD, Siemens NX, SolidWorks, Onshape and Fusion, and neither is peer reviewed.
  6. Vendor involvement is growing but uneven. Autodesk Research, Hugging Face, gNucleus, JD Industrial, Nomic, Markov and Normal released benchmarks. No public benchmarks were found from PTC/Onshape, Siemens, Dassault, Adam or Backflip.
Method

A benchmark counts if it has a defined evaluation protocol and a public release, whatever its date.

  • Dedicated CAD benchmark. The primary contribution, or a named evaluation suite, measures AI generation, editing, reconstruction, review, QA or operation of CAD models (B-rep, parametric programs, native CAD files or CAD GUIs). Benchmarks shipped inside methods papers count when they are named and reused.
  • Adjacent. Engineering drawing and documentation understanding, AEC/BIM, or general agent benchmarks that contain CAD subsets.
  • Non-academic, structured. Released by a company, lab or individual outside a paper, with a stated task set, a scoring method and results for several models. Not peer reviewed.
  • Non-academic, informal. A smaller vendor or individual comparison, often a single task or a blog post. Not peer reviewed.
  • Not counted. Foundational datasets used as test sets, methods papers that only report on existing benchmarks, vendor product reviews without a fixed protocol, hardware performance benchmarks, and non-CAD 3D code benchmarks.
Foundational datasets used as test sets (not counted)
  • DeepCAD (2021): 178,238 sketch-and-extrude command sequences. The most common source for text-to-CAD benchmarks, including Text2CAD, CADPrompt and CADFusion.
  • Fusion 360 Gallery (2021): human design sequences from Autodesk Fusion.
  • ABC (2019): over 1,000,000 B-rep models with design history discarded.
  • SketchGraphs (2020), MCB (2020), CC3D: sketch, mechanical-component and scan-based sets used as further test sources.

Get the full report

Enter your details to see the full benchmark tables (29 dedicated, 12 adjacent, 9 non-academic), frontier lab reporting, the papers critiquing these benchmarks, the two CD/DC 2026 talks, trends and caveats, and to download the PDF.


Have an Opinion and Data to back it up?

If you have access to deeper research into this area, have compiled (or able to compile the data) and are interested in discussing with others interested in the results, I would love to hear from you, and potentially have you present at an upcoming CDFAM event.

You can submit a presentation proposal using the online portal, or contact me and we can discuss your findings, and whether it might be a good fit.

If you are one of the researchers who developed the benchmark and want to discuss further, please reach out.


Recent Interviews & Articles