CDFAM Amsterdam 2025 · Amsterdam · 9–10 July 2025

AI Judges In Design: Statistical Perspectives On Achieving Human Expert Equivalence With VLMs

Abstract

The subjective evaluation of early stage engineering designs, such as conceptual sketches, traditionally relies on human experts. However, expert evaluations are time-consuming, expensive, and sometimes inconsistent. Recent advances in vision-language models (VLMs) offer the potential to automate design assessments, but it is crucial to ensure that these AI “judges” perform on par with human experts. However, no existing framework assesses expert equivalence.

Transcript

From YouTube’s automatic captions, lightly cleaned; expect some errors. Each timestamp opens the video at that moment.

Read the full transcript · 2,615 words

Fantastic. Hi everyone. Good afternoon. It’s a pleasure to be here. My name is Kristen Edwards and today I’ll be talking about research that we recently did within our group called AI judges in design, a statistical perspective on AI in design evaluation. And I did this work in collaboration with my adviser, Professor Faz Ahmed, along with professors Scarlett Miller and Fernas Terranchi. So, a little bit on my background.

0:26 I have my bachelor’s and my master’s degrees in mechanical engineering. And I’m currently a PhD candidate at MIT in the decode lab where my research is around AI in engineering design and manufacturing. So, I’d like to start with a question. Students were asked to design the most innovative milk frother that they could and created these. Which do you think is the most creative design? You might take a moment to look at these and compare the four and decide that this countertop jet engine is the most creative.

0:59 And it’s easy enough when you’re only judging four against each other. But how about when you’re dealing with tens of designs or hundreds or thousands? And this question is becoming more and more relevant because of the first of three things that are happening at once. So the first is that there has been this increased generation of content and design. And this has been going on for about the past two decades and we can see it across many fields.

1:24 The figure on the top is actually journal publications and you can see an explosion over the last two decades. But it’s also been bolstered by generative AI and this is kind of shown in a report from Harvard Business Review which showed that in 2024 the number one use case of generative AI was concept generation. So at the same time we’re seeing improvement in large language models or LLMs and their multimodal counterparts VLMs vision language models and this can be seen as they’re approaching human level on a number of benchmarks and baselines.

1:59 The validity of those benchmarks is constantly being questioned and moved as we need to choose things that are more representative of intelligence. But nonetheless, as training sizes, parameter sizes grow and as algorithms become better, we see improvements in these models. And kind of as a result of these two things, there has been the introduction of LLM as a judge. So if you haven’t heard of this, let me formally define it.

2:25 It is a framework by which a large language model evaluates an input X within a context C. And the evaluation can be a score of the input, a label or a selection among many inputs. And there have been a number of really interesting academic articles published in the past two years around LLM as a judge. Some around defining it, some about training better LLMs as a judge, others around understanding their biases and how they compare to humans.

2:53 But outside of academia, LLMs have been used as a judge in a lot of practical cases already. So, we’ve seen this in LLMs being used to review resumeumés or perform idea selection. And there are also some troubling cases. This recently received some uproar when people found that some scientists were putting terms in their papers in white font to try and trick any LLMs that were acting as peer reviewers.

3:20 And it brings up this general question of how should we ethically use LLMs as a judge and how can we guarantee that they’re performing in the ways that we hope. And in general, I propose that if you’re using an AI judge in any scenario, there should be some amount of assessment of how well it matches some ground truth or some baseline rating within within your field. And for me, this brings up the question, what are the implications of AI judges in design evaluation?

3:53 But first, let’s be clear, AI for evaluating designs is not new. There’s already a rich field of research around design creativity and design evaluation. In fact, Amobu in 1982 produced the paper that introduced the consensual assessment technique, which has been known as a gold standard in design evaluation for decades now. And there’s other work. Shaw, Vargas, Anernandez, and Smith released a survey based technique for evaluating designs.

4:22 And as the expert raider techniques like CAT became a gold standard, we see two main things happen. They are great because experts consider multiple factors and have high accuracy when evaluating designs, but they’re difficult because using an expert means that it’s slow and resource demanding. So in response to this there has been a lot of research around using AI for design evaluation trying to mimic those evaluators and so we saw symbolic and knowledge based AIS a couple of decades ago and still today and then more recently we’ve seen deep learning and neural netbased AI attempting to perform design evaluation and actually that’s my master’s thesis so I’m in the crowd.

5:09 So if AI for evaluating designs isn’t new, what is? Well, utilizing these large pre-trained models instead of bespoke models that you train yourself on a data set is new. And while there are powerful scaling implications to using pre-trained models because you don’t necessarily have to have a data set of labeled designs, using these requires that we properly evaluate whether an AI judge adheres to some ground truth.

5:42 And we can turn to literature that’s already been done around rating to get a jump start. So first of all, what metrics are people using in LLM as a judge research? The most common is agreement. And agreement can be pure agreement, which is just if you have a number of items that have been rated, the number where two raiders agree divided by the total number. Or it can be agreement weighted by chance or Cohen’s Kappa.

6:06 And that’s only good for ordinal or categorical data. If we turn then to group rating or crowd research, we see metrics usually around interrator reliability, like intraclass correlation and coins again. But these aren’t the only metrics that you can use for evaluating whether or not two judges or two raiders agree. And in fact, when my colleagues and I were researching what metrics are most appropriate for determining if an AI judge is agreeing with a human raider, we found all of these.

6:44 And generally these metrics can be grouped into these six categories. Agreement, interrator reliability, correlation or ranking, errorometrics, equivalence testing, and distributional similarity. And what I’m proposing today is that if you’re going to use an AI judge within design evaluation, you should assess how well that AI judge matches some ground truth, perhaps a human expert, along these six dimensions. And I’d like to first start by saying why any one of these alone isn’t sufficient.

7:19 So for example, if you had 10 ideas shown here, each mapping to one of the items and it was rated by multiple people. So raider one shown here in blue, raider two in orange, and raider three in green. We can look at different metrics to understand how well these raiders agree. First of all, it’s important to note that the agreement is zero for all three of these.

7:40 And this is the most common metric used in LLM as a judge research today. The reason it’s zero is that none of the raiders gave the exact same rating for any of the items. But this fails to capture that there is some trend among the ratings. So we might use a different metric like Spearman’s rank row which is a measure of relative ranking. And if we use blue raider one as the ground truth, we can see that both two and three actually have a p a perfect spearman’s rank of one with raider one.

8:12 And this indicates that all three of the raiders gave the same relative rank to the three different items. Sorry, to the 10 different items. But once again, Spearman’s rank alone doesn’t capture that there are vast differences in how all three raiders scored the designs. And so you might turn to mean absolute error which no longer looks at relative ranking but instead at absolute scores. And here we can see that raider 2 in orange agreed with raider one better than raider 3 agreed with raider one indicated by a lower mean absolute error.

8:48 But this becomes more interesting if we introduce a raider 4 here shown in red. And this raider provided the exact opposite linear ranking of the designs. And so that’s seen in the Spearman’s rank with a score of negative one. But it actually has the lowest mean absolute error of any of them. So if you were to use just one of these metrics, you would miss out on the full picture.

9:11 And none of these metrics, agreement, Spearmman’s rake or mean absolute error capture the fact that there is a huge difference in distribution here with Raider 2 having sort of a biodal approach whereas the rest of them have a linear more unimodal ranking. So this is to say we need a suite of statistical metrics to properly assess how well an AI judge matches a human raider. And among the six that I proposed, each category of metric measures a different aspect of raider similarity.

9:48 So for example, agreement measures whether ratings are close or are the same or are numerically close. Interrator reliability measures the consistency of ratings, accounting for chance. Correlation and ranking measures whether raiders agree on the relative ordering of different items. Error metrics is the magnitude of disagreement. Equivalence testing is whether differences in ratings are small enough to be considered practically insignificant or negligible. And finally, distributional similarity is about measuring whether rating distributions are similar in their shape or location.

10:25 And you might have noticed there are multiple different metrics and tests that can be used within each of these categories. And the way that you choose these is based on the features of your data set. So whether you’re dealing with categorical, ordinal, continuous, or discrete. And so as part of our research, we are outputting a flowchart that people can use in order to determine which of these metrics is most suitable for your specific design task.

10:56 So I’ve proposed the statistical suite and we wanted to display the utility of it by running a case study. So we performed a case study in which two human experts, three trained noviceses and four AI judges rated a data set of around 10,000 designs and this data set had been produced by Scarlett Miller and Christine Toe at Penn State. And the big question here was can AI judges learn to rate like an expert?

11:21 And so the problem that each of the raiders were given was given a design sketch provide a score for a specific design metric and the design metrics of interest were creativity, uniqueness, usefulness and drawing quality. And then the next big question for us as an experimental decision was what should the baseline be? And we determined that rather than choose random cut offs of where people should score on certain tests, since we had two experts, we would use the expert expert level of agreement as a baseline.

11:54 The methodology here was to provide a context in which a design description along with the expert raider rating was provided for a number of designs and then a query was given where a new design concept was given but there was no rating provided and the judge was asked to provide the rating themselves. These are encoded into input tokens and then run through a VLM either using reasoning or not.

12:19 And reasoning in this case is inference time reasoning via chain of thought thinking and then ultimately a rating was provided. So I won’t go too much into the details of the methodology here. I implore anyone who’s interested to look at the paper which I’ll link at the end or I’d be happy to chat later. But briefly the four AI judges differ in terms of the content of their context and their query and then also in terms of the vision language model that was used for them.

12:49 So three of them used a non-reasoning model just GPT40 and the last used a reasoning model open AAIS01 and a TLDDR oh because we were using ordinal ratings that are 1 through six because we didn’t assume that our data was normally distributed and because we were just comparing two raiders at a time these were the following statistical tests that we used and this is how they mapped to the six categories.

13:19 So a TLDDR on our findings is that using expert expert statistical scores as the baseline, we found that the AI judge with text, image, and reasoning best matched the expert expert agreement levels for three of the four categories, uniqueness, creativity, and it was actually tied in drawing quality. Something that I thought was more interesting was that in some cases, the AI judge that just used text performed just as well or even better than the AI judge that used both text and image.

13:48 And this is a bit counterintuitive as you’d expect that providing more context, text and image might improve scores. This actually confirms or is is confirmed by some other aspects of literature that are showing this trend too, where vision language models don’t always perform better when provided with image as context. And that’s an interesting future direction on why that’s occurring. And then finally, the AI judge using text, image, and reasoning had better overall performance than two of the three trained noviceses in terms of matching expert expert level agreements.

14:28 And this has powerful implications in terms of evaluation scaling. So if an AI judge can perform as well as trained noviceses who are typically used to replace experts because their time and resources are too expensive, this provides powerful implications for that line of thinking. But the biggest takeaway message was that statistical rigor matters. So the performance of the AI judge changed based on the judge and the metric.

14:54 So knowing that you must validate that the AI judge performs as expected before using it because some of them didn’t match experts agreements in one metric but did in others and so if you tried to immediately transfer one judge to a different metric you might get unsavory results. Finally, for future work, we only looked at four AI judges, but performing comprehensive in context learning and retrieval augmented generation research would be really useful for understanding how to best improve AI’s matching to experts.

15:31 Furthermore, performing these case studies on different design tasks and also studying the trends in which design metrics are easier for an AI judge to predict or not. And lastly, as food for thought, experts themselves differ on some ratings. So determining what the best baseline is and what ground truth really is is a non-trivial task. So all in all, today I introduced the concept of LLM as a judge and proposed that we need a suite of statistical metrics to properly assess how well any AI judge matches with a human raider.

16:07 I proposed these six categories and demonstrated a case study in which we saw the utility of different judges using these metrics. Thank you. To learn more about the CDFAM computational design symposium series, to see the archives of previous presentations, and to learn about future events, visit CDFAM.com.

More from CDFAM Amsterdam 2025

Computational Design, Evolutions

Computational Design, Evolutions

Mathew Vola · ARUP

Injecting AM into shoes

Injecting AM into shoes

René Medel · Framas

Computational design and optimization of vascular stents

Computational design and optimization of vascular stents

Dario Carbonaro · Politecnico di Torino

Open Source CDFAM

Open Source CDFAM

Aaron Porterfield · F=F

Design for Viscosity, Not Gravity

Design for Viscosity, Not Gravity

Hamilton Forsythe · RLP

Manufacturing Driven Design

Manufacturing Driven Design

Rhushik Matroja · Cognitive Design Systems

Strategic Urban Foresight

Strategic Urban Foresight

Ben Dru; Julia Barashkov · Urban Futures Lab

Register for Updates and Discounts on CDFAM events.