What Harry Potter, photo finishes, and Disney remakes can teach us about evaluating AII had just started my annual Harry Potter marathon again (don’t judge me) when my brain once again refused to watch a movie normally.
Right in the middle of the Sorting Ceremony, I turned to my husband with this existential question:
“Do you think the Sorting Hat has good recall?”
For those who aren’t familiar with the term, recall is a system’s ability to correctly find all the cases it is supposed to detect.
Anyway, he went to bed, and I stayed there with my questions.
After all, how can we be sure Harry was supposed to end up in Gryffindor? That Hermione wasn’t actually a Ravenclaw deep down? Is there some kind of ground truth, a committee of experts, a validated set of labels at Hogwarts? Or do we simply assume that, because the Sorting Hat said so, its prediction must be correct?
♦McGonagall discovering that the Sorting Hat has been running in production for a thousand years with no documentation, no benchmark, and no metrics. Image credit: Screenshot from Harry Potter and the Philosopher’s Stone (Warner Bros. Pictures, 2001).This is pretty much the kind of question that has been following me ever since I fell into machine learning a little over fifteen years ago.
Back then, I found it almost magical that statistical models could make complex phenomena somewhat predictable, without necessarily spending weeks looking for the right equation. I loved the idea that, with good data, well-chosen metrics, and a bit of methodology, you could tackle very different problems, from face detection to diagnostic support.
Then, like many others, I gradually drifted toward generative AI.
And if you work in this field, this scene will probably sound familiar: someone shows you a new AI agent with stars in their eyes, the demo is impressive, everyone is nodding along… until someone asks the uncomfortable question:
“Okay, but how do you prove that it works… every time?”
That’s usually the moment when we go from “this is incredible” to “well, when I tested it on one or two examples, it seemed pretty good.”
♦Very useful accessory when someone asks for your evaluation dataset. Image credit: Screenshot from Harry Potter and the Philosopher’s Stone (Warner Bros. Pictures, 2001).In my previous world, the world of “traditional” machine learning (I feel like a dinosaur when I write that), the evaluation framework was relatively clear: a training set, a test set, a few metrics, and one simple objective: show that a model performs well enough on representative data to be deployed into production.
Of course, every use case came with its own subtleties, but at least the playing field was clearly marked.
♦When I explain to junior data scientists that, “in my day”, we actually trained our own models. (AI generated image)With generative AI, many of our data scientist reflexes need to be reconsidered.
- How do you measure quality when there isn’t always a single correct answer?
- How do you compare two assistants that both produce acceptable results, but with different styles, reasoning processes, or trade-offs?
- And most importantly, how do you demonstrate that a new version is genuinely better than the previous one?
♦When OpenAI announces yet another new model and I realize we’ll have to reevaluate everything, again. Image credit: Screenshot from Up (Pixar Animation Studios, 2009)For me, this is one of the biggest challenges: the hardest part is no longer building the algorithm itself (thank you, pretrained models), but proving that it does what we expect it to do, and that it will continue doing so over time.
At its core, the problem goes far beyond AI. In everyday life, we spend our time evaluating things like exams, movies, restaurants, and hotels. And we always come back to the same questions: Is it good enough? And when several options seem good, which one is actually the best?
That’s exactly why the Sorting Hat fascinates me. In a way, it already does what we ask many AI systems to do: gather information, arbitrate between sometimes conflicting criteria, interact with the user, and then make a decision. So before talking about benchmarks or metrics, we first need to understand what it really means to make a good decision.
For the rest of this article, I’d like to propose a little game.
I’ll show you a few images. Each time, your mission will be very simple.
Answer the question: “Which one is the best?”
It shouldn’t be too difficult. Well… at least not at!
When Everyone Agrees on What “Best” MeansLet’s start with this photo.
♦Without a finish-line camera, who would you pick? SourceIf I ask you “who is the best?”, most of you will answer without hesitation: the one who crosses the finish line first (the torso, not the head).
When the differences become invisible to the naked eye, a photo finish settles the matter down to the thousandth of a second. That’s how, in the picture, Lyles beats Thompson at the 2024 Olympics, despite having the exact same official time.
This is the kind of thing that makes a data scientist happy: a clear evaluation criterion and a reliable measurement system. Once the race is over, there is rarely much room for debate: the best is the fastest.
Let’s move on to the second example.
♦Screenshot from Cars (Pixar Animation Studios, 2006).I‘ve always had a soft spot for Pixar, but since I have two boys, let’s just say I’ve probably watched Cars more often than is reasonable to admit publicly… At first glance, we’re facing the same problem as before. Three competitors arrive almost at the same time. We just need to see who crosses the line first, right?
Not so fast.
This image comes from the end of the first race in the movie. Lightning McQueen is in bad shape: he has just lost a tire, part of his bodywork is falling apart, and in one last desperate effort, he sticks out his tongue just before the finish line. The judges then examine the photo finish and conclude that all three cars finish in a tie.
Obviously, I couldn’t stop myself from going down the rabbit hole again. Can we consider the tongue as part of the car? And if the answer is yes, why do we completely ignore the tire that came off a few meters earlier? At what point does a car that is losing pieces stop being the car we are trying to measure?
The photo finish works perfectly well. The problem is that we first need to know what it is actually supposed to measure. As long as we stay with simple cases, the rule seems obvious. Then an edge case appears, and suddenly we discover all the ambiguities we never formally defined.
The lesson is simple: precisely defining what we are trying to measure, while thinking from the start about the weirdest possible situations, avoids a lot of debates later on. In AI, as elsewhere, it is often the exceptions that reveal how fragile the framework really is.
When You Need a Scoring GuideLet’s move on to the next example. Still the Olympics, but this time we’re going to talk about ice dancing.
♦Source: Wikimedia Commons, photos by Luu, CC BY-SA 4.0 and Rama, CC BY-SA 3.0 FRTo be honest, when I watch this kind of competition, I’m usually incapable of saying who truly deserves to win. Unless there is a spectacular fall or an obvious mistake, I quickly find myself ranking competitors based on highly scientific criteria such as “I liked that one, the music and costumes were nice.” And yet, the judges manage it without directly answering the question: who is the best?
They start by breaking it down into several smaller questions. They look at technical difficulty, quality of execution, choreography, interpretation, skating skills… then they assign scores for each of these aspects before combining them into a final score. The lesson here is that when a concept becomes too complex to be summarized by a single metric, we build a more detailed evaluation guide. This is exactly the idea behind the rubrics used for some AI systems (but, spoiler alert, that topic deserves an article of its own).
You have probably seen those 3D slow-motion replays shown during competitions. For the last few years, judges have had access to tools that allow them to analyze certain movements with far greater precision. When technology makes it possible, we can therefore delegate part of the measurements to automated tools. And yet, nobody has seriously suggested letting cameras alone decide the ranking.
Part of the evaluation remains deliberately human, especially everything related to artistic interpretation. And to limit subjectivity as much as possible, the most extreme scores are discarded before calculating the average.
Ice dancing reminds us that a good evaluation does not always rely on a single measurement. The more complex the problem, the more we need to multiply perspectives, accept a degree of judgment, and make our criteria as explicit as possible.
When the Real Answer Does Not Exist YetAfter that sporting digression, let’s return to my current favorite topic and take a look at these two posters.
♦If I ask you, “So, which one looks like the better adaptation?”, I’m almost certain I’ll trigger a debate longer than the seventh book. Harry Potter fans have already survived the great books versus movies battle, so let’s just say they’re well trained. And yet, we do have a few clues available.
The series will have more time to develop side plots and characters that the movies sometimes had to make disappear with a wave of a magic wand. The special effects will benefit from twenty years of progress, and we can already look at the cast, the writers’ experience, or the budgets involved.
In short, there is no shortage of signals.
Personally, I am waiting for the series with as much anticipation as I once waited for my Hogwarts letter. I’ve already subjected my husband to several perfectly objective analyses explaining why this adaptation has every chance of being better than the movies. He claims that I sometimes confuse prediction with personal preference.
I think he simply lacks confidence in my evaluation methods.
Even with the best possible preparation, there is one thing producers cannot fully anticipate: the audience’s reaction. You can prepare for quality, frame it, test it… but you cannot decide in advance how a series will be received.
In other words, the verdict will happen in production. Only then will we know whether the magic works once again.
For an AI system, it’s much the same story.
Upstream, we can consult experts, build evaluation datasets, run benchmarks, and multiply testing campaigns. All of this is essential for comparing solutions, anticipating problems, and reducing risk.
But even the best testing environment never completely reproduces reality.
The real evaluation begins when the product meets its real users, with their expectations, their habits, and above all, all the unexpected (and sometimes slightly ridiculous) ideas they will come up with for using it.
When “Good” Depends on Who You AskOne last image before we wrap up!
♦Non-exhaustive list of Disney live-action remakes. SourceEvery time a live-action remake is released, and there have been many of them, the same criticisms come back:
“Nobody asked for this remake.”
“The original wasn’t respected.”
“Disney was better before.”
Why does The Walt Disney Company keep making them if they are so disappointing?
Because the box office often tells a different story: despite mixed reviews, many of them have performed very well.
♦Left: critic scores for the animated movies (grey) and their remakes (color). Right: comparison of box office revenue. SourceAnd even when the immediate results disappoint, the objective may be elsewhere and over a longer period of time: relaunching a franchise, allowing parents to introduce the movies from their childhood to their own children, refreshing merchandise lines, and extending the life of these worlds in the parks. The movie is therefore not evaluated only as a movie, but as a piece of a much larger ecosystem.
As a fan, I will mostly judge its faithfulness to the original, the emotions it creates, its ability to bring something new, and, of course, to recreate a bit of magic. Disney is probably looking at different indicators: ticket sales, Disney+ subscriptions, merchandise sales, park attendance, and renewed interest in the franchise, probably among many other metrics we will never know about.
So, are these movies “good”? It all depends on the person doing the evaluation and the metric being used. We are watching the same movie, but we are not trying to answer the same question.
This brings us to our final lesson: sometimes, the success of an AI system is not measured by the most obvious metric at first glance, but through a broader reading of the business objectives.
So, Would the Sorting Hat Make a Good AI System?I’m still not sure I’ve answered the question. And what if, from the very beginning, it simply wasn’t the right question?
While trying to evaluate the Sorting Hat, I somehow ended up talking about photo finishes, Lightning McQueen’s tongue, ice dancing, HBO’s Harry Potter, and Disney remakes.
More than anything else, all of these examples show one thing: evaluation is not just about metrics.
In the 100 meters, everybody knows what needs to be measured.
In Cars, you first need to decide what actually crosses the finish line.
In ice dancing, multiple criteria matter.
And for a TV series or a remake, not everybody is looking for the same kind of success.
So before building a benchmark, it is worth asking:
- What are we actually trying to measure? With what measuring tool
- What does success look like?
- What are the edge cases?
- Who is judging the result?
- And most importantly: good for whom?
That is probably the real difficulty of AI evaluation: the hardest part is not always finding the right metric, but asking the right question.
As for the Sorting Hat’s recall, I’m still stuck.
But my next Harry Potter rewatch promises to be a very relaxing experience for my husband.
♦Would the Sorting Hat make a good AI system? was originally published in Code Like A Girl on Medium, where people are continuing the conversation by highlighting and responding to this story.