Search Captions & Ask AI

Rethinking How We Measure Soccer Performance

July 08, 2026 / 56:03

This episode of Wharton Moneyball features discussions on applied probability, expected goals (XG), and XG plus in soccer analytics. Guests include Eric Bradlow, Adi Weiner, and PhD student Jonathan Pippin.

Eric and Adi discuss the Wharton Moneyball Academy and its impact on students pursuing careers in data science. They highlight the importance of understanding applied probability and its applications in sports analytics.

Jonathan Pippin explains his research interests, focusing on applied probability and statistical machine learning. He discusses how modern computing has changed the approach to probability problems and the significance of simulation in research.

The conversation shifts to expected goals (XG) in soccer, with Jonathan detailing how XG is calculated and its relevance in evaluating player performance. He introduces XG plus, which aims to account for situations where shots are not taken, enhancing the analysis of attacking opportunities.

Listeners learn about the potential of XG plus to improve player evaluations and its implications for trade values in soccer. The episode concludes with student questions about expected assists and tackles, emphasizing the evolving landscape of sports analytics.

TLDR

Wharton Moneyball discusses applied probability, expected goals, and XG plus in soccer analytics with Jonathan Pippin and co-hosts Eric Bradlow and Adi Weiner.

Episode

56:03
00:00:00
Welcome, welcome everyone to this week's edition of Wharton Moneyball. I'm Eric Bradlow, Professor of Marketing Statistics and
00:00:07
Data Science here at the Wharton School. I'm joined today, it's a very special occasion
00:00:11
for us here at Wharton Moneyball. I'm joined by my colleague, long-time collaborator
00:00:16
and friend, Adi Weiner, Professor of Statistics and Data Science. He's actually right now in the classroom with
00:00:23
his Wharton Moneyball Academy. Today, we'll also be joined by a statistics third-year PhD student from the Statistics and
00:00:30
Data Science Department, Jonathan Pippin, who I think I'm allowed to call JP for the purposes
00:00:35
of today. And we're going to be asking him lots of questions. And of course, everyone knows we appear here
00:00:42
on the Wharton Podcast Network, and it's some combination of myself, Adi, Cade Massey, and Shane
00:00:47
Jensen each week here on Wharton Moneyball. So Adi, let's first start with, as a
00:00:53
parent of two alums of Wharton Moneyball, this is obviously an exciting time of year for
00:00:58
you. How are things going with the academy? How is everything going? How are you enjoying things?
00:01:04
As usual, we're having a great time. It's actually only our second day. And in the past, we actually have had
00:01:10
a couple of opportunities for our Wharton Moneyball Academy students to sit in on our podcast,
00:01:16
but that was in the studio. This is our first time with Riverside Technology and the video technology.
00:01:21
So we're here in the classroom. I think we're off to a terrific start. It's our second day.
00:01:26
Our visitor this morning was Neil Payne, who is our longtime, almost probably the most frequent
00:01:32
guest on our show. And maybe he'll make a pop in in the second half of our show today.
00:01:38
Well, just to do a little promotion, not that you need it because these students have
00:01:41
already are in the program, but for all of you, you're in for a real treat, not just as Adi's colleague and co-author,
00:01:48
but as I mentioned, my middle and youngest sons both went through Wharton Moneyball Academy.
00:01:53
My middle son is now has a career in data science, thanks to Adi, not just the Moneyball Academy, but the time he spent
00:02:00
at the Wharton School. He's coming back to Wharton for his MBA, but he was an undergrad with Adi in
00:02:05
his research lab, competed in the NFL, was a winner actually of the NFL Big Data
00:02:10
Bowl. So all of you are off to a great start to your careers of being there at Wharton Moneyball.
00:02:17
So Adi, why don't we jump in today? We're very fortunate to have a PhD student, as I mentioned, from the Department of Statistics
00:02:25
and Data Science, both here in my home department, Jonathan Pippin, JP. I'm just going to read what his research
00:02:31
interests say, and then Jonathan, I'm actually going to read them one at a time because
00:02:34
I'd like you to say to our listeners what each of these topics mean. So it says his research interests include, let's
00:02:41
start with the first one, which may not, it sounds like it should be obvious, but
00:02:44
may not be obvious to everybody, applied probability. What does it mean for someone to do
00:02:49
work in applied probability? Well, first of all, hi, Eric, and thanks for having me.
00:02:57
Well, yeah. Applied probability basically involves taking the concepts that we learn across math classes, probability classes in
00:03:07
statistics, and then taking those and applying those to particular problems in a bunch of different
00:03:12
fields. So today we're going to talk about soccer. We'll talk about how it's used in expected
00:03:17
goals. Additionally, expected goals, like a bunch of other things, it's probability models, and learning, I would
00:03:27
say, learning to, applying what you learn in a probability class to different contexts.
00:03:33
So that could be a particular sport. It could also be causal inference, as you'll
00:03:39
mention some of my other interests, and plenty of other fields that involve applying what you
00:03:45
learn, distributions, moments, variants, all the things you talk about on this show to a bunch
00:03:51
of different problems in real world problems. What would you say, Adi and I have
00:03:55
had this discussion on the air many times over the last 12 years, what would you
00:03:59
say to someone that said with GPUs, massive simulation possibilities today, what do we really need
00:04:06
math and probability for anymore? Can we just simulate out the entire world? We'll get all the intuition we need.
00:04:12
We don't have to rely on asymptotics. We can just do whatever finite sample inferences
00:04:18
we want. What would you say to people like that? By the way, I'm not in that camp,
00:04:23
but why would you waste your time with all this math and probability theory? Why not just learn how to code and
00:04:28
forget even coding? Why not just use AI, create, use Claude code, code everything up in a workflow and
00:04:36
then blast it through? That's a great point, Eric. Modern computing has totally changed how people can
00:04:43
approach these problems. Simulation, leveraging coding agents or what you can code up in Python.
00:04:50
It's tremendously useful. A lot of times it can give you a head start on finding an answer to
00:04:55
what you may be interested in in a probability problem. Professor Weiner and I actually, we did one
00:05:01
of- Just this. We did one of these for one of our research projects, which involved simulating something out
00:05:08
before we ultimately went to math. In fact, the way we approach that problem, and we've actually talked about it on our
00:05:12
Moneyball show before, was the question is, what's the maximum win probability obtained by the losing
00:05:18
team over the course of a game? Which is something that most people don't quite grasp because the losing team ends with a
00:05:25
win probability of zero, but they start in many situations with a win probability of close
00:05:30
to 50%. The win probability would be calculated at any point during the game. If you look at the trajectory of the
00:05:36
win probability, let's say it starts at 50%, and then it moves up and down depending
00:05:40
on what happens during the game. Eventually, the losing team ends at zero. You ask, what's the question?
00:05:45
What's the maximum win probability obtained? Well, we eventually solved that problem as a
00:05:51
theoretical problem, but not until we first simulated it. In fact, it was the simulation that gave
00:05:57
us the hints on how to solve it. That's fascinating to me because my training and
00:06:03
yours, Adi, I think, in the old days when we were doing math stat, would have been, you do math first, that gives you
00:06:10
a guide to what to vary in the simulation design, as opposed to, let's do simulation,
00:06:16
see the patterns, possibly come up with a set of results, and then see if the math, in some sense, can explain those patterns.
00:06:23
Well, in fact, one of the things we did by simulation, and also empirical evaluation.
00:06:27
Just to be clear for our listeners, a simulation begins with a model, a model for
00:06:32
where, say, football, or baseball, or basketball is actually played. It's like coin tosses, certain probability of a
00:06:39
score, and then you just simulate that and play fictitious games. It's almost like what we used to call
00:06:45
strat-o-matic. I don't think our kids back here have ever heard of strat-o-matic, but that's
00:06:50
from our time. We did that, but what we also did was- What sport did you do it? We did it, I think it was football,
00:06:58
and we did it also with basketball. Football, we simulated the outcome zero, two, three,
00:07:05
seven, eight, and for basketball, we simulated twos and threes. That gave us one way of doing it.
00:07:13
The other way we did it was empirically. Empirically means we just observed what was the
00:07:18
distribution of the maximum win probability in the sports that we've seen. That already told us, and by the way,
00:07:24
the fact that the two numbers were pretty close to start with was actually quite illuminating.
00:07:29
That's shocking to me. Here's what I would have guessed, and then this is why you guys can tell me
00:07:34
I'm entirely right. By the way, don't worry, we're going to get to XG and all that.
00:07:37
We got plenty of time here. We're going to get to soccer, don't worry about it. I would have guessed that here's what would
00:07:42
have happened. You would have posited some model of, let's call it how scores generate, a generative model.
00:07:50
You run a simulation, you match it to the empirical data, you notice it's deficient in
00:07:55
a number of ways, then you go back and you fix the model until it matches in some sense the marginal distribution that you
00:08:02
observed in the empirical data. That's what I would have guessed, and that the initial model would have been so far
00:08:08
off that you would have had to make, I'll call it more than just minor tweaks, it would have been like major surgical triage.
00:08:16
Well, I'll allow that JB to tell you, but the model for football that we used was actually pretty good right off the bat.
00:08:22
What happened with basketball? That's actually interesting. Yeah, it is. The model for football actually very closely matched
00:08:28
what we saw empirically. Our simple simulation model, once you tune a couple of parameters, it matched fairly closely.
00:08:36
What we saw in basketball were some major deviations from what we would expect in a
00:08:42
simulation based on team strength and what we would assume a team's shooting profile would be.
00:08:48
In basketball, what we found is that a lot of the win probability estimates that are
00:08:53
out there are systematically too extreme. We were trying to quantify the maximum win
00:08:59
probability attained by losing teams. What we found is that the win probability models that are out there largely overstate or
00:09:07
get too confident too quickly in paths or teams that eventually lose. So it was a pretty significant overshoot of
00:09:18
what we would expect. Just to be clear, when you say overshoot or too confident, do you mean that their
00:09:27
probability of winning the losing team never gets high enough or it's the opposite?
00:09:34
I'm just trying to know which direction you're claiming. It's the opposite. What we found was that when a team
00:09:40
is ahead, the win probability models in basketball tend to give, more often than not, an
00:09:46
overconfident estimate of how likely that team is to win. You'll see more collapses.
00:09:51
That's really what we were looking at. Collapses of teams that go on to lose had an abnormally high win probability more often
00:09:58
than we would expect under a well-calibrated model. I see. Actually, one of the things that's very fascinating,
00:10:09
when two equal teams play each other, the average win probability obtained by the losing team
00:10:15
is greater than 70%. I would have guessed the number was somewhere between 70 and 75, by the way.
00:10:22
That would have been where I would have guessed. It's interesting that I was at least somewhat
00:10:27
calibrated at all. And in half the games, the losing team is the two-to-one favorite at some
00:10:34
point. Right. That would have been my guess. I was shocked by that because you think,
00:10:40
God, if you've lost... So we call it actually the paradox of the blown lead, which means that we observe
00:10:46
that losing teams very frequently blow a big lead. Right now, the students, we had to interrupt
00:10:56
them. Egypt was winning two to nothing over Argentina. I don't know. By the time this is aired, you will
00:11:01
know the advantage of this. But this is a team that now has a very big lead probabilistically.
00:11:07
I don't know what it is, but I would imagine two to nothing, middle of the game in soccer is probably at least 80%.
00:11:15
And we'll see what happens and see whether they blow it. But that's the way to say it, by
00:11:19
the way. I think we'd all agree a two-to -nothing lead in soccer is calibrating leads.
00:11:24
This has been done for years in statistics, calibrating. One of my advisors, Hal Stern, who worked
00:11:30
with Tom Cover, who was one of your advisors, Adi, he wrote a paper about equating
00:11:36
leads in different sports. So this has been around in the statistics field for a long time.
00:11:40
Like two-nothing in soccer with 30 minutes to go is equivalent to this lead in basketball or football, et cetera.
00:11:48
It's actually one of those interesting calculations to do. Let me just go to the next one,
00:11:51
JP, just quickly. And we're going to get to soccer in two more minutes. So statistical machine learning.
00:11:58
So prior to generative AI and I'll call it AI agents and supportive tools for coding
00:12:05
and data science, we, the thing I started 20 years ago really was, I didn't call
00:12:12
it this, but it really was a center for big data and machine learning. So not just for myself and our audience,
00:12:18
could you talk to the students that are behind you also, why you still think having
00:12:23
training in machine learning skills are important and which skill do you think is going to
00:12:28
be more important for you going forward? I was just asking you about data engineering
00:12:32
versus the data science part of machine learning. Is it piping the right data so that
00:12:38
you can train an LLM or analyze it properly? Or do you think it's still the learning
00:12:43
models like XG boost and boosted trees and different types of predictive models?
00:12:51
Which part do you think will be more valuable for you in your academic career? And also as you try to work on
00:12:56
applied problems? That's a great question. Yeah, no, they're definitely, the presence of agentic
00:13:03
coding tools has definitely made it easier to implement these things, right? Anyone can open up a chat GPT window
00:13:09
or their favorite LLM and code up a machine learning model like that. But all the same steps that were still
00:13:17
important then are still important now for researchers. So, deciding what data you're going to use,
00:13:22
deciding the problem you want to solve, and then understanding what these algorithms are actually doing
00:13:28
under the hood is actually extremely important still to actually understand the inductive biases and all
00:13:35
the things that go into what's going to make your prediction come out. So, a black box model, you can now
00:13:41
ask a black box model to write you a black box model to give you some output, right?
00:13:47
But the same skill, if you actually want to understand where these predictions are coming from
00:13:53
or the architecture under which these models are giving you a prediction or an output, that
00:13:58
still requires deep study. And a lot of what I'm interested in is quantifying the uncertainty in these models and
00:14:05
actually trying to interpret them in a way that allows you to do something like inference
00:14:11
or something like I actually understand sort of what the parameters underneath the model are saying.
00:14:16
So, that's more my research interest, but I definitely think there is still a lot of
00:14:21
need. You still definitely need to know what your model is doing if you actually want to
00:14:27
- We actually, this actually came up over lunch. Maybe in a couple of weeks, we'll talk
00:14:31
about this, but everyone has been fascinated by the Jalen Brown trade. And one of the reasons why is that
00:14:37
the black box models, I mean, the most sophisticated black box models produced by analysts are
00:14:43
saying that Jalen Brown is very overrated. But they can't tell you where that, why
00:14:49
that conclusion comes from. One of the things that Neil thinks is that that's inaccurate because it has to do
00:14:54
with usage. And we don't have data for players who have such enormous amounts of usage, which means
00:14:59
that you're extrapolating with these models. And secondly, and this is what everyone admits,
00:15:03
the models can't tell you why. They can just give you a number. They don't say what component of his play,
00:15:10
what's causing this, or even forget about what fraction of the things you observe are causing
00:15:16
the prediction to be what it is. They don't work that way. And that's where deep knowledge of statistics can't
00:15:22
be replaced. And maybe also we would argue deep knowledge of basketball, because eventually you're going to have
00:15:28
to say what aspect of his game, is it a defensive liability? Is it rebounding? Is it guarding perimeter players?
00:15:35
I mean, you can, in other words, there are so many possible patterns that could lead
00:15:39
to his being overrated or not. And one would need some direction to learn. Let me just ask the last one.
00:15:45
I'm going to skip causal inference only because my advisor was Don Rubin. I understand causal inference.
00:15:50
I'm going to skip that one. What the hell is selective inference? I know it's a well-known term.
00:15:54
I just, I'm not as familiar with it. What is selective inference? Selective inference is a field that sort of
00:16:00
concerns itself with making valid inferences after you've already used the data to select, you know,
00:16:07
what you're testing or how you're testing it. And so, you know, you see this a lot of the time in sports, right?
00:16:13
You may be interested in, you know, evaluating who's the best player in the NBA or
00:16:19
by some metric. Well, you're using the data already to tell you who's good and bad.
00:16:25
At that point, how do you actually estimate who really is the best, right? And selective inference on top, you know, after
00:16:31
you use the data or some data-driven mechanism to create an inference problem, how do
00:16:35
you still use, you know, you still want to use all of your data. You don't necessarily want, you could split your
00:16:40
data. There's a lot of different, you know, methods, but ultimately it's about, you know, being honest
00:16:46
in the inference procedure about what inferences we really can say from data once you've already
00:16:51
used the data to sort of get you halfway there. Right. So, so the easiest example, most, many of
00:16:56
our listeners, I'm sure even quite a number of our students, even though they've started the
00:17:00
class, I've heard of what's called the confidence interval. And we all know what that is.
00:17:05
And that gives you a range of possible values. But if you've built the model using the
00:17:11
data and now you're using the data again to build the standard errors, you've just done
00:17:17
a selective process first. And now you need to figure out what the implications of that selections are on the
00:17:23
uncertainty. And it is dramatic. It can be, depending on how big a model process, how much you've used the data,
00:17:31
it can be huge. So if you consider a model with say 10 variables and you decide on two and
00:17:36
build confidence intervals for those two, those confidence intervals are way too narrow because of the
00:17:42
selective inference. Would you guys describe the winner's curse as a specific example of selective inference, which is
00:17:48
the classic, you know, you have a bunch of competing athletes or models, you pick the
00:17:52
winning model, then you do inference giving the winning model. And of course, the results are never as
00:17:58
predictive from that winning model as you would expect. Confidence intervals are always too narrow because it's
00:18:04
a form of selection first and utilization second. Absolutely. Yeah. The winner's curse is like the classic, classic,
00:18:12
you know, selective inference example. Well, we're at the World Cup time. So now let's move to the main topic
00:18:19
of today. And of course, we're going to leave time for students to ask questions of both Adi,
00:18:24
myself, and JP. So first it says in the notes here that you've been doing some work in XG
00:18:30
and computing something called XG+. But why don't we just start with, you know, the basics.
00:18:37
What is XG? And why do you think it's considered an advanced stat and a very popular statistics for
00:18:44
soccer? Yeah, expected goals are probably the most central, you know, analytic statistic in modern soccer.
00:18:53
It, you know, it involves basically, you know, first of all, expected goals, it's not this
00:18:58
absolute value, you know, this true value that you can observe. It's not a counting stat.
00:19:02
Yeah, it's not a statistic that you will just see. And it's like, ah, that's the expected goals.
00:19:07
It's a model-based statistic. So there's, you know, this involves a probability model, like we mentioned earlier.
00:19:12
And it, you know, involves estimating the probability that a given shot, once it's taken, is
00:19:19
scored. And these models vary in complexity. They can include anything from just the location
00:19:24
of the ball on the pitch, to how many other players are around the ball, to how open is the goal, where's the goalkeeper,
00:19:31
you know, maybe even what time is, how much time is left in the game, what, you know, what player is taking the shot.
00:19:36
They can get very detailed, but it ultimately is an estimate of how likely a given
00:19:41
shot is to be scored once it's taken. And I think that's, there's a couple reasons
00:19:46
why it's, you know, so commonly used. But one that's so nice is it basically tells you the quality of a shot.
00:19:52
And so, you know, a shot that is taken that's a 0.5 XG shot would mean, you know, we estimate that this shot
00:19:59
is about 50% to go in once it's taken. Who does it, I could imagine it being used to grade, if you'd like.
00:20:07
We're in a classroom here, I see a bunch of students. I could use to see it grading maybe
00:20:11
three different components. One is, how good is the player that shot the ball, right?
00:20:16
Because if they perform better than what the stats should say, given all the, that's one,
00:20:21
you could use it to score the goalie. You could use it to score the team at some level.
00:20:26
Where do you most see XG being used? I would imagine it's at the team level, but it could also be at the player
00:20:33
level. I mean, how do you see it being used? Yeah. So the main place that you and many
00:20:38
others will see it used is at the end of a match. You'll see, you know, this team accumulated two
00:20:43
and a half XG in this game. And this other team accumulated one and a half XG. It's sort of used to estimate, you know,
00:20:50
the cumulative quantity, quality of the chances a team took. A lot of times it's used to sort
00:20:55
of, in post-match analysis, like who deserved to win. Since goals are very rare in soccer, like
00:21:00
a lot of games end one nothing or even nil-nil, right? How do you grade a 0-0 tie?
00:21:06
Well, that's graded based on the quality of the chances created. And so you'll see it a lot with
00:21:10
teams using to compare or trying to say which team was better, which team created more
00:21:14
dangerous chances. But like you said, it's also used for individuals. So, you know, like you could residualize.
00:21:21
You could say, did a player over-perform their XG? Or did a goalkeeper over-perform their XG
00:21:25
in terms of saving more shots than you would expect to go in? It's definitely used for player analysis, but I
00:21:31
would say the main place it's used or the most common example is at the end of a match or even at the end
00:21:36
of a season, looking at the chances that a team racked up over that game to sort of determine maybe who deserved to win
00:21:42
or who was more threatening. It's just a way of averaging out some of the randomness, which is differentiated from the
00:21:52
skill component. Although it can be very fascinating on an incident base. So yesterday's America-U.S. match against Belgium,
00:22:00
the American goalie had one, at least to my eyes, looked like a spectacular save.
00:22:05
I didn't know because I don't, you know, I'm the first to admit I don't know that much about soccer.
00:22:10
Looked good, but the eye test isn't the same as an analytical model. So I want to know what was the
00:22:15
XG of that shot? If the XG of that shot was 0 .1, meaning it wasn't likely to go in,
00:22:21
I'd be, I should be less impressed by the save. If it was 0.9 and it didn't go in, then I should feel that maybe
00:22:28
that was a great save. So that's how it can be used. And maybe someone will tell me what the
00:22:32
XG on that shot was in this vast majority of students back here. Does anybody know the answer?
00:22:37
Do you know? You mean like before the shot happened or after the shot? That's a great question.
00:22:40
After the shot. Well, while he's taking it, XG is calculating. Before it's in the air.
00:22:45
Yes. I don't know about that one. Well, would you know any other number? Yeah, I think it was like a 0
00:22:50
.25 after it was in the air. After it was in the air, it was 0.25, which means it's pretty impressive, but
00:22:56
not particularly. So before we get into what's wrong with XG, because why would you need XG plus
00:23:02
if XG is brilliant, would you agree with the following? That if I had an infinite amount of
00:23:08
data, I would not need a model. Like I could literally just say, well, I could look at every place on the pitch.
00:23:15
I could look at every situation of players and their locations. I'd have just some big empirical lookup table.
00:23:22
I just pinpointed, and then I've got some empirical estimator. I wouldn't need any smoothing because there's no
00:23:27
discontinuity. I have infinite data. If I've got infinite data, I've got infinite
00:23:31
data, right? Yeah. Yeah, of course. If you have infinite data, then the problem solves itself.
00:23:36
You just look at, like you said, a binned average and you're good. We don't have that.
00:23:40
We have finite data and we have strategically chosen data. Players aren't shooting from random places on the
00:23:46
pitch in random configurations all the time. And so that's sort of where the model
00:23:51
comes in to sort of help us alleviate some of those biases and maybe smooth over
00:23:55
places where we don't actually have that much data. Okay. Now let's get to XG+. So what is XG plus?
00:24:03
XG plus came sort of as a response to some of the limitations of XG. XG, expected goals, is a great stat.
00:24:11
It tells you exactly what it tells you, which is how likely is a shot once it's taken to go in.
00:24:16
But many opportunities never have a shot realized. You see this in plenty of games.
00:24:23
Mexico played England just a couple of days ago in the World Cup. And Mexico crossed the ball dozens of times
00:24:29
into the box. And there were many times where a Mexican player did not get their head to the
00:24:34
ball. But if they had, it would have been a great chance, right? Expected goals, XG, only tabulates, it only accumulates
00:24:41
once a shot has been observed. So if there's an attack where there's no shot, but there was a very dangerous opportunity,
00:24:48
that looks like a zero in expected goals. And we've sort of wanted to set out to sort of fix that.
00:24:54
And also maybe mend another problem with XG, which is on a single attack, if you
00:24:59
have a bunch of rebounded shots, so imagine you shoot the goalkeeper saves, you shoot the
00:25:02
goalkeeper saves over and over and over in a short period of time, it's actually possible
00:25:06
to accumulate over one expected goals, which is not possible on a single attack.
00:25:12
And that's just, again, one of the maybe limitations of XG is accumulating XG. So we wanted to set out to sort
00:25:19
of solve some of these problems. Yeah. There's XG plus then, if it's not based on just whether a shot is taken or
00:25:25
not, is it based on, when you say the situation, is it the ball? Is it the ball and all the players?
00:25:32
Like what is the situation that you then score and make up XG plus? Yeah, great question.
00:25:40
The situation includes the ball and the other players and all the normal things that would
00:25:45
go into an expected goals model. The key distinction with expected goals plus or
00:25:50
XG plus is what we're calling it, is that we're not just modeling the probability that
00:25:55
there's a goal scored on a shot. We're also modeling the probability that a shot
00:25:59
is taken in a given situation. And so expected goals gives you something like a conditional probability, right?
00:26:05
You're estimating the probability of a goal given a shot. But the danger, where you're really interested in,
00:26:11
is the probability that the goal is scored on an attack. And that marginal probability involves not just the
00:26:17
conditional probability of scoring once you shoot, but also the probability of shooting in the first
00:26:22
place. And so the key development of our method is to estimate that probability that a shot
00:26:27
occurs from each sort of configuration, including the players, including the openness of the goal, including
00:26:33
the situation in general, and then factoring that into the expected goals to give a more
00:26:39
complete picture of the attack. So I have two follow-up questions to that, and maybe the students will have others
00:26:46
too. So one is, I think I know when expected goals is computed. It's computed at the time of a shot.
00:26:53
When is XG plus computed? That's number one. And then two, what kind of data requirements
00:27:01
must be there to be able to compute XG plus, possibly even greater data requirements that
00:27:08
are there then for XG? So when is it computed? And how has modern data tracking or video
00:27:15
tracking or other forms of tracking helped and now made it even feasible? Another great question.
00:27:22
So in order to estimate XG plus, the key development or the key addition is being
00:27:28
able to estimate the probability of a shot at a given position for an attack. And so that involves having estimates of where
00:27:37
the players are, where the ball is, continuous in time. And so we have some tremendous data providers
00:27:43
that are working with us to provide us with video-based tracking data, so derived from
00:27:48
broadcast footage. But essentially, you can imagine a 2D grid of little dots moving around on the pitch.
00:27:55
And so using those dots and really different features we can create from them. So we might imagine how far the closest
00:28:02
defender is from the ball, sort of matters not just for the likelihood of a goal
00:28:07
being shot and scored, but also the likelihood of a shot being taken. That closest defender goes into our model as
00:28:13
a feature, and as do some of the other defenders, as do the goalkeeper, and it does so continuous in time.
00:28:19
So continuous over time, we're estimating something like the probability that a shot occurs in some
00:28:24
fixed window. And so we have our data is estimates 30 times a second where each player is.
00:28:33
You might have coarser feeds or something like that, but ultimately we're just estimating in the
00:28:37
next second, in the next half second, what's the probability that a shot... Okay, so you're choosing a window, which is
00:28:43
helpful. Let me ask another question, which every student in that room and all of us have
00:28:48
to do this when we're working on applied problems. How did you think about the feature engineering
00:28:53
part of this problem? Because in some sense, anything could be a feature. And so did you use knowledge of soccer?
00:29:03
Did you use just brute force empiricism? Did you end up doing some sort of penalized criterion, some feature selection, whether it's some
00:29:11
sort of penalized lasso? Did you end up doing some sort of out of sample validation?
00:29:16
How did you decide what features to kind of jam into this XG plus model? Yeah, so I would say some combination of
00:29:23
those things. So first of all, a knowledge of soccer and how expected goals are generally calculated, which
00:29:29
is it uses the five closest defenders, the location of the goalkeeper, the location of the
00:29:33
person on the ball, the three closest attacking players. These are just common conventions.
00:29:38
You could change some of them. We have tried changing some of them for sensitivity, but ultimately that's just convention.
00:29:46
The critical thing and something I want to make sure that the students hear is how
00:29:50
we actually encode these features. It's very easy to just take tracking data, take the XY features and just throw them
00:29:55
into a model and just trust the model is going to figure out the patterns in the data.
00:30:01
That's where, like we mentioned earlier, knowing the architecture of these models, actually understanding how a
00:30:05
tree-based model, which is often used for expected goals, actually works. Those features, the more informative we can make
00:30:11
them, the better. So what we do is we encode not just the location of the players, the location
00:30:17
of the ball, we encode relative distances. So we're looking at how close each player
00:30:22
is in terms of distance, not just in terms of XY, and also bearing to the goal. The way we encode location and all these
00:30:29
features is very thoughtfully done so that when our high complexity model that may be coded
00:30:35
with an AI or LLM tool runs, it's ultimately picking up on informative features in the
00:30:41
data. We understand that just throwing a bunch of unnecessary features into a model class like this
00:30:48
actually degrades performance. And so that's why we try to be very careful to select what we want to
00:30:53
include, while at the same time, leaving open the possibility to cross-validate or do some
00:31:00
out-of-sample testing to determine if additional features are relevant or should be added or
00:31:05
dropped. Well, I mean, it sounds like one of the things that you learn from knowledge of
00:31:08
soccer is various orientations and coordinate systems matter, right? So if a player's back is to you,
00:31:15
that's very different when the player is front -facing. And you have that in the data?
00:31:19
You have the actual, because this is one of the things, Adi, I think we talked about a few weeks ago on the air,
00:31:25
is that having someone's XY location is different than knowing which way they're facing.
00:31:30
So you actually have that data? Yeah. Orientation data is included in the data set
00:31:37
that we have. In many 2D tracking data sets, it's not included. But there are also certain times where it
00:31:43
can be inferred based on their velocity, where they're running. It's very difficult to run 20 miles an
00:31:48
hour backwards. And so if you have a limited data set and you wanted to construct something like
00:31:54
we did, there are certain ways to try to infer at some of the features that we have.
00:31:58
One of the advantages that we had was we were able to get this data. The tracking data is very expensive.
00:32:04
And we have a partnership with a video -based tracking, which is not tracking with an
00:32:12
RFID trip or something, or using cameras that are in the stadium. This is just using the video feed to
00:32:19
create the tracking out of that. And that worked great. By the way, we'll talk about this.
00:32:24
The students may be interested in this. It just came up on my phone. You'll never believe what just happened.
00:32:29
Argentina scores three goals in the last 13 minutes to stun Egypt 3-2. So you want to talk about it.
00:32:35
I would have said there's no chance that Argentina is coming back to win that game.
00:32:40
See, their win probability had to be 0 % according to a model, according to the JT, the Yachty, the Weiner.
00:32:46
Actually, by the way, this is a fantastic example. The losing team ended up probably had a
00:32:53
max win probability that was quite high. I would guess it's above 80. 90-plus percent, easily.
00:32:59
Yeah, interesting. Although Egypt is how good a team? I'm not sure. You're up 2-0 with 13 minutes left.
00:33:07
But anyway, I want to get back to the XG plus for a second. But sorry, students, you missed it.
00:33:12
By the way, if it makes your students feel any better, I didn't get to watch the last 13 minutes of the game either,
00:33:17
because I've been on the air with you. But they have this thing called review. They have video.
00:33:22
I'll be able to watch it. But, did you actually apply this to actual players? Are there players that appear really good in
00:33:34
XG, but not as good in XG plus? And is there any names or something that might surprise us?
00:33:41
Absolutely. It's been well documented that there are very few players that actually overperform consistently their expected
00:33:49
goals. Just the shots they've taken, expected goals, residualizing that. Is that because of low power, or in
00:33:56
your belief they're just few players? There's a lot of goals, so it's not power.
00:34:02
Basically, just to be clear, we've actually talked about this before on the air, what we
00:34:06
call excess goals. JP has called residualizing, but I want to make sure that everyone understands.
00:34:13
To residualize means to take the numbers of goals, subtract the expected goals. And that difference reflects, it could be positive
00:34:20
or negative, how well you've done relative to the chances. So what you're saying is that nobody really
00:34:26
seems to outperform or even underperform their XG. The ability of a soccer player tends to
00:34:32
be wrapped up in their XG, just getting to the XG. Yeah. And that getting to the XG is what
00:34:40
our expected shots tries to get at. And so we would say the correlation from year to year in your overperformance to XG
00:34:47
is rather low. It's like 0.1. But in expected shots or XG plus, that correlation year to year
00:34:54
with players is like a 0.6 or a 0.7. So it's stickier. I love it. That to me, just for myself, that's the
00:35:04
highlight of what you've said so far in terms of the empirical findings. Because we've talked about the fact that certain,
00:35:11
I'll call it offensive or performance metrics, just do not correlate well across games or seasons,
00:35:17
et cetera. The fact that XG plus is more stable, I think is a strong case for its
00:35:24
validity as a useful metric. Because I do believe the creation of opportunities is probably a very stable construct.
00:35:32
Absolutely. And you asked about individual players. The player that stood out for us on
00:35:37
XG plus and expected shots is Erling Haaland, the striker for Norway, striker for Manchester City.
00:35:44
Well-established reputation of being able to take extremely acrobatic shots and get shots off where
00:35:50
other players couldn't. Like Adi was saying, if you looked at his expected shots, he would take way more
00:35:57
shots than his expected shots would indicate. So that residual would be very high, which
00:36:01
is to say he is creating more shooting opportunities or he is shooting more often than
00:36:06
we would estimate. And we consider that a very sticky skill. That's a skill that we would expect to
00:36:12
stay with him because of how consistent it is year to year. We're finishing variability, that's a little bit less
00:36:18
consistent, making and missing shots. Actually getting to take those shots and as
00:36:23
a team, creating those shooting opportunities, that is, we find to be more predictive than just
00:36:29
pure shot-making ability. So let me ask one last question. Then I want to, audience members, get ready
00:36:34
to start thinking about questions to ask JP or Adi about this. A natural, possibly erroneous conclusion from this would
00:36:46
be a player with high XG plus has to take more shots. They have to, because if they're good at
00:36:55
creating situations or maybe it's they have to get more playing time or et cetera, is
00:37:01
that like, I'm actually trying to combine your knowledge of selective inference with their knowledge of
00:37:06
causal inference to say, is the natural implication of this that a player that does really
00:37:13
well, like the guy you mentioned on Norway, Halon, my understanding is he doesn't take a
00:37:17
lot of shots, but when he takes them, he scores very often. So what's the implication of this for let's
00:37:26
someone that was a team manager or someone that was trying to give advice to players
00:37:30
that do well on this metric? Yeah, great question. So first of all, I just obligatory to
00:37:38
say this, we're working with observational data. So we're working with, you know, our models
00:37:42
are trained based on data that we actually observe in games, which has strengths, but also
00:37:46
its weaknesses in that this is not randomized. We don't necessarily, you know, know if something
00:37:52
is causal per se, but we can find strong associations. And the associations that we find in this
00:37:57
data is that if you were a team manager, for example, and you had a player who over-performed their expected shots, like a
00:38:03
Holland or like another, you know, some of our other elite strikers, like Mbappe, who consistently
00:38:09
over-performed their expected shots, it would be reasonable to assume that their performance in a
00:38:17
following season or in a following game will be more consistent on that metric than it
00:38:21
would be on expected goals, just based on the correlations. As for causality or whether players, you know,
00:38:28
who should take more shots or not take more shots or something that involves an intervention,
00:38:33
that involves a bit of a different mechanism than what we were trying to explore.
00:38:36
So I would caution against using it in an interventionist sort of way. But I would definitely say that, you know,
00:38:43
the correlation, the correlation structure in the data does indicate a much stronger correlation between expected
00:38:50
shot over-performance and XG plus over-performance relative to XG over-performance, which we find
00:38:55
to be very noisy. So JP, before we turn it over to the students that are part of the summer
00:39:00
Wharton Moneyball Academy summer program, can you speak maybe just for a second or two?
00:39:05
I know you have collaborators on the project. I know also I was shocked that you
00:39:09
have this data that you've talked about because I've been trying to get this data for
00:39:13
years. So maybe you could just say a few things about that and then we'll turn it
00:39:16
over to students for questions. Yeah, absolutely. I had some great collaborators on this project.
00:39:21
We had a Wharton sports research senior fellow, Paul Sabin, advised this project for me and
00:39:28
a fellow student, Tianxu Feng, who actually just graduated with a master's in data science from
00:39:34
the University of Pennsylvania. This project actually emerged and was done primarily
00:39:38
in our undergraduate sports research lab, which is every summer. I'm happy to oversee those projects and work
00:39:45
with these very talented students. That came out of that sort of process. It was a tremendous opportunity for undergrad students
00:39:55
here at Penn to get involved in sports research here. And that was where most of the thinking
00:40:01
for this project sort of came from. It's since grown well beyond that. Also, a major thanks to our data providers.
00:40:09
Gradient Sports provided us with this data. We have a partnership with them where they
00:40:12
provided us with this very difficult to get tracking data, something not everyone can get a
00:40:18
hold of. Our partnership with them has been really great. And so if anyone else is interested in
00:40:24
getting that data, they're probably someone you might want to reach out to. Sounds great.
00:40:28
Wadi, I'm going to turn it over to you to bring up the first student and let's start hearing some questions.
00:40:32
Okay. So JP is going to slide over and we're going to bring in our first student.
00:40:36
They've only been on campus for, I guess, about two days. So why don't you introduce yourself, say where
00:40:41
you're from, what your name is, and then ask your question. And of JP, me, you can even, if
00:40:49
we need to bring Neil in, we can bring him also to answer some of the questions.
00:40:53
Yeah. Thank you guys for having me. I'm Jack Brenner. I'm from New York. And my question's tailored more towards JP and
00:41:00
about his research towards XG+. And where you can kind of take that with expected assists and how with XGs and
00:41:09
XG+, you're dealing in a lot of hypotheticals of how, you know, will he be in the right position?
00:41:15
Will he convert it to the right part of the net? And with expected assists, you're taking that even
00:41:20
further in hypotheticals of where the ball will go, where he'll make the run and how,
00:41:24
and I was wondering about how, if there's any thoughts on how to improve expected assists?
00:41:29
Well, first of all, also, it's great to meet a fellow New Yorker. I grew up also in New York City,
00:41:34
went to Stuyvesant High School in New York. So it's always great to meet New Yorkers
00:41:39
as well. Actually, New York is our leading state representative. 20 of our students are from New York
00:41:44
state. All right. Not a surprise. Okay, go ahead, JP. So that's a great question, Jack.
00:41:50
Expected assist models, very interesting. They're trained very similarly to expected goals models,
00:41:55
which is they're trained on observed shots that are taken. An expected assist is only really recorded on
00:42:01
a shot that is taken. That's where you get the expectation is how likely is the shot that the person assists,
00:42:09
right? So you pass the ball to someone and they shoot. What's the likelihood that they score?
00:42:12
And it tries to credit that to you. But again, that only, that has the same problem as expected goals, which you have to
00:42:18
see a shot to actually realize an expected assist, where XG plus doesn't have to realize
00:42:24
a shot. You would still get credit for a dangerous attack, even if there's no shot, if the
00:42:28
likelihood that a shot would have happened in that situation is rather high. And so expected assist could also be taken
00:42:34
with a similar strain as we did, which is to say, what, maybe not the expected assist from observed shots, but the expected assist
00:42:44
from hypothetical shots, sort of like you said, and give people credit for creating those dangerous
00:42:48
opportunities, for making a pass that increases the XG plus, even if there is no shot
00:42:53
that actually comes out. Yeah, I think that is so crude. Your question, Jack, is phenomenal.
00:42:57
A couple things. One is, when I watch soccer, I'm not an expert in soccer. I'm thinking of passes as opportunities for someone
00:43:05
else to score. And then the second thing is, of course, the quality of the person at the other
00:43:09
end really matters. So another thing that's interesting, you know, JP talked about conditional probabilities.
00:43:16
I'm now passing. Well, first of all, did I make the optimal pass? If I have the location of everybody, I
00:43:22
made it. Here's another thing. I may pass it to where there's opportunities to score, but it wasn't the best pass
00:43:27
I could have made. Now, in some ways, if that's true, then I should be penalized for that.
00:43:33
And so then you need some measure of optimality. And then, of course, it's who is on
00:43:37
the other end of it. Like, I don't know, you guys probably watched the U.S. game last night.
00:43:40
Whoever that guy was at the top of the screen that they were leaving wide open, there was no way the U.S. was
00:43:46
even ever going to pass it to that guy because, you know, he wasn't very good. And so, you know, X model might say,
00:43:53
well, this guy's wide open. You've got to pass it to him. Yeah, but he's not going to score if
00:43:58
you pass it to him. So I think it's a great question. Well, that's a hard thing to do.
00:44:02
Once you start working into the XG models, individual player qualities, you start to get to
00:44:07
what we call over-parameterization and selection biases, and it's very hard. But great question.
00:44:12
So thank you. And now let's go one more student. At least one more. Hi, I'm Rachel.
00:44:18
I'm also from New York. I have a very similarly formatted question, which is about how would you do expected tackles
00:44:25
for, like, a defensive player? Expected tackles. Okay, so that's not something we've worked on,
00:44:32
but it's a great potential research question. JP, maybe to build on Rachel's question, like,
00:44:39
could you basically jam in any dependent variable you want on the left-hand side?
00:44:45
Like, Jack's interested in expected assists and Rachel's interested in expected tackles.
00:44:50
I'm interested in expected yellow cards because I like violence. You know, why not jam in anything on
00:44:56
the left? I do, I do. Let's jam in anything on the left-hand side and just build the model.
00:45:01
Is there any reason why you couldn't use the same machinery and just kind of plow
00:45:06
through it? On its face, there's no reason why you couldn't use the same machinery.
00:45:12
But for something like expected tackles, you might have a difficult time deciding what actually should
00:45:18
go into that model or what even, you know, how do you actually quantify, are you
00:45:23
training based on observed tackles? Are you training based on tackles that could
00:45:26
have been made? There's some very interesting questions that come with just throwing something on the left-hand side.
00:45:31
Your Twitter feed, if you're on analytics Twitter, is full of people putting X in front
00:45:34
of things and just saying, we have expected blank. I would say it's a little bit more
00:45:39
difficult than that. It requires, you need a bit of a structure. You need a little bit of a mechanism
00:45:46
under which you can do that. And more importantly, you need good data, data that's going to actually allow you to answer
00:45:51
that question. And so if you did have such data, like data, like the granularity that we have
00:45:57
at Wharton, it's very possible that you could build an expected tackles model and something like
00:46:02
residualizing over that. But also again, you get into the granularity of the data and how, you know, how
00:46:07
precise do you need your video estimate of location to be in order to make that estimate really reliable.
00:46:12
And so you run into some issues with reliability, stability of your estimates, but in theory,
00:46:17
right on its face, you can build an expected anything model, just how good that model
00:46:22
is going to be, how stable it's going to be. And do you have enough data for it?
00:46:25
Do you have the right data for it? Those are really the questions that... Yeah. And would your X tackles model be the
00:46:31
same as someone else's X tackled model? Because whose definition? And obviously this is something that matters a
00:46:36
lot. War is something that we talk about in lots of sports. There's war, which is wins above replacement.
00:46:44
But my war is different from someone else's war, most prominently in baseball. You know, Eric, I can't avoid talking about
00:46:50
baseball at some point. You got to. I had to get into it. But someone else's... Hey, the Yankees won one in a row,
00:46:57
but go ahead. Got to get that in. There's baseball reference war, which Neil was involved
00:47:04
with when he was there. And there's Van Graaff's war, and they're really radically different.
00:47:09
And it's just your assumption about what goes in matters. And so how you define those things can
00:47:14
be radically different. But that's a great reference. Thanks, Rachel. Let me just ask a question, actually, since
00:47:18
Neil is there. Let me just ask him a more philosophical... Come on, Neil. Neil, come on over here.
00:47:24
I can't wait without having to... Sorry, Neil. I'll have to ask you a question. It's always good to see you, as always.
00:47:30
Hey, Eric. So JP and Adi have talked about feature engineering that kind of gives kind of, let's
00:47:38
call them other explainable features or knowledge-based features. But in your experience working in the field,
00:47:44
if I'm a team, do I care about any of that? Can't I just take the massive raw data,
00:47:51
put it into some embedding space, some lower dimensional reduction space, jam it into some predictive
00:47:56
model? And if it beats your model, who cares whether I can explain it or understand it?
00:48:02
I'm in the prediction business. I'm not in the understanding business. I'm not in the causal inference business.
00:48:07
I'm not in any of those businesses. I just wanted to get your philosophy on that, because there are a lot of listeners
00:48:11
out there that are saying, I can just take this massive data, dimension reduce it in
00:48:16
some way, shove it in on the right -hand side of an equation, get some prediction,
00:48:20
and my prediction will beat yours. Yeah, that's kind of the Kaggle competition approach
00:48:25
to do it. But I think one of the things that's missing there, especially if you're working for a
00:48:29
team, is you have to be able to explain because you need them to buy in. You're dealing with a lot of people that
00:48:35
may be, they might be analytically skeptical, or they just may be curious about it, but
00:48:40
they don't know all of it. And if you come in guns blazing and say, my model, which is a total black
00:48:45
box, I can't tell you why it gives this result, but it does give this result. And it says you need to do this.
00:48:51
You're not going to get someone who's like a coach or a player to buy into that at all, because they're going to be
00:48:57
like, I don't trust you. I don't know what went into this. Do you even know why it's saying this?
00:49:01
If you can't explain it, then why am I going to put my reputation on the line by implementing this thing that you say,
00:49:07
trust me, bro, it's going to work. So that's the importance and why you can't just say the numbers say this, and that's
00:49:16
the end of the story. You have to be able to explain it. And there has to be a debate too,
00:49:20
I think, because there's a lot that the numbers are still missing. And you guys just got into that around
00:49:25
like finishing quality. Is that a real thing versus not XG and XG plus make different assumptions about that.
00:49:31
So I think that there's like elements of skill that are still left to be settled
00:49:36
in these things. All right. We have time for one more student question. We have two more lined up, but we'll
00:49:42
definitely take one. Maybe we'll do two. So let's do one. So we'll do quick, quickly say your name
00:49:48
or introduce and we'll do a quick, quick one minute answer. I'm Eli. I'm from Brooklyn, New York.
00:49:55
And my question was, how do you think pro teams are currently using this data? And how do you think they should be
00:50:01
using it? Soccer data or? All right. That's a great question. We can come down. In fact, why don't you sit down JP
00:50:09
and we'll try to answer that. I'm going to start by saying it's very hard to find out.
00:50:14
Soccer is one of the most tight lipped among the analytics community. There isn't a public facing analytics world like
00:50:23
you'll see in certainly in baseball, football, even in basketball. Basketball is also a little tight lipped because
00:50:28
it's tracking database. So we don't really know. We've had members of some of the American
00:50:35
professional leagues have come to talk to us and they're pretty tight lipped about how they
00:50:39
use it. I can only say they're using it. And we could also see a little bit on the field.
00:50:44
Do we have any to add JP? Yeah, absolutely. Analytics are notoriously, soccer has been notoriously slow
00:50:52
relative to baseball, relative to some of the other American sports. It's growing fairly quickly now.
00:50:57
You'll see job openings for most of the major clubs in the analytics and machine learning
00:51:01
engineer, data scientist space. But a lot of these teams are very small. And so a lot of them rely on
00:51:09
some of this public work that we're doing. And so I can't speak to necessarily what
00:51:14
X team is doing or what Y team is doing. But in general, I would say that the public space still has a big effect on
00:51:21
what a lot of clubs do, especially clubs that don't have as big of an analytics team.
00:51:26
They have to adopt more from the public space. So they're definitely some of our public work.
00:51:31
Would I be surprised if it pops up, a particular team is using this? No, I wouldn't be surprised in that at
00:51:37
all. But I think what you're saying is true, that the teams are looking towards what you
00:51:43
guys are doing. And I hope that's true. It makes me sleep better at night knowing
00:51:48
that the research enterprise that Adi has built around sports analytics, I hope it's actually being
00:51:53
used by teams. That would be nice. Maybe one more question. One more question. We have our last question.
00:51:59
It's a Chicago Cubs t-shirt. Let's make sure that that's seen. We got it. We see the Cubs.
00:52:08
My name is Henry. I'm from Chapel Hill. My question is for JP about how you see the use in the future of XG
00:52:13
plus to evaluate players like trade value and transfer value. When rubber meets the road is the expression.
00:52:24
So how does it actually going to end up in evaluating players and maybe potentially change
00:52:28
how much you pay for them? Absolutely. The cool thing about XG plus is that decomposition we talked about, right?
00:52:33
The expected goals portion and the expected shots portion, those can be viewed in conjunction.
00:52:37
So you could look at a player's XG plus, or you could look at their expected goals or just their expected shots.
00:52:42
And we understand that each of those maybe more or less correlated with huge performance, right?
00:52:47
They're over or under performance on those metrics. And so teams are going to be able
00:52:51
to say, all right, we may have a player who ran it on a very hot finishing streak.
00:52:56
We may not believe that, right? Everything we know about expected goals data will
00:53:00
tell us that there are very few players that consistently over perform their expected goals.
00:53:04
But this other player that I'm interested in has a very high expected shots or a
00:53:09
very high over performance of their expected shots. We might believe that a little more.
00:53:13
We might believe that might carry over. That might be more skill-based, right? There might be more signal in that number
00:53:18
than in their expected goals number or their over performance over expected goals.
00:53:23
And so I think it may have a similar effect to certain things in the NFL or in the MLB.
00:53:29
You may believe a certain skillset is more likely to continue in the future. There are plenty of examples of that.
00:53:35
And that's where I think it may have an effect is breaking down, first of all, creating a decomposition, breaking down the shooting and
00:53:42
the scoring into two distinct components that are roughly orthogonal, but both very informative.
00:53:49
And then first of all, that decomposition and secondarily being able to say, how much do
00:53:54
we this? How much do we actually believe a player is worth? How much do we believe they're actually going
00:53:58
to produce in the future? It may change those estimates and improve our ability to estimate what a player may actually
00:54:04
be worth on the transfer market. And so to that end, that's where I think teams might implement what we're doing for
00:54:10
transfers, for selecting players and scouting players in general. So the last thing I would add is
00:54:15
that to summarize, that last point would prevent you from overpaying for performance that happened in
00:54:20
the past. So you don't want to pay for performance that happened because that's not what you're buying.
00:54:25
You're buying performance in the future and allow you to acquire players who you expect to
00:54:31
overperform their market value. That sounds very familiar. Buying low, selling high, market inefficiencies, using statistics.
00:54:37
That sounds to me like moneyball. I agree with that. And then the last thing I would say
00:54:42
is that it also leads to the question of the optimal combination of players. Because having players that create opportunities on a
00:54:49
team with good finishers, it could be even more advantageous, which is why I like tree
00:54:55
-based models for lots of reasons to do these types of evaluations. Well, Adi, this has been another show.
00:55:00
It's been a very unique show in all of our 12 years. I don't know, is this maybe the second
00:55:05
time I think only in 12 years we've done something live with a classroom? Because remember, we've done live in a studio
00:55:11
audience before. Yes, we have. But I don't know if a classroom. So lots of people to thank today.
00:55:15
Let me start with the most important person. I want to thank JP. This, to me, is what research is about.
00:55:22
And so I want to thank you for joining us on Wharton Moneyball. Of course, I want to thank, even though
00:55:26
he's only on the air briefly, our longtime guest, Neil Payne. Adds one more time to his long history
00:55:32
with us here on Wharton Moneyball. And of course, on behalf of myself, Eric Bradlow, my co-host for today, Adi Weiner,
00:55:38
some combination of us, Kate Massey, Shane Jensen, are here on the Wharton Podcast Network.
00:55:43
I'd like to thank our producer for today, Jacob Grodnick in Broad Media. And of course, the big boss woman, Dee
00:55:48
Patel. Between now and next week, enjoy your sports, enjoy your statistics. We'll see you next week here on the
00:55:54
Wharton Podcast Network.

Episode Highlights

  • Wharton Moneyball Academy Launch
    The Wharton Moneyball Academy kicks off its second day with excitement and new technology.
    “It's our second day.”
    @ 01m 06s
    July 08, 2026
  • The Importance of Machine Learning Skills
    JP discusses the ongoing relevance of machine learning skills in the age of AI.
    “That's a great question.”
    @ 13m 00s
    July 08, 2026
  • Understanding Selective Inference
    JP explains selective inference and its implications in data analysis.
    “Selective inference is a field that sort of concerns itself with making valid inferences.”
    @ 15m 58s
    July 08, 2026
  • Understanding Expected Goals (XG)
    Expected goals models estimate the likelihood of a shot being scored based on various factors.
    “A shot that is a 0.5 XG shot would mean it's about 50% to go in.”
    @ 19m 54s
    July 08, 2026
  • Introducing XG Plus
    XG Plus addresses the limitations of XG by also considering the probability of a shot being taken.
    “XG plus is not just modeling the probability that there's a goal scored on a shot.”
    @ 24m 00s
    July 08, 2026
  • The Limitations of XG
    XG only counts shots taken, missing opportunities that could have been dangerous.
    “Expected goals only tabulates once a shot has been observed.”
    @ 24m 41s
    July 08, 2026
  • Player Analysis with XG Plus
    XG Plus provides a more stable metric for evaluating player performance over time.
    “The correlation year to year in overperformance to XG is rather low.”
    @ 34m 45s
    July 08, 2026
  • The Importance of Correlation in Sports Analytics
    Understanding the correlation between expected shots and actual performance can guide team strategies.
    “The correlation structure in the data does indicate a much stronger correlation...”
    @ 38m 43s
    July 08, 2026
  • Innovative Research in Sports Analytics
    Students at the Wharton Moneyball Academy engage in groundbreaking sports research.
    “It was a tremendous opportunity for undergrad students here at Penn to get involved in sports research.”
    @ 39m 51s
    July 08, 2026
  • Future of Player Evaluation with XG+
    XG+ could revolutionize how teams evaluate player trade and transfer values.
    “It may change those estimates and improve our ability to estimate what a player may actually be worth.”
    @ 54m 00s
    July 08, 2026
  • A Unique Show After 12 Years
    This episode stands out as one of the most unique in 12 years of broadcasting.
    “It's been a very unique show in all of our 12 years.”
    @ 55m 00s
    July 08, 2026
  • Thanking the Team
    A heartfelt thank you to the team and contributors of the show.
    “Let me start with the most important person. I want to thank JP.”
    @ 55m 17s
    July 08, 2026

Episode Quotes

  • That's shocking to me.
    Rethinking How We Measure Soccer Performance
  • That's a great question.
    Rethinking How We Measure Soccer Performance
  • If I had an infinite amount of data, I would not need a model.
    Rethinking How We Measure Soccer Performance
  • XG plus is more stable, I think is a strong case for its validity.
    Rethinking How We Measure Soccer Performance
  • Your question, Jack, is phenomenal.
    Rethinking How We Measure Soccer Performance
  • Buying low, selling high, market inefficiencies, using statistics. That sounds to me like moneyball.
    Rethinking How We Measure Soccer Performance

Key Moments

  • Introduction of Guests00:09
  • Model-Based Statistics19:07
  • Estimating Probability19:12
  • Quality of a Shot19:50
  • Post-Match Analysis20:55
  • XG Plus Explained24:00
  • Player Performance33:45
  • Show Wrap-Up55:53

Tension Over Time

Words per Minute Over Time

Vibes Breakdown