Search Captions & Ask AI

AI in Sports – Wharton Professors Adi Wyner and Cade Massey | AI in Focus Series

November 10, 2023 / 27:27

This episode features discussions on artificial intelligence, statistics, and their applications in sports analytics with guests Kade Massie and AI Wier from Wharton.

Kade Massie, a practice professor in operations, information, and decisions, and AI Wier, a professor of statistics and data science, share insights on the differences between AI and statistics. They explain how AI emerged from the need to analyze large datasets rather than reconstruct human thought processes.

The conversation also touches on the challenges organizations face in trusting AI models for decision-making, highlighting the concept of algorithm aversion. Massie and Wier discuss how human decision-makers often hold models to a higher standard than their own judgments.

They further explore the current state of sports analytics, emphasizing the blend of statistics and AI in player evaluations and decision-making processes. The guests share their thoughts on the future of AI in sports, particularly in injury prediction and player evaluation.

Listeners will gain a better understanding of how AI and statistics are transforming sports analytics and the importance of integrating human expertise with data-driven models.

TLDR

Kade Massie and AI Wier discuss AI, statistics, and their impact on sports analytics, emphasizing trust in models and future developments.

Episode

27:27
00:00:00
welcome welcome to the next episode of the analytics at Wharton AI at Wharton podcast series on artificial
00:00:08
intelligence while I've enjoyed all the episodes today's really a special one
00:00:12
for me not just because of my love of sports but of course cuz I'm interviewing two of my colleagues and
00:00:17
co-hosts of another show here on SiriusXM Wharton Moneyball so it's my honor to introduce my colleague Kade
00:00:23
Massie Kate is a practice professor in our operations information and decisions Department he teaches researchers and
00:00:29
consult on how to improve decision-making in organizations and probably the part you could not have
00:00:34
written it better than this thank you Kate especially by blending models and experts he's also the faculty director
00:00:40
of Wharton people lab faculty co-director of the Wharton Sports analytics and business initiative and of
00:00:45
course as I mentioned co-host of Wharton Moneyball so Kade Welcome to our podcast
00:00:50
thanks Eric delighted to be here I'm also joined by my longtime friend and colleague in the statistics and data
00:00:55
science department AI weiner I really think it's important to read ai's uh bio
00:01:00
here a short bio because of all the different things he's doing around data science statistics artificial
00:01:05
intelligence at the school ai's a professor of statistics and data science whose research has spanned many areas
00:01:11
including applied probability modeling information Theory and machine learning he's written many articles in
00:01:16
statistical methodology and also applications including other episodes like we've had just now Neuroscience
00:01:22
medicine climate science and of course extensively in sports analytics we're
00:01:27
along with Cade he's the co-director of our Sports Analytics business initiative
00:01:30
he also runs the Penn Sports research seminar the summertime Moneyball Academy and he's also one of my co-hosts of
00:01:37
Wharton Moneyball and he's the director of the undergraduate program in statistics and data science so AI
00:01:41
welcome to the podcast yes thank you Eric great to be here and that completes our 28 minutes of talking today but but
00:01:47
let's get serious so AI why don't we start with the following and this is a
00:01:50
question I get all the time and I imagine you do too what is AI and how does it differ if any with what you and
00:01:59
I might call statistics well it is actually a great great blending of the two I mean statistics have been around a
00:02:05
long time um if we go back historically AI was designed to try to figure out how
00:02:10
people think and build models that do that and then sometime in the 90s people realized that the smart way to do
00:02:17
artificial intelligence is not try to reconstruct the thought process but just get tons of data and this was a huge
00:02:24
breakthrough that I think it really happened in the '90s in machine translation and in machine speech
00:02:28
recognition and then eventually vision and then eventually vision and so what happened was is in instead of trying to
00:02:34
actually build machine learning and artificial intelligence people discovered that the right way to do that
00:02:40
was statistics now fast forward to today we still have statistics and we have machine learning and Ai and they are
00:02:45
still somewhat similar and also different so I have a rule of thumb about what kinds of problems are machine
00:02:51
learning and what kinds of problems are statistics although you have to remember
00:02:55
there's an enormous um feedback and and and intersection between the two and we
00:02:59
all use each other's methods and the way I like to think about it is this what is
00:03:04
the difficulty in solving the problem and if the difficulty is complexity so for example Vision self-driving cars
00:03:12
image recognition um trying to take the entirety of the web and use that to answer a question like you think of
00:03:18
large language models that's AI it's something that a human can do and now
00:03:23
you trying to get the machine to do it statistics is all about what I would think of as noisy problems we have a
00:03:30
colleague in in criminology who tries to predict whether or not someone commits a
00:03:33
crime after they've been in jail for a while recidivism well you're never going
00:03:38
to get a perfect prediction Sports is a classic place where a lot of our problems are statistics for example
00:03:45
trying to figure out you know what your batting average is going to be next year
00:03:48
or things like that but also um but other problems come up in both areas which are common to both where where
00:03:55
there is a lot of signal um but also um and and but you can handle them with statistical methods so the way I
00:04:02
generally describe it is what we call signal versus the noise and the Machine learning problems have high signal low
00:04:08
noise and the statistics problem have high noise and lower signal so Kate let me ask you related to that so um whether
00:04:15
it's let's maybe one could argue that uh statistics has gotten its footing in
00:04:21
companies and decision-making maybe you could argue machine learning has because
00:04:24
people like predict things that can help predict the future uh AI is kind of new
00:04:30
how do you see companies reacting to these different methods and are we now in such I'll call it an enlightened
00:04:37
explosion world like everyone's just saying of course I have to adopt it or is that like not the reality of
00:04:43
today well certainly organizations are are looking at opportunities for it and they're excited about it and and
00:04:49
certainly vendors who think they have a a new application for it or selling the potential of that but ultimately it
00:04:57
always comes back to is somebody going to to rely on that algorithm or model for a decision they have to use the D
00:05:05
thing it's it's one thing for it to spin and do pretty things it's another for a
00:05:09
human to depend on it and what we've seen reliably is that that's a pretty
00:05:15
steep hill to climb that people are loathed to trust decisions with to to models when models aren't perfect and in
00:05:24
many many applications models are inevitably imperfect as soon as we see them be imperfect we're an to lean on
00:05:30
them even if we know humans are imperfect there's an asymmetry between the penalties humans apply to models
00:05:37
that are imperfect versus humans that are imperfect as soon as they see that imperfect performance they're reluctant
00:05:43
they're they'd rather lean on a human for you know if I'm going to use an
00:05:47
algorithm and kind of offload some of it that has to be much more accurate than in some sense we as fallible
00:05:55
humans that's right that's what that's what we've observed we we along with
00:05:59
some colleagues at at at pen we've we've called that algorithm aversion and again
00:06:04
it's not to say that people just innately don't trust models we're happy
00:06:08
to to to play with models to lean on models but if we see them performing perfectly we're much harsher in our
00:06:14
treatment of them than we are humans so AI I know you wanted to jump in here and
00:06:18
we'll let's get to you before we jump into the main topic of today which is of
00:06:21
course our all our passion which is AI and sports but please you had a followup so I guess my followup to you Kate is
00:06:27
that I think um we might want to privilege human versus a model when they both make mistakes at approximately the
00:06:32
same rate but how much is that disadvantage do we really trust we really kill models even if they're
00:06:37
better than humans just because they make mistakes yeah so I can't speak to the
00:06:42
exact threshold where you might go back to the model but where we've run experiments we've manipulated exactly
00:06:48
that and and and people can observe that the model outperforms but in let but they hold models to a higher standard
00:06:54
there seems to be a model of a standard of perfection for models that they don't
00:06:57
use when it comes to humans well AI let's jump in first with you with kind of the application of AI in sports
00:07:05
analytics so what do you see as you mentioned kind of Statistics being the field of I let me just see if I got this
00:07:12
right clearly AI is you're going to have high low noise massive data statistics
00:07:18
are going to have typically smaller data sets High noise which means you have to
00:07:21
rely more on you know mathematical models for decisionmaking how do you see Sports analytics today how much of it do
00:07:27
you see the use of statistics how much do you see the use of AI is there some combination what do you see happening
00:07:33
right now okay so certainly there's an incredible amount of statistics and the
00:07:37
reason why statistics is being used so much today primarily as people see the value
00:07:42
in it I think that's been a giant sea change over the years brought about by
00:07:46
the Moneyball Revolution Oakland a etc etc people realized that using data to make decisions is a great thing to do
00:07:54
and that data wasn't Advanced I mean Bill James really revolutionized Sports
00:07:57
analytics with just counting and percentages and clever ways of adding and subtracting things not fancy so
00:08:03
statistics has certainly um become essential to the operation of successful sports teams and in lots of ways player
00:08:10
evaluations none of this is particularly complex um there are of course machine learning advances that have been made
00:08:17
that have been useful certainly the tracking data and that's for decision makings on something like we argue about
00:08:23
on our radio show all the time should we have a robotic umpire we certainly have
00:08:27
the Hawkeye data in baseball we have all this backing data from Sports vision and
00:08:30
sports View and that tracks where the ball is and where the players are and that leads to kind of mixture models um
00:08:37
statistics mixed with with sort of machine learning which is these massively giant tracking data terabytes
00:08:43
of of information that need to be processed so you can still build statistical models so there still high
00:08:49
noise in the sense that you can't predict whether or not we're not able yet to predict whether or not say a
00:08:54
running back is going to break through and get a large gain um that's still high noise but we have so much
00:09:01
information so we have to use kind of machine likee models to to fit them so that's what we're happening I'd still
00:09:05
see say that we're probably 90% in the statistics domain although a lot of the
00:09:09
flashy stuff is coming out of the the machine learning and the AI the things on TV that where you everything is
00:09:15
labeled and you see all these great the these great uh predictions they'll tell
00:09:19
you things like what's the probability of a catch um that's kind of a mixture
00:09:22
of a statistics and a and a and an ml or a machine learning model so Kade one of
00:09:26
the things you even talked about in your own bio was the blend of models and experts how do you see that blend
00:09:32
happening in the field let's talk about sports since I know you do work with a
00:09:35
lot of sports teams how do you see that blending of machine learning Ai and human experts today what are some good
00:09:43
examples well it's sadly it's largely uncomfortable that that that historically these come from very
00:09:49
different communities and they don't necessarily play well together and so the organizations that actually are
00:09:56
blending them most successfully are are are in some way forcing the groups to work together because you know
00:10:03
traditional decision maker in NFL say isn't super interested in dialing in the
00:10:08
computer science guy who just graduated from MIT that's not a that's not the way
00:10:12
they usually make decisions the organizations that get it well have leadership to say we are going to do
00:10:16
this they don't make it one person's model or another it's more the team's
00:10:21
model the sports that are furthest along especially on the Personnel side are baseball and baseball they they do have
00:10:28
models that are like the teams model they they lean on them heavily when they evaluate Personnel they've built those
00:10:34
things up over years these things are not something someone writes down one time and they're off and running they
00:10:40
tend to be highly iterative the modelers get lots of input from the traditional decision makers they get hypotheses from
00:10:47
the decision decision makers the decision makers point out things the models are missing these things are they
00:10:52
should be iterative and a dialogue between the two sides that's tough to pull off and so it's relative L rare So
00:11:00
kid let me just follow up with a question on that you mentioned the idea of kind of using whether it's humans or
00:11:05
Theory to come up with hypotheses that are tested just you as a scientist how important is that or can't you know
00:11:12
can't I just use Ai and machine learning algorithms to just find patterns in the
00:11:18
data and that's what it is like why do I I'm saying this in a factious way
00:11:21
because I have an answer but this is about you guys not me why not why need Theory why not just explore the data and
00:11:26
see what we find so ai's going to have a deeper answer than I am for this because
00:11:31
it's very much in line with what he's been talking about but but but sadly
00:11:35
sadly for some of us in many domains of sports we just don't have the data to
00:11:39
support that kind of exploration so we have to bring more structure to the conversation we can't just set the
00:11:45
algorithm loose and see what it tells us there are certain places where that can
00:11:49
happen but by and large we just don't have enough data for an algorithm to learn reliably so
00:11:57
for example I'm I'm off often worrying about personnel and we just don't see
00:12:01
enough people there's a lot of variables that you could lean on we don't have
00:12:05
that many observations we need to bring a little structure we need that human hypothesis in order to make traction
00:12:11
that's the way I would go about it I'm curious how OD he's gonna talk about it
00:12:14
well I'll start with just simply agreeing I mean we think we have a lot of data because we have lots and lots
00:12:19
and lots of observations but it's not nearly as many as you think you'd want
00:12:22
because you have so many variables as well and a lot of times the data is not new data it's just more and more data of
00:12:29
more or less the same thing and what you need in statistics to do good forecasting and good model building is
00:12:34
lots of independent data we really don't have that so you think about the incredible amounts of observations you
00:12:39
can take using tracking data it's terabytes in a game that's but that's
00:12:43
still one game it's still only a couple hundred plays and only one Victory right
00:12:47
so all of this is is highly correlated which means we just don't have as much
00:12:51
data as you think so I'll start with that the second thing is when you're
00:12:55
dealing with noisy models and these are player performances are fundamentally noisy meaning people you cannot look at
00:13:02
a player in the minor leagues or or watch them swing or pitch or or run and predict what they're going to do on the
00:13:07
field it's just not possible they're human beings they're not they're not
00:13:10
robots which means that it the and this is a technical word overfitting is easy to do and if you throw a machine
00:13:17
learning AI model at a high noise not so many independent observations you are going to get it wrong this is a
00:13:26
fundamental area of research and this is what we're doing and what I'm working on
00:13:29
with my grad students and my students to figure out how to use Ai and statistics
00:13:34
in this exact setting to build good models and I'm just it's just not trivial you just can't throw it in and
00:13:40
crank it as if it's uh uh just automatically run for you so we're here on the AI at Wharton and analytics at
00:13:46
Wharton podcast series we're on our episode on AI and sports this is Eric bradow professor of marketing statistics
00:13:51
and data science here at the arton school vice dean of analytics and I'm here with my friends and colleagues Kade
00:13:56
Massie practice professor in the oid department and AI weiner who's professor
00:14:01
of statistics and data science so Kade let me ask you um I would think that a lot of the work that maybe firms are
00:14:10
doing today is about collecting these novel data sets whether it's you know we
00:14:15
have an episode on Neuroscience so maybe it's brain data or maybe it's motion
00:14:20
tracking data that AI said or maybe it's eye tracking data how much do firms
00:14:24
recognize that in some sense we don't have the data we need and that we need
00:14:29
to kind of do you know whether it's putting sensors on players or measuring sleep or rest how much do firms realize
00:14:35
that in some sense it's better data that's going to solve these problems not
00:14:39
necessarily more correlated data I think there's pretty healthy appreciation of
00:14:43
that now of course there's variation the teams teams vary in their quality of
00:14:48
ownership quality of management and not everyone sees the opportunity there but increasingly teams do see the
00:14:53
opportunity and of course some sports are again farther ahead on this than others I mean sport like baseball they
00:15:00
are way down the path and they because they're so far down the path they are
00:15:04
getting quite inventive in the in the kinds of data they're looking for in search of an edge sports like football
00:15:10
and hockey there's much more variance but the sophisticated teams in those sports are in fact being creative and
00:15:18
one way to think about it we I think we tend to think that there's one big model
00:15:22
sitting inside the firm somewhere that's that's spitting out answers whether it's
00:15:25
play calls or who to draft but in fact these organizations have a bunch of little models and they're they're mostly
00:15:32
going out and solving one problem getting one Insight adding one you know measure to a player's evaluation using
00:15:39
different little models whatever they can get their hands on it's much more the the the propagation of these smaller
00:15:46
models and eventually they'll they'll be integrated but right now it's the
00:15:49
propagation of smaller models All In Search of small edges yeah the one thing we hope at least I don't know maybe you
00:15:54
guys would disagree with this I'd love your thoughts um there can be different
00:15:57
models for different purpos as a matter of fact there should be probably um but there should be a common data set so
00:16:03
could you talk about I know you work with a lot of students how do you guys let's say you want to solve a problem
00:16:08
like for example I know we're going to be doing a high school sports competition and you're going to doing
00:16:12
something around soccer how like how do you collect a data set like suppose you have like one data set which is motion
00:16:19
tracking and you have another data set which might be training and another data set like how do how do you even think
00:16:24
about integrating these disparate data sets together when you're trying to solve a problem so interesting uh I
00:16:31
actually would call that data science so the idea that a now wait we've got statistics but and now we've got machine
00:16:37
learning we've got AI now you want to add data science yeah so we our department has been the statistics
00:16:42
department for many years that's that was my appointment and all of a sudden I
00:16:44
got an extra an extra title on statistics and data science and uh people still wonder what what is it and
00:16:51
what exactly does how is it different from statistics and obviously these things overlap and I would say that the
00:16:56
new um the the new direction is that because we have so much data it's so large and it needs to be integrated and
00:17:03
managed and um curated if you will uh wrangled sometimes people is the word they use that's that task is data
00:17:10
science and that pushes Us in the direction of Cs computer science engineering and that's a hard task um I
00:17:18
personally don't work in it our my colleagues in engineering and Cs and some of my colleagues and statistics do
00:17:23
more of that it's a challenge and I would say that teams invest heavily in the kinds of personnel who can do
00:17:31
those things and it's a big it's a big Direction and it's expensive because
00:17:35
they're highly in demand in lots of areas so having people who can who can collect data integrate them make
00:17:42
dashboards software this is not modeling this is not statistics it's not even
00:17:47
machine learning or AI it's just the support structure expensive and and needs to be done we don't even have a
00:17:53
lot of the the basic data set so one of the things that we in the in the University land we have to deal with
00:17:59
only what's public and sometimes that's hard and there's often very good data
00:18:03
with the individual teams and some consulting firms have it and you can of course buy it but a lot of that data set
00:18:09
is data cost is outside of our our budgets um and so this is always a challenge and we're we're working on it
00:18:16
so Kate I know you have some thoughts on this as well I just wanted to emphasize AI got
00:18:21
there eventually but I wanted to emphasize that there's the backend side of that and there is the front end side
00:18:26
as well and as analy list we tend to lose sight of both of those things so the computer science part of integrating
00:18:32
these data sets building the the plumbing and the infrastructure is vital you can't do anything and often the
00:18:37
first hire in this space for an organization going down this road is on the Cs side and then Audi eventually
00:18:43
talked about like the dashboard the front end these organizations all have these they do have these database
00:18:49
systems that people drop reports into and people pull reports out of and it is the single primary way that that the non
00:18:58
analyst interact with the data so it's that front end that dashboard and related issues on the front end is vital
00:19:05
and that's development I mean this is again this is not this is not the traditional data science that we think
00:19:10
about and yet it's vital to the way data are used in the organizations so AI let
00:19:15
me start with you um when you think about the most sophisticated use or interesting use for you of
00:19:23
AI now I guess I have to say AI machine learning statistics or data science today
00:19:28
what are what is it for you what's the one that you find the most interesting
00:19:33
and if you want you could rely on what's interesting to you is what you and your
00:19:36
students are doing research on what's most interesting to you that's going on
00:19:39
today and then Kate I'd like to ask you if you know after that you know if we're
00:19:43
sitting here in 10 years what do you see us talking about well we'll still be on
00:19:47
our Wharton Moneyball show in 10 years and we'll just talk about it there but
00:19:50
AI what what's the most sophisticated application you see today well there's a
00:19:54
lot of really sophisticated ones that I'm not sure have really borne fruit yet
00:19:58
so so there have been a bunch of competitions in football that have released tracking data and some of that
00:20:04
information has been used and Incorporated by teams by by um ESPN or or the NFL in what they call Next
00:20:11
Generation stats to provide you the kind of information you couldn't have dreamed
00:20:14
of you know years ago so one example would be um you see a running back has handed a ball and you want to know well
00:20:21
on average what would an average running back do in that situation and then you can compare what the actual running back
00:20:26
in that situation did and you and you post that that Delta that differential and that's an example of either poor
00:20:32
performance or great performance but you wonder whether or not that model was really high quality and and and some of
00:20:38
our students have had real have looked at this and and and and it's been interesting but I don't think we're
00:20:42
there yet to see that these things have really provided value because the data is is hard to make good sense out of it
00:20:49
so that's one kind of problem um the problem that I've been working with with
00:20:52
our our students is actually on uncertainty so a model produces a a forecast all right well let me get right
00:20:57
to it then I was going we we'll probably talk about this on our radio show tomorrow but let's talk about
00:21:01
uncertainty for a second the New York Giants are fourth in one from the 17 yard line in the game yesterday and the
00:21:08
question is should they kick the field goal forget that they missed it there has to be a lot of people like you
00:21:13
always go fourth in one and you've shown that's not true so could you talk about
00:21:17
briefly the role of uncertainty in these models and that most people think it's a
00:21:22
yes no decision when it's really not okay so there is a couple of issues here
00:21:27
so when you build a model the model will give you what we call a point estimate and the point estimate will tell you for
00:21:33
example the probability of winning the game if you go for it and the probability of winning the game if you
00:21:38
don't go for it and then you can easily just pick the one that's higher now the
00:21:43
problem with that is we don't really know the probability we've had to estimate it using a model and there is a
00:21:48
whole bunch of actual probabilities that are all kind of equally supported by the
00:21:53
data and what we want to do is take all that information and say wait a minute if I had gotten a different set of games
00:21:59
different Universe of of of information and tried to build the same model with that would I've come up with the same
00:22:06
decision and if every single time I do that it's always go for it then you can
00:22:10
be clear that that's a a firm decision if on the other hand only 55% of the historical somewhat what we call
00:22:17
simulations the word actually is called bootstraps of the of the his historical data produce one result then you have to
00:22:23
throw up your hands and say I don't really know and you have to communicate to that to the coaches on on the field
00:22:28
to say you know we don't really know you decide so Kade let me ask you since I
00:22:33
believe in some of your research you've covered the topic and then we'll get to
00:22:36
the the future in 10 years we you've covered the topic or you've done some
00:22:39
research on you talked about algorithm aversion but even how about risk aversion so how does uncertainty play
00:22:46
into you know you're asking a human decision maker to trust a model when there's massive potential uncertainty he
00:22:54
she they see the situation in the game and then decide hey you know what my read of this is the uncertainty does not
00:23:02
swamp out what I've historically done how do decision makers think about risk
00:23:06
aversion in using models well I think the primary role of risk aversion there is is to there's an asymmetry in the way
00:23:17
you you get it wrong if you get it wrong doing the conventional thing the punishment is less harsh than if you get
00:23:24
it wrong doing the unconventional thing that's that's the strongest Dynamic
00:23:27
there that affects especially coaches using these models or not in decision- making
00:23:32
on the field but what I love about what um Audie and Ryan are doing in this research is they're introducing this
00:23:39
idea to model output that we've been talking about for humans for a long time
00:23:44
people need to know what they know and what they don't know they need to know
00:23:48
when they're sure and when they're less sure and they need to be able to say and
00:23:52
own and explain as part of my recommendation this is one with high confidence this is one moderate
00:23:57
confidence or this this is one with light confidence we've talked about that
00:23:59
with humans for decades and now Audi and Ryan are saying hey we should do the same thing with our models and it's
00:24:07
vital because those guys on the field are going to be incorporating whatever comes out of the model with other
00:24:12
factors and we as analysts have to recognize there's always other factors the models rarely if ever have every
00:24:18
possible consideration so there's always other factors and so it does depend on
00:24:22
how sure the model is if the model is absolutely certain then it's going to swamp other factors but it's not always
00:24:27
is going to be absolutely certain and we haven't talked about models in these
00:24:30
terms in the in the past it's fantastic fantastic new development from these
00:24:35
guys so in the last minute or so so AI tell me uh if we were sitting here 10 years from now what have you been
00:24:42
working on and what do you think the field of AI and machine learning in sports has done over the next 10 years
00:24:48
or if we're sitting here 10 years from now over the past 10 years what are the
00:24:51
big new advances well I'm fairly certain that we're going to have unbelievable
00:24:54
Graphics we're going to have great statistics that are going to be computed on the fly I think when our broadcast
00:25:00
experience will be really different um I'm hoping that we've made progress on
00:25:04
some of the really big questions which are still amazingly unanswered which is injuries we have really no clue how to
00:25:11
predict an injury although there are lots of startups trying to offer that that that product if you will even in
00:25:16
it's infant stage to to teams because we just really have a hard time with that
00:25:20
um I think player evaluation and the complex sports will become much better really good at it in baseball there's
00:25:27
still some things to work on I'm working on them can't can't keep keep my hands
00:25:30
off of it but I think in football and in basketball and in soccer um you'll see
00:25:35
much better ways of evaluating a players building rosters and this is going to become necessary every team's going to
00:25:42
have to have it it'll be Baseline right now in some sports you don't have to
00:25:45
have anyone and you're not automatically behind ex except in baseball I think
00:25:49
it's going to be everywhere every team will have a substantial uh staff doing
00:25:53
the routine things that just need to get done um that's the extent of my imagination I'm I'm curious to know what
00:25:59
Kate has to say me too well I on the same page with you think the newest Frontier is biomechanics and the
00:26:06
greatest opportunity there is in injury reductions that's the biggest development we're gonna see the steepest
00:26:11
development will be in biome mechanics over the next five years um the one that I hope I I agree with AI that player
00:26:18
evaluation is going to get better I think in particular the reason baseball's always been ahead is that we
00:26:23
can kind of just add up the players in a linear model and get a reasonal output for the for the team we know that we
00:26:30
suspect there are interactions on basketball courts football fields soccer pitches and we don't have a great way of
00:26:36
modeling those interactions right now that means there are players that we think are doing more than they actually
00:26:42
are to to advance the team success and importantly there are players who are making very important contributions that
00:26:47
are being underappreciated right now and I'm hopeful I believe that the models
00:26:51
will get better at identifying those in the next five maybe 10 years well I'd like to thank my
00:26:58
colleagues friends Wharton Moneyball co-hosts uh Kade Massie practice Professor here at oid also the faculty
00:27:05
director of Wharton people lab faculty co-director of Wasabi and co-host of Wharton Moneyball and my colleague AI
00:27:11
Wier professor of statistics and data science and while he's running everything we're doing here in sports
00:27:16
and statistics I'd like to thank both Ai and Cade for joining me for this episode
00:27:19
on AI and sports

Episode Highlights

  • AI and Statistics: A Blending
    AI and statistics have a complex relationship, blending methods for better decision-making.
    “AI was designed to try to figure out how people think.”
    @ 02m 10s
    November 10, 2023
  • Algorithm Aversion in Decision-Making
    People are often reluctant to trust algorithms over human judgment, even when models outperform.
    “There’s an asymmetry between the penalties humans apply to models versus themselves.”
    @ 05m 34s
    November 10, 2023
  • The Importance of Data in Sports Analytics
    Sports teams are increasingly recognizing the value of quality data over quantity.
    “Increasingly teams do see the opportunity.”
    @ 14m 53s
    November 10, 2023
  • The Role of Data Science
    Data science is crucial for integrating datasets and building infrastructure in organizations.
    “This is not traditional data science; it’s vital to how data is used.”
    @ 19m 13s
    November 10, 2023
  • Uncertainty in Decision Making
    Understanding uncertainty is key in sports decision-making, especially when using models.
    “Most people think it’s a yes/no decision when it’s really not.”
    @ 21m 25s
    November 10, 2023
  • Future of AI in Sports
    Experts predict significant advancements in player evaluation and injury prediction in the next decade.
    “Every team will have a substantial staff doing the routine things that need to get done.”
    @ 25m 55s
    November 10, 2023

Episode Quotes

  • AI was designed to try to figure out how people think.
    AI in Sports – Wharton Professors Adi Wyner and Cade Massey | AI in Focus Series
  • Statistics is all about what I would think of as noisy problems.
    AI in Sports – Wharton Professors Adi Wyner and Cade Massey | AI in Focus Series
  • We just don’t have enough data for an algorithm to learn reliably.
    AI in Sports – Wharton Professors Adi Wyner and Cade Massey | AI in Focus Series
  • AI got there eventually but I wanted to emphasize the backend side.
    AI in Sports – Wharton Professors Adi Wyner and Cade Massey | AI in Focus Series
  • We don’t really know the probability; we’ve had to estimate it using a model.
    AI in Sports – Wharton Professors Adi Wyner and Cade Massey | AI in Focus Series
  • Every team’s going to have to have it; it’ll be baseline.
    AI in Sports – Wharton Professors Adi Wyner and Cade Massey | AI in Focus Series

Key Moments

  • AI vs Statistics01:50
  • Algorithm Aversion06:04
  • Integration of Data Sets16:33
  • Data Science Challenges17:28
  • Importance of Infrastructure18:31
  • Uncertainty in Models20:59
  • Future Predictions24:40
  • Biomechanics Breakthroughs26:04

Tension Over Time

Words per Minute Over Time

Vibes Breakdown