Search Captions & Ask AI

Diversity Is Critical for the Future of AI — Leading Diversity at Work

November 09, 2023 / 46:36

This episode discusses responsible and fair AI in the workplace and marketplace, featuring guests Dr. Broderick Turner and Dr. Kareem Gan.

Dr. Broderick Turner, an assistant professor at Virginia Tech, shares his journey into AI ethics, emphasizing the importance of representation in data and technology development. He highlights his work at the Technology Race and Prejudice Lab, focusing on how race and racism impact consumer and managerial decisions.

Dr. Kareem Gan, founder of RA AI Audit, discusses his experience in AI governance and the need for companies to invest in responsible AI practices. He addresses issues of bias in AI systems, particularly how they can misrepresent various demographic groups and the implications for industries like healthcare and criminal justice.

The conversation also covers the challenges of transparency in AI, the hallucination problem in language models, and the importance of diverse teams in technology development. Both guests stress the need for companies to engage with stakeholders and prioritize ethical considerations in AI.

Listeners are encouraged to understand the implications of AI in their lives and advocate for responsible practices in technology development.

TLDR

Experts discuss the importance of responsible AI, representation, and transparency in technology's impact on society.

Episode

46:36
00:00:01
this podcast is brought to you by knowled of [Music] Warton hello my name is Stephanie Cy and
00:00:18
I'm an assistant professor of management at the Wharton School of the University
00:00:21
of Pennsylvania and I'm delighted to welcome you to today's episode of the
00:00:25
knowledge at War and leading diversity at work podcast series which is focus focused on responsible and fair air AI
00:00:32
in the workplace and the marketplace joining me today are two very special guests first we have Dr broer Turner who
00:00:41
is an assistant professor of marketing at Virginia Tech Pamplin School of Business and a visiting fellow at
00:00:47
Harvard Business school's Institute for business and Global Society he also runs
00:00:52
the technology race and Prejudice lab or the Trap lab where he and his fellow Trappers are pushing the boundaries on
00:01:00
understanding how race and racism underly many consumer and managerial decisions his main research area focuses
00:01:07
on the intersection of marketing technology racism and emotion next we have Dr kareim Gan who is the founder of
00:01:16
RA AI audit an AI governance and research consultancy with 16 years of experience in ethics and governance
00:01:24
spanning both industry and Academia he recently served as the founding user experience researcher on meta's
00:01:31
responsible AI team leading AI fairness user research at the company he has helped meta our AI team scale its
00:01:40
products across the company and bolster the responsible adoption of AI Dr ganana
00:01:46
holds a PhD in management from the University of Virginia's Darden School of Business specializing in business
00:01:52
ethics and organizational behavior and I just realized we've got a Virginia connection uh across the two of you here
00:01:59
who would have who would have thought that that would happen but in any case welcome bro and Kareem I am so honored
00:02:05
to have you with me today for a conversation on responsible and fair AI in the workplace and in the marketplace
00:02:12
so we're hoping to cover you know a lot of interesting ground here today and as
00:02:16
I as talked to brck and cre about this earlier I like to think of myself as a as a novice on these topics so today I'm
00:02:24
going to represent the average uh consumer um and worker who doesn't have a lot of experience talking about um AI
00:02:32
other than what I read about and hear about um on mainstream media and from my colleagues so so we're hoping today that
00:02:39
um this conversation will be helpful to others of you who are just like me who are trying to figure out how relevant uh
00:02:46
a conversation about responsible and fair AI is to us as workers but also as consumers so let's get started talking
00:02:55
about broadly about responsible and fair Ai and so project I'm I'm going to go to
00:03:00
you first and what I'm hoping you can do is is share a bit about how and why you
00:03:05
became involved in this topic and summarize for us some of your current work related to this topic yeah so I got
00:03:13
involved in this topic area uh honestly more than a decade ago when I was teaching high school math uh at a public
00:03:22
high school in Atlanta Georgia where I taught like 90% uh black students and what I wanted
00:03:29
to do while I was there was figure out ways to make those kids' lives better
00:03:34
fairer more Equitable uh and simultaneously I used to teach linear regression and students asked me Mr turn
00:03:41
am I ever going to use this and now I have an answer the answer is now that uh you know so I've moved into this space
00:03:49
where I'm thinking a lot about how race and racism underly uh Market Systems and Technology
00:03:57
being a system that matters and our research group the technology race and Prejudice lab is doing active research
00:04:05
now on what levers can be moved to lead to uh more equality in systems to lead to better outcomes in technology uh this
00:04:16
has extended some into some uh company advisory work where uh we're trying to
00:04:21
get company consider that moving uh communities earlier into the development process leads to better uh Products that
00:04:29
come out out and then on the public knowledge side uh we're writing white papers on uh how different uh identity
00:04:37
groups may be impacted by gender of AI for instance so we just uh finished a report on for Hispanic herited Munch
00:04:45
Hispanic herit Heritage Month thank you on how uh in uh how folks with Hispanic Origins and background should be
00:04:54
thinking about gender of AI how do they fit into this interlan system of data the classification code okay great thank
00:05:03
you so much uh so Kareem same question to you can you share a bit about how and why you became involved in this topic
00:05:09
around responsible and fair Ai and summarize some of your current work on this topic yeah absolutely firstly thank
00:05:17
you so much for inviting me I'm excited about this dialogue with yourself and
00:05:21
brri and uh love the Virginia connection Virginia is well represented here today
00:05:27
um so i' I've been involved in ethics and governance um for quite some time you know prior to
00:05:32
pursuing my PhD so I'm not surprised that I ended up in kind of AI governance
00:05:38
uh what I'm a little bit surprised about is actually post pursuing my PhD going
00:05:42
back to uh industry initially when I started off my PhD program I had no intentions uh to go back to industry but
00:05:50
in my last year of my PhD which was my sixth year I had received offers basically from Academia and Industry and
00:05:58
uh I went for back to Industry so um firstly as someone who belongs to a minority group um I was very passionate
00:06:07
and still am about having a front seat to these conversations that are happening in Tech and uh helping shape
00:06:14
the uh the technology using you know the latest and greatest research but also interacting with product teams to be
00:06:22
able to do that policy and legal and so forth um the other reason that I chose to go back was
00:06:30
you know as I was thinking through my impact to be honest uh joining a company like meta's responsible AI team where
00:06:38
the company has over three three billion users I think you know it's very difficult to kind of argue with the
00:06:44
extent of impact that you can do uh when you uh you have three billion users and
00:06:49
any kind of like product changes that you make can impact really a lot of people worldwide and and lastly honestly
00:06:56
uh one of the concerns that I've had about uh Academia was you know being stuck in a small college town somewhere
00:07:05
uh where there's not very much diversity and uh that would not be appropriate for
00:07:11
the upbringing of my kids and so that was a a a very uh strong concern uh not having much of an autonomy about where I
00:07:19
live or where I work um so it drove me back to actually go to um to Industry now as far as my research
00:07:30
uh much of what I do for clients is obviously confidential but on the public side of things I've been testing out uh
00:07:36
some popular AI image generators um and I'm finding a few patterns so uh most of
00:07:42
the images uh that are being generated from uh the uh at least the AI generators uh that I have tested tend to
00:07:51
be of people from the white race uh people of Co color are uh greatly overlooked uh these models tend to
00:07:58
associate uh a professional headshot or a professional dress code or a professional hairstyle with those of uh
00:08:06
white people uh women are underrepresented uh and uh when presented as professionals they're often
00:08:14
limited to gendered occupations such as teachers nurses graphic designers and much less likely to produ to appear as
00:08:22
CEOs for example or medical doctors or lawyers um images also tend to spew on the younger side of things and Senior
00:08:31
managers for example are associated with having gray hair which is a sign of aism um and finally you know these
00:08:40
models tend to portray certain populations in demeaning ways for example uh if if you ask Del Bing to
00:08:48
portray Turks uh it will uh provide you with a picture of a Stern turkey uh a Stern turkey dressed in a turban right
00:08:58
um whereas if you're ask it to produce for example an image of an American or a
00:09:02
French person or something like that will do a much better job right um they might also refuse to produce images of
00:09:09
certain populations For No Good Reason such as if you ask it to produce um an image of a Muslim or a Jew uh while
00:09:17
presenting followers of other religions so definitely this technology is extremely powerful uh I am not a
00:09:25
doomsayer per se as it uh relating to the technology but I believe uh that you know we have to do
00:09:32
our prudent and due diligence in order for us to be able to direct the trajectory of this technology in ways um
00:09:39
such that we're maximizing the benefits while minimizing the harms my my favorite example of uh that gender of AI
00:09:47
uh art is when people do prompts for Jesus in the temple flipping the table that they literally get Jesus doing like
00:09:56
gymnastics doing a backflip over a table uh so yeah text text has no meaning it's it's
00:10:03
just funny absolutely so we're gonna definitely dive into there's a lot of to
00:10:08
unpack there and certainly the implications of this I mean I think of where I sit I can automatically connect
00:10:13
those dots but I think as we begin to talk about some of the challenges it'll
00:10:16
be important for people to understand what could possibly go wrong when you you're not represented in how these
00:10:22
models are being uh trained I I think to me that's an obvious answer but I don't
00:10:27
think that's obvious to everybody what happens when you lack representation in
00:10:32
in the data amongst the data or you're misrepresented right who you are and the
00:10:36
the the groups that you're part of are misrepresented in in how the model is
00:10:40
being trained so we're going to talk about that in a minute but let me just and probably this is part of the answer
00:10:44
to your question here Kareem um let's let's start with business and employers
00:10:49
first right and and I can imagine I have some sense of this from my own work that
00:10:54
businesses and employers um are likely and can be struggling to understand their conversation not just in a role
00:11:02
about AI use but fair and responsible AI use and so obviously without you know saying anything confidential can you
00:11:11
talk generally about how you through your work are helping businesses employers make sense of their roles and
00:11:18
responsibilities on this topic yeah I think it's important to be clear that um companies are accountable
00:11:24
for how their AI systems operate and that responsibility can be offloaded right it can be considered an
00:11:31
externality of some sort they have obligations to their customers and it's just a matter of time uh before
00:11:38
legislation comes into effect in this space you know obviously we have right now the uh EU AI act which we're waiting
00:11:46
it's just a matter of time before enforcement happens and there's a lot of
00:11:50
talk uh about legislation uh Congress and so forth here in the United States um and honestly if come companies want
00:11:59
to build successful products in this AI era era they must invest in responsible AI infrastructure right um doing so
00:12:07
requires a commitment from the board of directors from senior leadership but ALS
00:12:12
ultimately what it does is it pays off uh in terms of earning customer trust and confidence right it it just doesn't
00:12:20
cut it for you to actually have an AI strategy uh without thinking through the the trust layer or the responsible layer
00:12:29
um you know issues of like fairness and privacy and um robustness and and so forth right because if if you're a
00:12:38
product manager for example you're trying to produce the best products no one really wants to deal with a company
00:12:43
whose products are reaching you know their privacy uh you know their their privacy obligations and stuff like that
00:12:52
like if you're using personal data uh of your users to train your models or if
00:12:58
your products don't work for certain segments of the population or exposes your users to specific types of harms
00:13:04
you know cutting Corners like that really does a disservice uh to your stakeholders of course uh but it also
00:13:12
exposes you to legal and reputational uh risk and it's just a bad way of doing
00:13:19
business it's just a matter of time before you know uh better companies are able to gain more market share because
00:13:27
they're taking this more um seriously and investing in setting up the infrastructure yeah you know a lot of
00:13:34
what I hear you saying is you know sometimes it's easy to think that ethics and responsibility questions of
00:13:42
responsibility isn't the job of a corporation but you take a strong stance and certainly through your own
00:13:48
dissertation work and your own scholarship is that these are things that are not nice to haves um there must
00:13:54
haves and and you know while it's it shouldn't always come down to this we do
00:14:01
have a system called the legal system which is set in place to help adjudicate these issues when it does feel like it's
00:14:08
a great area even though a gray area excuse me even might be black and white with respect to businesses and their
00:14:14
responsibility to ensure that they're not inducing harm uh through their activities brri I mean certainly I know
00:14:21
that you do work with companies as well I I would also like you to put on your researcher and educator hat um and help
00:14:28
us to understand like the nature of this conversation from a scientific perspective and also in the work that
00:14:36
you're doing you know as a professor as part of your fellowship I'm actually
00:14:41
thinking a lot about chat GPT these days certainly because as an educator this is
00:14:46
sort of the one this is probably the most I know about this conversation around Should students be able to use
00:14:52
GPT to complete their assignments and it seems to me like it's such a polarizing
00:14:56
sets of issues um so I am just just curious as you think about the the places where you set project um how
00:15:03
you're thinking about the issue around businesses and employers and Ed education's role in this conversation so
00:15:11
I think a couple of thoughts right so first uh to piggy back on Kareem when I talk to both students and businesses I
00:15:19
think of myself as an educator no matter where I am right and I have the same thing I tell them uh which is if you get
00:15:27
this wrong right you don't include uh a wide SWA of human beings in the creation
00:15:33
of your technology products when you fail because it's not an if but when you fail you lose money
00:15:41
because you spent all this money on development without considering uh the human beings at the end and when the
00:15:47
product gets released the product fails again because people don't use it right
00:15:53
so as a really simple example uh those automatic uh faucets where you put put your hand under it did not
00:16:01
include uh like darker melanated people in the data set in the training set for when they were testing out this
00:16:08
automatic hand washer and so while it sailed sold fine in North America when it went further out into uh places
00:16:17
closer to the Equator where people were Browner regardless of uh ethnic origin they didn't sell any because the product
00:16:24
did not work they would install it and people put their hand under the thing and I know if you're a black or a brown
00:16:30
person you have done this at the airport bathroom been upset right but if your country is majority melanated then no
00:16:39
one bought them and so have they included those folks earlier in the development process like all that money
00:16:47
they spent on development uh would have been worth it because they could have sold uh more products and so when I'm
00:16:53
talking to my students and Business Leaders who in some ways are my students uh uh I go look this is why we do it
00:17:01
like I'm I'm going to speak to the same incentives that you have you want to
00:17:04
make more money you want to keep your job uh you want to do right by your shareholders then you need to move uh
00:17:12
the process or move the people earlier into the development of this process now in terms of this question around chat
00:17:20
GPT and uh the way that uh people are using this inside and outside the classroom uh we developed the triap lab
00:17:29
this these three questions call the 3D model of Equitable Tech adoption right and I'm going to share those questions
00:17:36
with you and this will think help you understand help people understand what do I do with any Tech that comes into my
00:17:42
business or into my classroom those three questions question one what does this product actually do not what do you
00:17:50
want it to do not what it might do in the future not what it could do if we had uh a billion more hours of Compu
00:17:58
like what does it actually do today second question is who does this product disempower all right every technology
00:18:07
will increase power for some and decrease power for others so ask yourself who does this disempower and
00:18:13
then the third question is what is the daily use of this product not the edge case don't tell me that you know we're
00:18:22
going to uh I give you a per example for Hispanic Heritage Month there is uh some
00:18:29
question that maybe we should include in facial recognition whether or not uh this person is Hispanic or this person
00:18:37
is uh Latino and the sales pitch for this is that this will help us when we do the census to identify who is
00:18:46
Hispanic or Latino so we have a better count right now I can get into like the research and on this and why this is
00:18:52
problematic or some of the philosophy of like how this goes arve but let's talk
00:18:56
about that case they're saying we're going to use this for the census the
00:19:01
census happens every 10 years all right there is no government service that's
00:19:06
going to buy this expensive technology and then only use it once every 10 years so then what's the most likely daily use
00:19:16
case right of a technology that uh like says this person Hispanic or not right are they going to install it at Borders
00:19:26
probably uh would they install where you hand over your uh passport probably right and is that problematic definitely
00:19:34
right and so this is how I'm thinking about these things uh and in the case of
00:19:38
chat GPT your students should ask themselves what does this actually do right is this writing papers or is it
00:19:47
just creating plausible sounding sentences right and if it's just creating plausible sounding sentences o
00:19:53
that could get real bad real fast uh if that plausibility has no relationship to
00:19:59
accuracy all right so you know feel free to use it and feel free to fail that's
00:20:04
on you uh that's come down to it so you're doing a great job both of you already of
00:20:10
raising some of these key challenges but I'd like to explore these more and certainly you know we don't have time
00:20:17
for tap tens so we'll do and I'm sure that there are 10 on each of your lists
00:20:21
but let's think about this in terms of the top two or three issues from where
00:20:25
you sit so project let's go back to you you can those that you've already talked
00:20:29
about or you know invite others into the conversation what would you say from where you sit are the top two to three
00:20:36
issues with respect to responsible and fair AI in the workplace and the marketplace uh two things uh the number
00:20:47
one issue is that the human beings that end up using the AI or being users of machine learning or being users of this
00:20:58
generative AI aren't necessarily the people that are included in the data all
00:21:04
right aren't necessarily the people that are included in the classification of
00:21:07
the data and definitely aren't the people that are setting the rules and doing the coding of the data and so the
00:21:13
most pressing uh issue is to get those people into those spaces right to get representative data to get
00:21:23
representative uh classification of that data to get representativeness in the actual coders that decide the rules for
00:21:31
what we see and then that way the stuff that actually comes out will be closer to fair and Equitable because those
00:21:38
people would be in the room so that's one the second thing is that all of my
00:21:43
research uh is around this space is really to demystify this black box there is nothing magic all right going on inside
00:21:54
of these statistical models they are statistical models and if you learned in high school y equal MX plus b then you
00:22:04
like have the building block that you need to understand how these systems work uh we can talk about this if you
00:22:12
ever come hang out on the Trap lab but trust me if you know y equals MX plus b all right then you too can start to
00:22:20
understand that it's not magic going on inside of these systems but just a bunch
00:22:25
of opinions uh around this commodify human labor and the data the classification and the code and it's
00:22:32
interesting you know we've sort of hinted around sort of like language here um in this case I think I find myself
00:22:39
personally feeling like it does sound like it's magic because the the language
00:22:44
of AI and technology that we're using to talk about these things sounds very
00:22:49
foreign to me and so I think that is what gives it its Mystique is if I can talk about it in a way that doesn't
00:22:56
coincide with lay ter terminology that regular people use then it sounds like it's inaccessible to me right it's like
00:23:03
it's a is that a power move I don't know but language is powerful and language
00:23:09
can be used as a way to exclude and in this case you know I think what you're
00:23:13
helping us to understand brck is you many of you because you learned y equals MX plus b you know that's a language
00:23:21
that you spoke at some point and to the extent that we give people a grounding for understanding this new technology
00:23:28
through something that they already understand language that they already possess then the I think the Mystique
00:23:34
arounds it disappear the power differential between the creators of the technology and the consumers and the
00:23:40
employees who were trying to figure out what does this mean for them that starts
00:23:44
to be reduced as well is that sort of along some of the same lines R of of what you're suggesting that's exactly it
00:23:50
right so like if you hear the term oh chat gbd says they have a billion parameter model right you'll go oh are
00:23:58
big words billion parameter what does that mean let's explain it right let's
00:24:02
break it down real quick so let's go back to this yals MX plus b b is call it
00:24:08
our Y intercept some error term we're not going to worry about that we're just
00:24:11
going to focus on the yals MX Y is an output every computer every machine right they're all uh based on the
00:24:19
touring machine nothing really changes they can only do what they've been told
00:24:22
to do why is going to be the output that comes out all right that could be plausible sounding sentences
00:24:28
that could be art of Jesus flipping over a table whatever all right and then M and X is what matters X is inputs so if
00:24:37
I'm deciding I'm building a system that's going decide whether or not uh
00:24:42
someone gets parole for instance those X's might be ZIP code right problematic
00:24:47
those X's might be past criminal history those X's might be height right those X
00:24:52
could be anything then the m is what matters right the m is a essentially the slope of that line we learned this in uh
00:25:02
uh 10th grade but that slope is an opinion it is some developers opinion or if it's unsupervised it's still some
00:25:09
developers opinion on how much that X matters how much does it matter that I live in the zip code how much does it
00:25:17
matter that you know I'm six foot6 right how much does it matter that my name is
00:25:21
broaderick and then like that opinion gets added in so each one of those mxs we can call that
00:25:28
a parameter so if it's a billion things has a billion opinions but they all get
00:25:32
filtered out and come out to this output why that's it we have now learned machine learning congratulations
00:25:39
everybody right like Pat yourselves on the back if you learned why goes MX plus b you too uh can start to understand
00:25:47
what's going on these systems you can do like Kareem or myself and run audits
00:25:51
where you basically test these systems uh with a bunch of stimuli to see what comes out to explain uh this thing is
00:26:00
maybe leading to inequality because of uh this weirdness right and so let's just demystify the whole thing it's not
00:26:07
I mean it's complex let me not poo poo my computer science folks out there but
00:26:13
uh we can share a language that inside of you know this increasing complexity comes down to a pretty simple building
00:26:21
block and you know the building block already if you made it through high school and our Wharton grads or NBAs and
00:26:28
undergrads I know uh that you learned y equals MX plus b I have to tell you it's
00:26:33
been a long time since I learned why equals MX plus b I'm not sure my high school teacher was as great as you were
00:26:40
explaining that but I feel like I have in the last three minutes a much better understanding of what it is that people
00:26:46
are trying to suggest is sophisticated um which often I think sometimes feels inaccessible but I think you made me
00:26:53
believe and I'm sure others believe that they too can understand sort of like
00:26:57
what the big deal here is and assess for themselves is it a big deal or is it just more uh of what we already know
00:27:04
Kareem let me turn to you and let's get your top two to three issues I don't
00:27:08
know how much they overlap I know earlier you spoke about you know issues around representation in the data not
00:27:14
sure if that's one of your top two or three issues if you have others can you
00:27:17
share with us um from your Vantage Point what are some of the key challenges yeah so first thing I agree
00:27:24
with h bradrick in that you know transparency is a is a major issue right the Splat boox problem and trying to
00:27:31
basically understand uh what companies are doing are companies being transparent about how they're training
00:27:37
these models what data sets they're using are they explaining to their users
00:27:41
what is happening behind the scenes obviously there's an optimal level of transparency as well like transparency
00:27:47
after specific um you know uh particular level might get too in depth to the extent that the uh average user might
00:27:56
zone out right it's irrelevant in information right you don't want to inundate your user with with uh too much
00:28:02
technical details that they really zone out um so for me the the first pressing challenge is addressing uh unfairness in
00:28:09
AI systems right so if you're training data like Rodrick has mentioned is missing certain populations or if your
00:28:16
data is Mis mislabeled obviously that can give rise to bias and can have adverse effects on um you know certain
00:28:24
segments of the population uh this particularly becomes uh you know problematic in issues of like Health
00:28:32
Care employment uh you know uh criminal justice where these decisions are consequential for people right um so if
00:28:43
obviously if these uh these issues of bias are not left unaddressed uh they can perpetuate uh unfairness in society
00:28:50
at a very high rate right we're not just talking about your prototypical kind of
00:28:57
BU we're talking at like an exponential rate with these automated decision systems which is why they um they can be
00:29:05
very dangerous right so the second problem I see is what is called the hallucination problem which is pretty
00:29:11
much like making up stuff right like rodick had mentioned these llms learn to predict the next word right or phrase in
00:29:19
a sentence so um they can misrepresent facts uh they can also tell you a very good story uh that has a very good
00:29:30
narrative uh that is extremely plausible uh but is misleading right and I think this is a particularly big problem uh
00:29:39
given that we as human beings we have like an automation bias where uh we have a propensity to kind of favor
00:29:45
suggestions from an automated decision system uh and to ignore um contradictory information that uh that you know we
00:29:55
might know right we just can't kind of defer to uh this automated decision system right um the third one I would
00:30:04
say is kind of like data privacy and security concerns and this involves things like unauthorized access uh of
00:30:12
data scraping of data that for example the the the company or the llm uh might not have uh access to uh you know or
00:30:22
consent to use uh things like data leakage for example where a large language model might um you know
00:30:31
present some information that a user had used um and we have cases like that coming up in in the media where Samsung
00:30:40
for example has banned uh the use of chat GPT VI its uh employees because of the fact that uh the the llm was uh you
00:30:49
know revealing some of these Trade Secrets uh so things like malicious use and manipulation attempts by nefarious
00:30:57
actors these are all very serious concerns as well yeah so I as as I listen to you all
00:31:04
talk about these concerns it sort of raises the concerns that I've just had again just as an an employee right and
00:31:12
certainly as as a consumer and and certainly what came across my uh social media feed recently was uh in the last
00:31:21
couple weeks there was a uh lawsuit class action lawsuit uh filed in the Northern District of
00:31:29
California um the plaintiffs are open Ai and chat GPT and the Microsoft cor Corporation um and one of the big areas
00:31:40
in included in this um lawsuit is invasion of privacy and it's it's pretty
00:31:47
compelling again I'm not a lawyer no and we none of us knows how this is going to
00:31:51
pan out but just looking at it to understand the the lengths at which people people um don't understand how
00:32:00
their data is being accessed and utilized without their permission um is it can be concerning and so I would
00:32:08
encourage people to it's a it was a nice tutorial I would say if you will for me
00:32:12
in privacy concerns around this topic but I'm going to certainly um put punt
00:32:18
back to the two of you you come and if you've had a chance to like look at anything in relation to this class
00:32:25
action lawsuit I'm wondering if you're surprised by this lawsuit or if you just
00:32:31
thought that this was going to somebody it was just a matter of time before it happen any thoughts on that baric and
00:32:37
then I'll come back to you kareim uh yeah so clearly it was only a matter of
00:32:42
time and we think about again what do these systems actually do all right so these large language models train uh on
00:32:52
previous data and training just means that they take in a bunch of data so they then connected uh together but
00:33:00
whose data did they take where did that data come from now there is some indication that you know when you're
00:33:07
training on a huge Corpus of uh data that you're goingon to tend towards the
00:33:14
cheapest and freest sources right this is why we get some a lot of weirdness that comes out of these systems in terms
00:33:21
of bias because they're trained on the free parts of the internet uh a lot of
00:33:26
the time and what is over represented in the free Parts the internet Stephanie the free parts are people who
00:33:35
I'm trying to think about social media as an example as a source of data and
00:33:39
then I'm thinking about who the users might be um I'm thinking about um not
00:33:44
sure if this is where you're going but I think about all the young people who use
00:33:48
social Med put all of their information resources yeah so we so we have that right we have an over
00:33:54
representation of younger folks all right so we're going to have weird age related things that don't really exist
00:34:01
in the data the other thing that gets over represented on the internet on the free Parts at least is propaganda right
00:34:08
so some websites where they'll say really negative things about presidents for instance are free whereas if I go to
00:34:16
I don't know the New York Times I can read four articles before I get to a firewall and they're like no sir you're
00:34:21
done right uh and so there's going to be some weirdness that's in the data from
00:34:26
that the other that's free on the internet or over represented for free on the internet is pornography right so if
00:34:32
I'm taking in porn and propaganda into these systems right then I'm going to
00:34:37
get all types of weirdness out and the only way to improve these things is to get access to uh you know data that has
00:34:48
closer to accuracy that has time spin into it you want to read one of our research articles Stephanie how much
00:34:53
does it cost if you're like not affiliate with school it's like thousands of dollars right and so if
00:34:58
you're one of these companies and you want to improve your data how do you do
00:35:02
it do you spend thousands of dollars per article from Stephanie Cur or you steal
00:35:07
it right is that like a a rhetorical question you can't say but I'll answer
00:35:15
the question the answer is you're probably you're probably not gonna spend
00:35:19
thousands of dollars per article per Professor If instead you can steal it right and this goes for any of we'll
00:35:27
call it better data uh because I need a bunch of it to get to some version of it's not accurate but a better uh
00:35:36
distribution of data and so yeah of course they're getting sued because to make the system better they had to take
00:35:43
this stuff in and that we do have intellectual property laws somebody's going to have to pay or they're gonna
00:35:49
have to change the law Kareem do you want to chime in on this conversation before we move to the how do we fix it
00:35:56
yeah absolutely I I think um you know lawsuits are just going to mushroom from here onwards uh it's just a matter of
00:36:03
time honestly you know every day we're hearing about another lawsuit uh you know whether it's relating to privacy or
00:36:11
discrimination or um other aspects um so to me I'm I'm not too surprised I think
00:36:19
uh these companies are you know obviously I I can't speak on on their behalf but you know we have a responsib
00:36:26
ility to make sure that we're training our llms on reliable data sources uh we
00:36:32
need to kind of fine-tune these llms we need to have like some kind of a factchecking mechanism to ensure uh you
00:36:39
know if they're producing uh garbage that you know they're they're being kind
00:36:44
of they're receiving feedback and and they're improving we have to have human
00:36:48
reviewers in the loop uh that can play a role in in correcting like kind of the the trajectory of of these llms um you
00:36:59
know perhaps even including like confidence scores to allow users to kind of understand Engage The reliability of
00:37:06
the information that they're getting right and uh and even to uh provide users with a mechanism to provide
00:37:15
feedback for the llm as to whether you know the response that I got was terrible or not so that uh they're
00:37:23
taking in another signal from their users so I think in conjunction all of these different measures uh hopefully
00:37:30
over the course of time and as long as there's you know strong Buy in from leadership to kind of improve these
00:37:37
models I think um companies can do a better job at uh protecting the privacy of people but also ensuring that you
00:37:45
know whatever gets propagated uh in terms of outputs is more reliable okay uh any other suggestions Solutions um
00:37:54
around that or the other challenges that you and brck brought up today um in like
00:38:01
we think about like one or two key things that can be done uh for various audiences and I think I'm thinking of
00:38:08
lots of audiences I'm thinking of um you know employers and institutions I'm
00:38:14
thinking of consumers and employees as well yeah there's I let kareim go first
00:38:22
on that one yeah I think there's the grass is so green um there is so much that can be done in this space um so
00:38:32
firstly there needs to be uh legislation there needs to be comprehensive AI legislation and enforcement to protect
00:38:38
public interests uh not sufficient for companies to have voluntary commitments that's good but it uh it's not
00:38:46
sufficient second thing is you know there needs to be um when we're thinking
00:38:50
about data protection laws uh they need to be strengthened to include you know AI systems and Ure that um the Privacy
00:38:58
rights of users and so forth uh you know we spoke about transparency and that companies uh need to disclose how AI
00:39:07
systems have been built what data sources they've been using and to kind of um to take a crack at the blackbox
00:39:15
problem uh stakeholder engagement is really important you know as we're thinking through diverse audiences and
00:39:22
data uh sources uh you need to engage with different stakeholders I always say when you're working in this
00:39:29
space you really need to get it like be cross functional because there's a lot
00:39:34
of barbwire in this space you know whether it is you know privacy legal regulatory you're working with ethicists
00:39:40
product managers you're working with uh data scientists researchers uh you know
00:39:45
Eng it is extremely cross functional space and it has to be that way because it is a soot technical problem that uh
00:39:54
takes all the great minds to be able to attack it from different ways you need to have diverse teams for example so
00:40:00
that you have representation people can identify and rectify bias effectively right training data needs to be more
00:40:07
inclusive make sure you know like bradrick was saying who's missing out from this data like who are we not
00:40:13
seeing you know I always say that um this technology is not neutral um this technology is already has a point of
00:40:21
view and its point of view is what it has been trained on right so if it has been missing people of color for example
00:40:30
then that's its point of view right it's not like it's coming as a blank slate no
00:40:36
it is already has a point of view and so uh that's a little bit of mythbusting so
00:40:42
I mean we need to kind of like take inclusive data and inclusive uh inclusivity into consideration uh
00:40:51
throughout the product development life cycle right we need to have like audits that are done on uh these these models
00:41:00
and and ensure that you know we have uh results and that can kind of feed into how we can improve uh human oversight
00:41:08
for example I mentioned this earlier that there needs to be a human in the loop to be able to kind of uh you know
00:41:15
spot and uh improve the trajectory of uh outcomes I mean there's so much that can
00:41:21
be done I can go on and on and on but start yeah so much brck your turn what would
00:41:30
you say um are some key specific things that can be done so I'm going to break
00:41:35
this down into uh different market segments because that's how my brain is wired so we're going to talk about what
00:41:41
can companies do what can consumers do what can researchers do and then what can your students do all right so first
00:41:49
for companies consider that 3D model that I laid out earlier before you roll out a product what is this product do
00:41:56
who does it disempower uh and what uh is its daily use all right if you need help
00:42:04
answering these questions then call RA audits uh you know Kore will pick up the phone uh or you can holler at the folks
00:42:12
over here in the Trap lab right and we can help you think through some of these things uh as we're trying to make this
00:42:18
technology better and more Equitable all right two consumers what can you do do not accept that It's Magic there is
00:42:28
nothing magical about these systems these systems are just people and their opinions of people if you learned yals
00:42:37
MX plus b then you have learned the building blocks of every machine Learning System all right so if there is
00:42:45
some weird outcome that comes out when you're on Facebook or Twitter or you notice some weirdness on your uh when
00:42:53
you you put out an application for a loan trust your gut all right it is weird and
00:42:59
say something because again the only way to improve these things as well is uh uh
00:43:04
for them to update their model and it's you you are right that something's wrong
00:43:09
all right it is not magic is not your fault is them uh don't accept magic for
00:43:16
researchers if you are interested in working on these topics and in this space go to jointhe trap.com and come
00:43:23
hang out with us uh we meet online once a week over Zoom Wednesdays at 1 pm uh and then finally for students this is
00:43:34
going to be a weird one uh because in some ways this rise of uh generative Ai and writing and art uh makes people
00:43:46
believe that we are like this close to uh having AI that can write beautiful books or make beautiful art uh and I'm
00:43:57
going to challenge that all right we they also told us we were this close to having fully self-driving cars uh within
00:44:04
a year 10 years ago and they said this every year and it's not happening anytime soon all right I'm going to say
00:44:10
the same thing is the case for art and writing and uh creativity I think that they will be a
00:44:18
premium uh on people that can actually Express themselves clearly and accurately and honestly and so if you
00:44:27
are currently a student uh anywhere if you're in a b school if you're in a
00:44:31
college if you're in high school and you're listening to this I don't know
00:44:34
why you're listening to a Wharton podcast in high school but maybe you know you're one of those kids uh get
00:44:40
really into building your uh toolbox of creativity right get really into uh creative writing right get really into
00:44:51
making more art because there will be a value on the ACT actual human element that comes out of this if uh what
00:45:00
everyone else is doing is just plugging in some chatbot that gives you uh plausible sentences that aren't that
00:45:07
good if you can write better sentences you'll win thank you and and I'm just going to
00:45:13
tell you project we actually do have knowledge at war in high school so your high school learners are going to learn
00:45:20
a lot from what you all share today as much as your you know fully experienced senior leaders will um as well uh so so
00:45:29
broject and Kem this has been fantastic I feel so much more knowledgeable as somebody who believes that I know a lot
00:45:36
about a lot of things I I know that I don't know a lot about this topic and so
00:45:40
the past amount of time 45 minutes that we've been chatting have been amazing
00:45:44
and just I think helping me to feel more empowered um as a researcher a consumer
00:45:49
and a worker around sort of what exactly is happening and and how I need to pay attention um to to what is being shared
00:45:58
um so I want to thank you uh so much for sharing your insights and your expertise
00:46:02
with us uh and to all the uh leading diversity at work podcast listeners we truly appreciate you for being here uh
00:46:09
so that's all for now thanks to our audience for joining us and listening to
00:46:14
this episode of the knowledge at Warton leading diversity atw work podcast series goodbye for
00:46:19
now for more insight from knowledge at Wharton please visit knowledge. won. [Music]
00:46:34
EDU

Episode Highlights

  • Introduction to Responsible AI
    Stephanie Cy welcomes guests to discuss responsible AI in the workplace.
    @ 00m 23s
    November 09, 2023
  • Dr. Broer Turner's Research
    Dr. Turner shares insights on race and technology's impact on consumer decisions.
    @ 01m 07s
    November 09, 2023
  • Dr. Kareem Gan's AI Governance
    Dr. Gan discusses his role in AI governance and the importance of ethical practices.
    @ 01m 16s
    November 09, 2023
  • The Importance of Representation
    The conversation highlights the need for diverse representation in AI development.
    @ 21m 01s
    November 09, 2023
  • Demystifying AI Models
    Understanding AI isn't magic; it's about basic math principles like y = MX + b.
    “If you learned y equals MX plus b, you can understand these systems.”
    @ 25m 39s
    November 09, 2023
  • The Hallucination Problem
    AI can misrepresent facts, leading to dangerous consequences in decision-making.
    “These LLMs can misrepresent facts and tell plausible but misleading stories.”
    @ 29m 11s
    November 09, 2023
  • Class Action Lawsuit
    A lawsuit against OpenAI raises serious privacy concerns regarding data usage.
    “People don’t understand how their data is being accessed and utilized without permission.”
    @ 31m 27s
    November 09, 2023
  • The Illusion of AI Creativity
    Generative AI may not be as close to creating art as we think. 'They told us we were this close to self-driving cars... and it’s not happening anytime soon.'
    “They also told us we were this close to having fully self-driving cars.”
    @ 44m 01s
    November 09, 2023
  • Empowering Future Creatives
    Students are encouraged to hone their creativity as it will be valued in the future. 'Get really into building your toolbox of creativity.'
    “Get really into building your toolbox of creativity.”
    @ 44m 40s
    November 09, 2023

Episode Quotes

  • I like to think of myself as a novice on these topics.
    Diversity Is Critical for the Future of AI — Leading Diversity at Work
  • This technology is extremely powerful.
    Diversity Is Critical for the Future of AI — Leading Diversity at Work
  • It's not magic going on inside these systems, just a bunch of opinions.
    Diversity Is Critical for the Future of AI — Leading Diversity at Work
  • If you learned y equals MX plus b, you can understand these systems.
    Diversity Is Critical for the Future of AI — Leading Diversity at Work
  • This technology is not neutral; it already has a point of view.
    Diversity Is Critical for the Future of AI — Leading Diversity at Work
  • Trust your gut; it is weird and say something.
    Diversity Is Critical for the Future of AI — Leading Diversity at Work

Key Moments

  • Introduction00:15
  • Guest Introductions00:35
  • AI Governance Discussion01:16
  • Understanding AI22:25
  • Transparency Issues27:24
  • Bias in AI28:07
  • Privacy Concerns31:27
  • Empower Future Generations44:40

Tension Over Time

Words per Minute Over Time

Vibes Breakdown