Search Captions & Ask AI

Making Statistics Work in the Real World

June 29, 2017 / 04:28

This episode features Wharton statistics Professor Bashor Bataria discussing his research on graph-based methods in statistics, computational complexity, and applications in business.

Professor Bataria explains the intersection of statistics, probability, and combinatorics, highlighting the emergence of graph-based methods due to network data. He emphasizes the efficiency of these methods in handling large datasets and their near-optimal performance.

He discusses practical applications, such as the two-sample problem in genetics, where researchers compare gene expression levels between patients with diabetes and healthy individuals. This research aims to provide theoretical understanding and method comparisons.

Another application mentioned is in natural language processing, particularly in analyzing word similarity, which can benefit businesses using social media data. Bataria also shares his ongoing work in high-dimensional data analysis.

Overall, the episode provides insights into the practical implications of Bataria's research for businesses and future directions in statistical methods.

TLDR

Wharton Professor Bashor Bataria discusses graph-based methods in statistics and their applications in genetics and natural language processing.

Episode

4:28
00:00:02
we're here today with Wharton statistics Professor bashor bataria to talk about
00:00:05
some of his latest research bashor thanks for being with us today could you first of all talk to us a little bit
00:00:10
about give us a brief summary of your research what kind of question you were trying to answer right so my research
00:00:15
interests are the intersection of Statistics probability and combinatorics so recently numerous very interesting
00:00:20
combinatorial and graph theoretic problems of emerging statistics mainly because of the ubiquitous presence of
00:00:26
network data and the increasing use of graph based methods in modern analytics and as a consequence many interesting
00:00:32
connections have emerged between modern statistical methods and classical Concepts in geometry and probability and
00:00:39
you know as a consequence you can use them to solve interesting problems in statistics great and tell us a little
00:00:44
bit about in this study what were the key takeaways that you took away that you took away from the study the key
00:00:50
takeaways of my research basically the interplay between computational complexity which is the time it takes to
00:00:56
to implement one of these methods and its statistical performance perance which is how close it is to the
00:01:01
mathematically based procedure it turns out that many of these graph based methods are computationally very
00:01:07
efficient so they can be aply used to solve uh and apply in large data sets and we have also shown that in many
00:01:14
situations these tests have near Optimal Performance guarantees which provide the
00:01:19
theoretical justification required for using these procedures now had graph based methods is that something that has
00:01:25
not been used as much in the past or had been looked at with more skepticism mean
00:01:29
why is this right so so graph based methods have been used in practice for a long time
00:01:34
but I think what comes out of my research is the the answer to the question that why they work so they were
00:01:38
used and they were working fine before but here is so here we provide some theory behind why it works giving the
00:01:44
justification of using it in real problems great and so getting to that actually like if I'm a business if I'm a
00:01:50
business person or I own a business I mean how can I apply This research practically in my life right so one of
00:01:58
the recent projects we're looking at is what is known as the two sample problem
00:02:01
so imagine I have a situation where I want to find out whether a set of genes regulate or affect the occurrence of a
00:02:08
disease so for example to just illustrate suppose I have uh uh 20 genes and I have the gene expression level
00:02:15
data from 100 patients who have diabetes and I have the same expression level data for 100 patients who are healthy
00:02:22
and the goal is to find out whether these 20 genes are expressed differentially so by that I mean that
00:02:28
their expression levels are sign signicantly different between these two sets of patients and uh our research
00:02:34
sort of aims to provide the theoretical understanding and comparison of the different methods that are deployed to
00:02:41
understand or answer such questions what are some other applications for This research right so another interesting
00:02:46
application of our work is in natural language processing mainly in problems which try to understand similarity
00:02:52
between words so imagine the word color which can be spelled in two ways one with the letter U one without the letter
00:02:58
U and they are the same words but the words for instance wolf and fox are both animals but they are very different
00:03:05
words so in this case in spite of the amount of data we have the support size which is basically the collection of all
00:03:12
words is far larger than the than the data set itself so one of our methods is what we are studying can be used to
00:03:19
analyze such problems as well which I would think would be pretty interesting to businesses nowaday just because so
00:03:24
many people are like taking social media posts and things like that trying to get
00:03:28
data about customers that way right perfectly yeah yeah yeah sure and so what's next for This research right so
00:03:34
currently I'm trying to understand or or analyze the methods for analyzing data
00:03:39
in what is known as the high dimensional setting where you have say 10,000 genes
00:03:43
and only a few hundred patients and uh you want to find out something about how the genes affect the disease or the
00:03:49
patients and for these cases different new techniques are required and I'm trying to understand the theoretical uh
00:03:56
the uh background of these results and how these can be used to find to get new methods and new algorithms great BR
00:04:02
thanks for being with us today thank you thank you [Music]

Episode Highlights

  • Exploring Graph-Based Methods
    Professor Bashor Bataria discusses the efficiency of graph-based methods in statistics.
    “Many of these graph-based methods are computationally very efficient.”
    @ 01m 05s
    June 29, 2017
  • The Two Sample Problem
    A practical application of the research involves analyzing gene expression levels.
    “The goal is to find out whether these 20 genes are expressed differentially.”
    @ 02m 06s
    June 29, 2017
  • Applications in Natural Language Processing
    The research extends to understanding word similarities in language processing.
    “Our methods can analyze problems in natural language processing.”
    @ 02m 46s
    June 29, 2017

Episode Quotes

  • Graph-based methods have been used in practice for a long time.
    Making Statistics Work in the Real World
  • Our research aims to provide theoretical understanding of different methods.
    Making Statistics Work in the Real World
  • Imagine the word color, spelled with or without a U.
    Making Statistics Work in the Real World

Key Moments

  • Introduction to Research00:02
  • Graph Theory Insights00:19
  • Key Takeaways00:48
  • Natural Language Processing02:46
  • Future Directions03:34

Tension Over Time

Words per Minute Over Time

Vibes Breakdown