{"id":808,"date":"2023-08-08T09:26:59","date_gmt":"2023-08-08T09:26:59","guid":{"rendered":"https:\/\/tbekk.com\/devstream\/?p=808"},"modified":"2023-08-08T09:27:36","modified_gmt":"2023-08-08T09:27:36","slug":"fundamentals-of-statistics-for-data-scientists-and-analysts","status":"publish","type":"post","link":"https:\/\/tbekk.com\/devstream\/2023\/08\/08\/fundamentals-of-statistics-for-data-scientists-and-analysts\/","title":{"rendered":"Fundamentals Of Statistics For Data Scientists and Analysts"},"content":{"rendered":"\n<p><em>Key statistical concepts for your data science or data analysis journey.<\/em><\/p>\n\n\n\n<hr class=\"wp-block-separator has-text-color has-light-gray-color has-alpha-channel-opacity has-light-gray-background-color has-background is-style-wide\"\/>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong><em>Link:<\/em><\/strong><em><strong> <\/strong><\/em><a href=\"https:\/\/www.kdnuggets.com\/023\/08\/fundamentals-statistics-data-scientists-analysts.html\"><em>kdnuggets<\/em><\/a><\/li>\n\n\n\n<li><em><strong>Author:<\/strong><\/em> <strong><a href=\"https:\/\/www.kdnuggets.com\/author\/tatevkaren-aslanyan\"><em>Tatev Karen Aslanyan<\/em><\/a><\/strong><\/li>\n\n\n\n<li><em><strong>Publication date:<\/strong><\/em> <em>August 7, 2023<\/em><\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-text-color has-light-gray-color has-alpha-channel-opacity has-light-gray-background-color has-background is-style-wide\"\/>\n\n\n\n<p>As Karl Pearson, a British mathematician has once stated,&nbsp;<strong>Statistics<\/strong>&nbsp;is the grammar of science and this holds especially for Computer and Information Sciences, Physical Science, and Biological Science. When you are getting started with your journey in&nbsp;<strong>Data Science<\/strong>&nbsp;or&nbsp;<strong>Data Analytics<\/strong>, having statistical knowledge will help you to better leverage data insights.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>\u201cStatistics is the grammar of science.\u201d&nbsp;<strong>Karl Pearson<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p>The importance of statistics in data science and data analytics cannot be underestimated. Statistics provides tools and methods to find structure and to give deeper data insights. Both Statistics and Mathematics love facts and hate guesses. Knowing the fundamentals of these two important subjects will allow you to think critically, and be creative when using the data to solve business problems and make data-driven decisions. In this article, I will cover the following Statistics topics for data science and data analytics:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code><strong>- Random variables\n\n- Probability distribution functions (PDFs)\n\n- Mean, Variance, Standard Deviation\n\n- Covariance and Correlation \n\n- Bayes Theorem\n\n- Linear Regression and Ordinary Least Squares (OLS)\n\n- Gauss-Markov Theorem\n\n- Parameter properties (Bias, Consistency, Efficiency)\n\n- Confidence intervals\n\n- Hypothesis testing\n\n- Statistical significance \n\n- Type I &amp; Type II Errors\n\n- Statistical tests (Student's t-test, F-test)\n\n- p-value and its limitations\n\n- Inferential Statistics \n\n- Central Limit Theorem &amp; Law of Large Numbers\n\n- Dimensionality reduction techniques (PCA, FA)<\/strong><\/code><\/pre>\n\n\n\n<p><em>If you have no prior Statistical knowledge and you want to identify and learn the essential statistical concepts from the scratch, to prepare for your job interviews, then this article is for you. This article will also be a good read for anyone who wants to refresh his\/her statistical knowledge.<\/em><\/p>\n\n\n\n<h1 class=\"wp-block-heading\">Before we start, welcome to LunarTech!<\/h1>\n\n\n\n<p>Welcome to&nbsp;<a href=\"http:\/\/lunartech.ai\/\" rel=\"noreferrer noopener\" target=\"_blank\"><strong>LunarTech.ai<\/strong><\/a>, where we understand the power of job-searching strategies in the dynamic field of Data Science and AI. We dive deep into the tactics and strategies required to navigate the competitive job search process. Whether it\u2019s defining your career goals, customizing application materials, or leveraging job boards and networking, our insights provide the guidance you need to land your dream job.<\/p>\n\n\n\n<p>Preparing for data science interviews? Fear not! We shine a light on the intricacies of the interview process, equipping you with the knowledge and preparation necessary to increase your chances of success. From initial phone screenings to technical assessments, technical interviews, and behavioral interviews, we leave no stone unturned.<\/p>\n\n\n\n<p>At&nbsp;<a href=\"http:\/\/lunartech.ai\/\" rel=\"noreferrer noopener\" target=\"_blank\">LunarTech.ai<\/a>, we go beyond the theory. We\u2019re your springboard to unparalleled success in the tech and data science realm. Our comprehensive learning journey is tailored to fit seamlessly into your lifestyle, allowing you to strike the perfect balance between personal and professional commitments while acquiring cutting-edge skills. With our dedication to your career growth, including job placement assistance, expert resume building, and interview preparation, you\u2019ll emerge as an industry-ready powerhouse.<\/p>\n\n\n\n<p>Join our community of ambitious individuals today and embark on this thrilling data science journey together. With&nbsp;<a href=\"http:\/\/lunartech.ai\/\" rel=\"noreferrer noopener\" target=\"_blank\">LunarTech.ai<\/a>, the future is bright, and you hold the keys to unlock boundless opportunities.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\">Random Variables<\/h1>\n\n\n\n<p>The concept of random variables forms the cornerstone of many statistical concepts. It might be hard to digest its formal mathematical definition but simply put, a&nbsp;<strong>random variable<\/strong>&nbsp;is a way to map the outcomes of random processes, such as flipping a coin or rolling a dice, to numbers. For instance, we can define the random process of flipping a coin by random variable X which takes a value 1 if the outcome if&nbsp;<em>heads&nbsp;<\/em>and 0 if the outcome is&nbsp;<em>tails.<\/em><img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_01.png\" width=\"100%\"><\/p>\n\n\n\n<p>In this example, we have a random process of flipping a coin where this experiment can produce&nbsp;<strong><em>two<\/em><\/strong>&nbsp;<strong><em>possible outcomes<\/em><\/strong>: {0,1}. This set of all possible outcomes is called the&nbsp;<strong><em>sample space&nbsp;<\/em><\/strong>of the experiment. Each time the random process is repeated, it is referred to as an&nbsp;<strong><em>event.&nbsp;<\/em><\/strong>In this example, flipping a coin and getting a tail as an outcome is an event. The chance or the likelihood of this event occurring with a particular outcome is called the&nbsp;<strong><em>probability<\/em><\/strong>&nbsp;of that event. A probability of an event is the likelihood that a random variable takes a specific value of x which can be described by P(x). In the example of flipping a coin, the likelihood of getting heads or tails is the same, that is 0.5 or 50%. So we have the following setting:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_02.png\" width=\"100%\"><\/p>\n\n\n\n<p>where the probability of an event, in this example, can only take values in the range [0,1].<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>The importance of statistics in data science and data analytics cannot be underestimated. Statistics provides tools and methods to find structure and to give deeper data insights.<\/p>\n<\/blockquote>\n\n\n\n<h1 class=\"wp-block-heading\">Mean, Variance, Standard Deviation<\/h1>\n\n\n\n<p>To understand the concepts of mean, variance, and many other statistical topics, it is important to learn the concepts of&nbsp;<strong><em>population<\/em><\/strong>&nbsp;and&nbsp;<strong><em>sample<\/em>.&nbsp;<\/strong>The<strong>&nbsp;<em>population<\/em><\/strong>&nbsp;is the set of all observations (individuals, objects, events, or procedures) and is usually very large and diverse, whereas a&nbsp;<strong><em>sample<\/em><\/strong><em>&nbsp;<\/em>is a subset of observations from the population that ideally is a true representation of the population.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_03.png\" width=\"100%\"><br>Image Source: The Author<\/p>\n\n\n\n<p>Given that experimenting with an entire population is either impossible or simply too expensive, researchers or analysts use samples rather than the entire population in their experiments or trials. To make sure that the experimental results are reliable and hold for the entire population, the sample needs to be a true representation of the population. That is, the sample needs to be unbiased. For this purpose, one can use statistical sampling techniques such as&nbsp;<a href=\"https:\/\/github.com\/TatevKaren\/mathematics-statistics-for-data-science\/tree\/main\/Sampling%20Techniques\" rel=\"noreferrer noopener\" target=\"_blank\">Random Sampling, Systematic Sampling, Clustered Sampling, Weighted Sampling, and Stratified Sampling.<\/a><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Mean<\/h2>\n\n\n\n<p>The mean, also known as the average, is a central value of a finite set of numbers. Let\u2019s assume a random variable X in the data has the following values:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_04.png\" width=\"100%\"><\/p>\n\n\n\n<p>where N is the number of observations or data points in the sample set or simply the data frequency. Then the&nbsp;<strong><em>sample mean<\/em><\/strong><em>&nbsp;<\/em>defined by&nbsp;<strong>?<\/strong>, which is very often used to approximate the&nbsp;<strong><em>population mean<\/em><\/strong><em>,&nbsp;<\/em>can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_05.png\" width=\"100%\"><\/p>\n\n\n\n<p>The mean is also referred to as&nbsp;<strong><em>expectation<\/em>&nbsp;<\/strong>which is often defined by&nbsp;<strong>E<\/strong>() or random variable with a bar on the top. For example, the expectation of random variables X and Y, that is<strong>&nbsp;E<\/strong>(X) and&nbsp;<strong>E<\/strong>(Y), respectively, can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_06.png\" width=\"100%\"><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>import numpy as np\nimport math\nx = np.array(&#91;1,3,5,6])\nmean_x = np.mean(x)\n# in case the data contains Nan values\nx_nan = np.array(&#91;1,3,5,6, math.nan])\nmean_x_nan = np.nanmean(x_nan)<\/code><\/pre>\n\n\n\n<h2 class=\"wp-block-heading\">Variance<\/h2>\n\n\n\n<p>The variance measures how far the data points are spread out from the average value<em>,<\/em>&nbsp;and is equal to the sum of squares of differences between the data values and the average (the mean). Furthermore, the&nbsp;<strong><em>population variance<\/em><\/strong><em>,&nbsp;<\/em>can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_07.png\" width=\"100%\"><\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>x = np.array(&#91;1,3,5,6])\nvariance_x = np.var(x)\n\n# here you need to specify the degrees of freedom (df) max number of logically independent data points that have freedom to vary\nx_nan = np.array(&#91;1,3,5,6, math.nan])\nmean_x_nan = np.nanvar(x_nan, ddof = 1)<\/code><\/pre>\n\n\n\n<p>For deriving expectations and variances of different popular probability distribution functions,<a href=\"https:\/\/github.com\/TatevKaren\/mathematics-statistics-for-data-science\/tree\/main\/Deriving%20Expectation%20and%20Variances%20of%20Densities\" rel=\"noreferrer noopener\" target=\"_blank\">&nbsp;check out this Github repo<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Standard Deviation<\/h2>\n\n\n\n<p>The standard deviation is simply the square root of the variance and measures the extent to which data varies from its mean. The standard deviation defined by&nbsp;<strong><em>sigma<\/em><\/strong><em>&nbsp;<\/em>can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_08.png\" width=\"100%\"><\/p>\n\n\n\n<p>Standard deviation is often preferred over the variance because it has the same unit as the data points, which means you can interpret it more easily.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>x = np.array(&#91;1,3,5,6])\nvariance_x = np.std(x)\n\nx_nan = np.array(&#91;1,3,5,6, math.nan])\nmean_x_nan = np.nanstd(x_nan, ddof = 1)<\/code><\/pre>\n\n\n\n<h2 class=\"wp-block-heading\">Covariance<\/h2>\n\n\n\n<p>The covariance is a measure of the joint variability of two random variables and describes the relationship between these two variables. It is defined as the expected value of the product of the two random variables\u2019 deviations from their means. The covariance between two random variables X and Z can be described by the following expression, where&nbsp;<strong>E<\/strong>(X) and&nbsp;<strong>E<\/strong>(Z) represent the means of X and Z, respectively.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_09.png\" width=\"100%\"><\/p>\n\n\n\n<p>Covariance can take negative or positive values as well as value 0. A positive value of covariance indicates that two random variables tend to vary in the same direction, whereas a negative value suggests that these variables vary in opposite directions. Finally, the value 0 means that they don\u2019t vary together.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>x = np.array(&#91;1,3,5,6])\ny = np.array(&#91;-2,-4,-5,-6])\n#this will return the covariance matrix of x,y containing x_variance, y_variance on diagonal elements and covariance of x,y\ncov_xy = np.cov(x,y)<\/code><\/pre>\n\n\n\n<h2 class=\"wp-block-heading\">Correlation<\/h2>\n\n\n\n<p>The correlation is also a measure for relationship and it measures both the strength and the direction of the linear relationship between two variables. If a correlation is detected then it means that there is a relationship or a pattern between the values of two target variables. Correlation between two random variables X and Z are equal to the covariance between these two variables divided to the product of the standard deviations of these variables which can be described by the following expression.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_10.png\" width=\"100%\"><\/p>\n\n\n\n<p>Correlation coefficients\u2019 values range between -1 and 1. Keep in mind that the correlation of a variable with itself is always 1, that is&nbsp;<strong>Cor(X, X) = 1<\/strong>. Another thing to keep in mind when interpreting correlation is to not confuse it with&nbsp;<strong><em>causation<\/em><\/strong>, given that a correlation is not causation. Even if there is a correlation between two variables, you cannot conclude that one variable causes a change in the other. This relationship could be coincidental, or a third factor might be causing both variables to change.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>x = np.array(&#91;1,3,5,6])\ny = np.array(&#91;-2,-4,-5,-6])\ncorr = np.corrcoef(x,y)<\/code><\/pre>\n\n\n\n<h1 class=\"wp-block-heading\">Probability Distribution Functions<\/h1>\n\n\n\n<p>A function that describes all the possible values, the sample space, and the corresponding probabilities that a random variable can take within a given range, bounded between the minimum and maximum possible values, is called&nbsp;<strong><em>a probability distribution function (pdf)<\/em><\/strong>&nbsp;or probability density. Every pdf needs to satisfy the following two criteria:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_11.png\" width=\"100%\"><\/p>\n\n\n\n<p>where the first criterium states that all probabilities should be numbers in the range of [0,1] and the second criterium states that the sum of all possible probabilities should be equal to 1.<\/p>\n\n\n\n<p>Probability functions are usually classified into two categories:&nbsp;<strong><em>discrete<\/em><\/strong>&nbsp;and&nbsp;<strong><em>continuous<\/em><\/strong>. Discrete<em>&nbsp;<\/em>distribution<em>&nbsp;<\/em>function describes the random process with&nbsp;<strong><em>countable<\/em><\/strong>&nbsp;sample space, like in the case of an example of tossing a coin that has only two possible outcomes. Continuous<em>&nbsp;<\/em>distribution function describes the random process with&nbsp;<strong><em>continuous<\/em><\/strong>&nbsp;sample space. Examples of discrete distribution functions are&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Bernoulli_distribution\" rel=\"noreferrer noopener\" target=\"_blank\">Bernoulli<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Binomial_distribution\" rel=\"noreferrer noopener\" target=\"_blank\">Binomial<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Poisson_distribution\" rel=\"noreferrer noopener\" target=\"_blank\">Poisson<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Discrete_uniform_distribution\" rel=\"noreferrer noopener\" target=\"_blank\">Discrete Uniform<\/a>. Examples of continuous distribution functions are&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Normal_distribution\" rel=\"noreferrer noopener\" target=\"_blank\">Normal<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Continuous_uniform_distribution\" rel=\"noreferrer noopener\" target=\"_blank\">Continuous Uniform<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Cauchy_distribution\" rel=\"noreferrer noopener\" target=\"_blank\">Cauchy<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Binomial Distribution<\/h2>\n\n\n\n<p><a href=\"https:\/\/brilliant.org\/wiki\/binomial-distribution\/\" rel=\"noreferrer noopener\" target=\"_blank\">The binomial distribution&nbsp;<\/a>is the discrete probability distribution of the number of successes in a sequence of&nbsp;<strong>n<\/strong>&nbsp;independent experiments, each with the boolean-valued outcome:&nbsp;<strong><em>success<\/em><\/strong>&nbsp;(with probability&nbsp;<strong>p<\/strong>) or&nbsp;<strong><em>failure<\/em><\/strong>&nbsp;(with probability&nbsp;<strong>q<\/strong>&nbsp;= 1 ? p). Let&#8217;s assume a random variable X follows a Binomial distribution, then the probability of observing<em>&nbsp;<\/em><strong><em>k&nbsp;<\/em><\/strong>successes in n independent trials can be expressed by the following probability density function:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_12.png\" width=\"100%\"><\/p>\n\n\n\n<p>The binomial distribution is useful when analyzing the results of repeated independent experiments, especially if one is interested in the probability of meeting a particular threshold given a specific error rate.<\/p>\n\n\n\n<p><strong>Binomial Distribution Mean &amp; Variance<\/strong><img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_13.png\" width=\"100%\"><\/p>\n\n\n\n<p>The figure below visualizes an example of Binomial distribution where the number of independent trials is equal to 8 and the probability of success in each trial is equal to 16%.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_14.png\" width=\"100%\"><br>Image Source: The Author<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code># Random Generation of 1000 independent Binomial samples\nimport numpy as np\nn = 8\np = 0.16\nN = 1000\nX = np.random.binomial(n,p,N)\n# Histogram of Binomial distribution\nimport matplotlib.pyplot as plt\ncounts, bins, ignored = plt.hist(X, 20, density = True, rwidth = 0.7, color = 'purple')\nplt.title(\"Binomial distribution with p = 0.16 n = 8\")\nplt.xlabel(\"Number of successes\")\nplt.ylabel(\"Probability\")\nplt.show()<\/code><\/pre>\n\n\n\n<h2 class=\"wp-block-heading\">Poisson Distribution<\/h2>\n\n\n\n<p><a href=\"https:\/\/brilliant.org\/wiki\/poisson-distribution\/\" rel=\"noreferrer noopener\" target=\"_blank\">The Poisson distribution<\/a>&nbsp;is the discrete probability distribution of the number of events occurring in a specified time period, given the average number of times the event occurs over that time period. Let&#8217;s assume a random variable X follows a Poisson distribution, then the probability of observing<em>&nbsp;<\/em><strong><em>k&nbsp;<\/em><\/strong>events over a time period can be expressed by the following probability function:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_15.png\" width=\"100%\"><\/p>\n\n\n\n<p>where&nbsp;<strong><em>e<\/em><\/strong>&nbsp;is&nbsp;<a href=\"https:\/\/brilliant.org\/wiki\/eulers-number\/\" rel=\"noreferrer noopener\" target=\"_blank\"><strong><em>Euler\u2019s number<\/em><\/strong><\/a>&nbsp;and&nbsp;<strong><em>?&nbsp;<\/em><\/strong>lambda, the&nbsp;<strong><em>arrival rate parameter&nbsp;<\/em><\/strong>is<strong><em>&nbsp;<\/em><\/strong>the expected value of X. Poisson distribution function is very popular for its usage in modeling countable events occurring within a given time interval.<\/p>\n\n\n\n<p><strong>Poisson Distribution Mean &amp; Variance<\/strong><img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_16.png\" width=\"100%\"><\/p>\n\n\n\n<p>For example, Poisson distribution can be used to model the number of customers arriving in the shop between 7 and 10 pm, or the number of patients arriving in an emergency room between 11 and 12 pm. The figure below visualizes an example of Poisson distribution where we count the number of Web visitors arriving at the website where the arrival rate, lambda, is assumed to be equal to 7 minutes.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_17.png\" width=\"100%\"><br>Image Source: The Author<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code># Random Generation of 1000 independent Poisson samples\nimport numpy as np\nlambda_ = 7\nN = 1000\nX = np.random.poisson(lambda_,N)\n\n# Histogram of Poisson distribution\nimport matplotlib.pyplot as plt\ncounts, bins, ignored = plt.hist(X, 50, density = True, color = 'purple')\nplt.title(\"Randomly generating from Poisson Distribution with lambda = 7\")\nplt.xlabel(\"Number of visitors\")\nplt.ylabel(\"Probability\")\nplt.show()<\/code><\/pre>\n\n\n\n<h2 class=\"wp-block-heading\">Normal Distribution<\/h2>\n\n\n\n<p><a href=\"https:\/\/brilliant.org\/wiki\/normal-distribution\/\" rel=\"noreferrer noopener\" target=\"_blank\">The Normal probability distribution<\/a>&nbsp;is the continuous probability distribution for a real-valued random variable. Normal distribution, also called&nbsp;<strong><em>Gaussian distribution<\/em><\/strong>&nbsp;is arguably one of the most popular distribution functions that are commonly used in social and natural sciences for modeling purposes, for example, it is used to model people\u2019s height or test scores. Let&#8217;s assume a random variable X follows a Normal distribution, then its probability density function can be expressed as follows.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_18.png\" width=\"100%\"><\/p>\n\n\n\n<p>where the parameter&nbsp;<strong>?&nbsp;<\/strong>(mu)<strong>&nbsp;<\/strong>is the mean of the distribution also referred to as the&nbsp;<strong><em>location parameter<\/em><\/strong>, parameter&nbsp;<strong>?&nbsp;<\/strong>(sigma)<strong>&nbsp;<\/strong>is the standard deviation of the distribution also referred to as the&nbsp;<em>scale parameter<\/em>. The number&nbsp;<a href=\"https:\/\/www.mathsisfun.com\/numbers\/pi.html\" rel=\"noreferrer noopener\" target=\"_blank\"><strong>?<\/strong><\/a>&nbsp;(pi) is a mathematical constant approximately equal to 3.14.<\/p>\n\n\n\n<p><strong>Normal Distribution Mean &amp; Variance<\/strong><img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_19.png\" width=\"100%\"><\/p>\n\n\n\n<p>The figure below visualizes an example of Normal distribution with a mean 0 (<strong>? = 0<\/strong>) and standard deviation of 1 (<strong>? = 1<\/strong>), which is referred to as<strong>&nbsp;<em>Standard Normal&nbsp;<\/em><\/strong>distribution which is&nbsp;<em>symmetric.<\/em><img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_20.png\" width=\"100%\"><br>Image Source: The Author<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code># Random Generation of 1000 independent Normal samples\nimport numpy as np\nmu = 0\nsigma = 1\nN = 1000\nX = np.random.normal(mu,sigma,N)\n\n# Population distribution\nfrom scipy.stats import norm\nx_values = np.arange(-5,5,0.01)\ny_values = norm.pdf(x_values)\n#Sample histogram with Population distribution\nimport matplotlib.pyplot as plt\ncounts, bins, ignored = plt.hist(X, 30, density = True,color = 'purple',label = 'Sampling Distribution')\nplt.plot(x_values,y_values, color = 'y',linewidth = 2.5,label = 'Population Distribution')\nplt.title(\"Randomly generating 1000 obs from Normal distribution mu = 0 sigma = 1\")\nplt.ylabel(\"Probability\")\nplt.legend()\nplt.show()<\/code><\/pre>\n\n\n\n<h1 class=\"wp-block-heading\">Bayes Theorem<\/h1>\n\n\n\n<p>The Bayes Theorem or often called&nbsp;<strong><em>Bayes Law<\/em><\/strong>&nbsp;is arguably the most powerful rule of probability and statistics, named after famous English statistician and philosopher, Thomas Bayes.<img loading=\"lazy\" decoding=\"async\" width=\"220\" height=\"236\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_21.gif\"><br>Image Source:&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Thomas_Bayes\" rel=\"noreferrer noopener\" target=\"_blank\">Wikipedia<\/a><\/p>\n\n\n\n<p>Bayes theorem is a powerful probability law that brings the concept of&nbsp;<strong><em>subjectivity<\/em><\/strong>&nbsp;into the world of Statistics and Mathematics where everything is about facts. It describes the probability of an event, based on the prior information of&nbsp;<strong><em>conditions<\/em><\/strong>&nbsp;that might be related to that event. For instance, if the risk of getting Coronavirus or Covid-19 is known to increase with age, then Bayes Theorem allows the risk to an individual of a known age to be determined more accurately by conditioning it on the age than simply assuming that this individual is common to the population as a whole.<\/p>\n\n\n\n<p>The concept of&nbsp;<strong><em>conditional probability<\/em>,&nbsp;<\/strong>which plays a central role in Bayes theory, is a measure of the probability of an event happening, given that another event has already occurred. Bayes theorem can be described by the following expression where the X and Y stand for events X and Y, respectively:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_22.png\" width=\"100%\"><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em>Pr<\/em>&nbsp;(X|Y): the probability of event X occurring given that event or condition Y has occurred or is true<\/li>\n\n\n\n<li><em>Pr<\/em>&nbsp;(Y|X): the probability of event Y occurring given that event or condition X has occurred or is true<\/li>\n\n\n\n<li><em>Pr&nbsp;<\/em>(X) &amp;&nbsp;<em>Pr&nbsp;<\/em>(Y): the probabilities of observing events X and Y, respectively<\/li>\n<\/ul>\n\n\n\n<p>In the case of the earlier example, the probability of getting Coronavirus (event X) conditional on being at a certain age is&nbsp;<em>Pr<\/em>&nbsp;(X|Y), which is equal to the probability of being at a certain age given one got a Coronavirus,&nbsp;<em>Pr<\/em>&nbsp;(Y|X), multiplied with the probability of getting a Coronavirus,&nbsp;<em>Pr<\/em>&nbsp;(X), divided to the probability of being at a certain age.,&nbsp;<em>Pr<\/em>&nbsp;(Y).<\/p>\n\n\n\n<h1 class=\"wp-block-heading\">Linear Regression<\/h1>\n\n\n\n<p>Earlier, the concept of causation between variables was introduced, which happens when a variable has a direct impact on another variable. When the relationship between two variables is linear, then Linear Regression is a statistical method that can help to model the impact of a unit change in a variable,&nbsp;<strong><em>the<\/em><\/strong>&nbsp;<strong><em>independent variable<\/em><\/strong>&nbsp;on the values of another variable,&nbsp;<strong><em>the dependent variable<\/em><\/strong>.<\/p>\n\n\n\n<p>Dependent variables are often referred to as&nbsp;<strong><em>response variables<\/em><\/strong>&nbsp;or&nbsp;<strong><em>explained<\/em>&nbsp;<em>variables<\/em>,&nbsp;<\/strong>whereas independent variables are often referred to as&nbsp;<strong><em>regressors<\/em><\/strong>&nbsp;or&nbsp;<strong><em>explanatory variables<\/em><\/strong>. When the Linear Regression model is based on a single independent variable, then the model is called&nbsp;<strong><em>Simple Linear Regression<\/em><\/strong>&nbsp;and when the model is based on multiple independent variables, it\u2019s referred to as&nbsp;<strong><em>Multiple Linear Regression<\/em>.&nbsp;<\/strong>Simple Linear Regression can be described by the following expression:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_23.png\" width=\"100%\"><\/p>\n\n\n\n<p>where&nbsp;<strong>Y<\/strong>&nbsp;is the dependent variable,&nbsp;<strong>X<\/strong>&nbsp;is the independent variable which is part of the data,&nbsp;<strong>?0&nbsp;<\/strong>is the intercept which is unknown and constant,&nbsp;<strong>?1<\/strong>&nbsp;is the slope coefficient or a parameter corresponding to the variable X which is unknown and constant as well. Finally,&nbsp;<strong>u<\/strong>&nbsp;is the error term that the model makes when estimating the Y values. The main idea behind linear regression is to find the best-fitting straight line,&nbsp;<strong><em>the regression line,<\/em><\/strong>&nbsp;through a set of paired ( X, Y ) data. One example of the Linear Regression application is modeling the impact of&nbsp;<em>Flipper Length<\/em>&nbsp;on penguins\u2019&nbsp;<em>Body Mass,&nbsp;<\/em>which is visualized below.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_24.png\" width=\"100%\"><br>Image Source: The Author<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code># R code for the graph\ninstall.packages(\"ggplot2\")\ninstall.packages(\"palmerpenguins\")\nlibrary(palmerpenguins)\nlibrary(ggplot2)\nView(data(penguins))\nggplot(data = penguins, aes(x = flipper_length_mm,y = body_mass_g))+\n  geom_smooth(method = \"lm\", se = FALSE, color = 'purple')+\n  geom_point()+\n  labs(x=\"Flipper Length (mm)\",y=\"Body Mass (g)\")<\/code><\/pre>\n\n\n\n<p>Multiple Linear Regression with three independent variables can be described by the following expression:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_25.png\" width=\"100%\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Ordinary Least Squares<\/h2>\n\n\n\n<p>The ordinary least squares (OLS) is a method for estimating the unknown parameters such as ?0 and ?1<strong>&nbsp;<\/strong>in a linear regression model. The model is based on the principle of&nbsp;<strong><em>least squares<\/em>&nbsp;<\/strong>that<strong>&nbsp;<\/strong>minimizes the sum of squares of the differences between the observed dependent variable and its values predicted by the linear function of the independent variable, often referred to as&nbsp;<strong><em>fitted values<\/em><\/strong>. This difference between the real and predicted values of dependent variable Y is referred to as&nbsp;<strong><em>residual<\/em>&nbsp;<\/strong>and what OLS does, is minimizing the sum of squared residuals. This optimization problem results in the following OLS estimates for the unknown parameters ?0 and ?1 which are also known as&nbsp;<strong><em>coefficient estimates<\/em><\/strong>.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_26.png\" width=\"100%\"><\/p>\n\n\n\n<p>Once these parameters of the Simple Linear Regression model are estimated, the&nbsp;<strong><em>fitted values<\/em><\/strong><em>&nbsp;<\/em>of the response variable can be computed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_27.png\" width=\"100%\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Standard Error<\/h2>\n\n\n\n<p>The&nbsp;<strong><em>residuals<\/em><\/strong>&nbsp;or the estimated error terms can be determined as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_28.png\" width=\"100%\"><\/p>\n\n\n\n<p>It is important to keep in mind the difference between the error terms and residuals. Error terms are never observed, while the residuals are calculated from the data. The OLS estimates the error terms for each observation but not the actual error term. So, the true error variance is still unknown. Moreover, these estimates are subject to sampling uncertainty. What this means is that we will never be able to determine the exact estimate, the true value, of these parameters from sample data in an empirical application. However, we can estimate it by calculating the&nbsp;<strong><em>sample<\/em><\/strong><em>&nbsp;<\/em><strong><em>residual variance<\/em>&nbsp;<\/strong>by using the residuals as follows.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_29.png\" width=\"100%\"><\/p>\n\n\n\n<p>This estimate for the variance of sample residuals helps to estimate the variance of the estimated parameters which is often expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_30.png\" width=\"100%\"><\/p>\n\n\n\n<p>The squared root of this variance term is called&nbsp;<strong>the standard error<\/strong>&nbsp;of the estimate which is a key component in assessing the accuracy of the parameter estimates. It is used to calculating test statistics and confidence intervals. The standard error can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_31.png\" width=\"100%\"><\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>It is important to keep in mind the difference between the error terms and residuals. Error terms are never observed, while the residuals are calculated from the data.<\/p>\n<\/blockquote>\n\n\n\n<h2 class=\"wp-block-heading\">OLS Assumptions<\/h2>\n\n\n\n<p>OLS estimation method makes the following assumption which needs to be satisfied to get reliable prediction results:<\/p>\n\n\n\n<p><strong>A1: Linearity&nbsp;<\/strong>assumption states that the model is linear in parameters.<\/p>\n\n\n\n<p><strong>A2:<\/strong>&nbsp;<strong>Random<\/strong>&nbsp;<strong>Sample&nbsp;<\/strong>assumption states that all observations in the sample are randomly selected.<\/p>\n\n\n\n<p><strong>A3: Exogeneity&nbsp;<\/strong>assumption states that independent variables are uncorrelated with the error terms.<\/p>\n\n\n\n<p><strong>A4: Homoskedasticity&nbsp;<\/strong>assumption states that the variance of all error terms is constant.<\/p>\n\n\n\n<p><strong>A5: No Perfect Multi-Collinearity&nbsp;<\/strong>assumption states that none of the independent variables is constant and there are no exact linear relationships between the independent variables.<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>def runOLS(Y,X):\n\n   # OLS esyimation Y = Xb + e --&gt; beta_hat = (X'X)^-1(X'Y)\n   beta_hat = np.dot(np.linalg.inv(np.dot(np.transpose(X), X)), np.dot(np.transpose(X), Y))\n\n   # OLS prediction\n   Y_hat = np.dot(X,beta_hat)\n   residuals = Y-Y_hat\n   RSS = np.sum(np.square(residuals))\n   sigma_squared_hat = RSS\/(N-2)\n   TSS = np.sum(np.square(Y-np.repeat(Y.mean(),len(Y))))\n   MSE = sigma_squared_hat\n   RMSE = np.sqrt(MSE)\n   R_squared = (TSS-RSS)\/TSS\n\n   # Standard error of estimates:square root of estimate's variance\n   var_beta_hat = np.linalg.inv(np.dot(np.transpose(X),X))*sigma_squared_hat\n   \n   SE = &#91;]\n   t_stats = &#91;]\n   p_values = &#91;]\n   CI_s = &#91;]\n   \n   for i in range(len(beta)):\n       #standard errors\n       SE_i = np.sqrt(var_beta_hat&#91;i,i])\n       SE.append(np.round(SE_i,3))\n\n        #t-statistics\n        t_stat = np.round(beta_hat&#91;i,0]\/SE_i,3)\n        t_stats.append(t_stat)\n\n        #p-value of t-stat p&#91;|t_stat| &gt;= t-treshhold two sided] \n        p_value = t.sf(np.abs(t_stat),N-2) * 2\n        p_values.append(np.round(p_value,3))\n\n        #Confidence intervals = beta_hat -+ margin_of_error\n        t_critical = t.ppf(q =1-0.05\/2, df = N-2)\n        margin_of_error = t_critical*SE_i\n        CI = &#91;np.round(beta_hat&#91;i,0]-margin_of_error,3), np.round(beta_hat&#91;i,0]+margin_of_error,3)]\n        CI_s.append(CI)\n        return(beta_hat, SE, t_stats, p_values,CI_s, \n               MSE, RMSE, R_squared)<\/code><\/pre>\n\n\n\n<h1 class=\"wp-block-heading\">Parameter Properties<\/h1>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>Under the assumption that the OLS criteria A1 \u2014 A5 are satisfied, the OLS estimators of coefficients \u03b20 and \u03b21 are&nbsp;<strong>BLUE<\/strong>&nbsp;and&nbsp;<strong>Consistent<\/strong>.<\/p>\n\n\n\n<p><strong>Gauss-Markov theorem<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p>This theorem highlights the properties of OLS estimates where the term&nbsp;<strong><em>BLUE<\/em><\/strong>&nbsp;stands for&nbsp;<strong><em>Best Linear Unbiased Estimator<\/em><\/strong>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Bias<\/h2>\n\n\n\n<p>The&nbsp;<strong>bias<\/strong>&nbsp;of an estimator is the difference between its expected value and the true value of the parameter being estimated and can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_32.png\" width=\"100%\"><\/p>\n\n\n\n<p>When we state that the estimator is&nbsp;<strong><em>unbiased<\/em><\/strong>&nbsp;what we mean is that the bias is equal to zero, which implies that the expected value of the estimator is equal to the true parameter value, that is:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_33.png\" width=\"100%\"><\/p>\n\n\n\n<p>Unbiasedness does not guarantee that the obtained estimate with any particular sample is equal or close to ?. What it means is that, if one&nbsp;<strong><em>repeatedly<\/em><\/strong>&nbsp;draws random samples from the population and then computes the estimate each time, then the average of these estimates would be equal or very close to \u03b2.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Efficiency<\/h2>\n\n\n\n<p>The term&nbsp;<strong><em>Best<\/em><\/strong>&nbsp;in the Gauss-Markov theorem relates to the variance of the estimator and is referred to as&nbsp;<strong><em>efficiency<\/em><\/strong><em>.&nbsp;<\/em>A parameter can have multiple estimators but the one with the lowest variance is called efficient<strong>.<\/strong><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Consistency<\/h2>\n\n\n\n<p>The term consistency goes hand in hand with the terms&nbsp;<strong><em>sample size<\/em><\/strong>&nbsp;and&nbsp;<strong><em>convergence<\/em><\/strong>. If the estimator converges to the true parameter as the sample size becomes very large, then this estimator is said to be consistent, that is:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_34.png\" width=\"100%\"><\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>Under the assumption that the OLS criteria A1 \u2014 A5 are satisfied, the OLS estimators of coefficients \u03b20 and \u03b21 are&nbsp;<strong>BLUE<\/strong>&nbsp;and&nbsp;<strong>Consistent<\/strong>.<br><strong>Gauss-Markov Theorem<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p>All these properties hold for OLS estimates as summarized in the Gauss-Markov theorem. In other words, OLS estimates have the smallest variance, they are unbiased, linear in parameters, and are consistent. These properties can be mathematically proven by using the OLS assumptions made earlier.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\">Confidence Intervals<\/h1>\n\n\n\n<p>The Confidence Interval is the range that contains the true population parameter with a certain pre-specified probability, referred to as the&nbsp;<strong><em>confidence level<\/em><\/strong><em>&nbsp;<\/em>of the experiment, and it is obtained by using the sample results and the<strong>&nbsp;<em>margin of error<\/em><\/strong>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Margin of Error<\/h2>\n\n\n\n<p>The margin of error is the difference between the sample results and based on what the result would have been if one had used the entire population.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Confidence Level<\/h2>\n\n\n\n<p>The Confidence Level describes the level of certainty in the experimental results. For example, a 95% confidence level means that if one were to perform the same experiment repeatedly for 100 times, then 95 of those 100 trials would lead to similar results. Note that the confidence level is defined before the start of the experiment because it will affect how big the margin of error will be at the end of the experiment.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Confidence Interval for OLS Estimates<\/h2>\n\n\n\n<p>As it was mentioned earlier, the OLS estimates of the Simple Linear Regression, the estimates for intercept ?0 and slope coefficient ?1, are subject to sampling uncertainty. However, we can construct CI\u2019s<em>&nbsp;<\/em>for these parameters which will contain the true value of these parameters in 95% of all samples. That is, 95% confidence interval for ? can be interpreted as follows:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>The confidence interval is the set of values for which a hypothesis test cannot be rejected to the level of 5%.<\/li>\n\n\n\n<li>The confidence interval has a 95% chance to contain the true value of ?.<\/li>\n<\/ul>\n\n\n\n<p>95% confidence interval of OLS estimates can be constructed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_35.png\" width=\"100%\"><\/p>\n\n\n\n<p>which is based on the parameter estimate, the standard error of that estimate, and the value 1.96 representing the margin of error corresponding to the 5% rejection rule. This value is determined using the&nbsp;<a href=\"https:\/\/www.google.com\/url?sa=i&amp;url=https%3A%2F%2Ffreakonometrics.hypotheses.org%2F9404&amp;psig=AOvVaw2IcJrhGrWbt9504WTCWBwW&amp;ust=1618940099743000&amp;cd=vfe&amp;ved=0CAIQjRxqFwoTCOjR4v7rivACFQAAAAAdAAAAABAI\" rel=\"noreferrer noopener\" target=\"_blank\">Normal Distribution table<\/a>, which will be discussed later on in this article. Meanwhile, the following figure illustrates the idea of 95% CI:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_36.png\" width=\"100%\"><br>Image Source:&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Standard_deviation#\/media\/File:Standard_deviation_diagram.svg\" rel=\"noreferrer noopener\" target=\"_blank\">Wikipedia<\/a><\/p>\n\n\n\n<p>Note that the confidence interval depends on the sample size as well, given that it is calculated using the standard error which is based on sample size.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>The confidence level is defined before the start of the experiment because it will affect how big the margin of error will be at the end of the experiment.<\/p>\n<\/blockquote>\n\n\n\n<h1 class=\"wp-block-heading\">Statistical Hypothesis testing<\/h1>\n\n\n\n<p>Testing a hypothesis in Statistics is a way to test the results of an experiment or survey to determine how meaningful they the results are. Basically, one is testing whether the obtained results are valid by figuring out the odds that the results have occurred by chance. If it is the letter, then the results are not reliable and neither is the experiment. Hypothesis Testing is part of the&nbsp;<strong><em>Statistical Inference<\/em><\/strong>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Null and Alternative Hypothesis<\/h2>\n\n\n\n<p>Firstly, you need to determine the thesis you wish to test, then you need to formulate the&nbsp;<strong><em>Null Hypothesis<\/em><\/strong>&nbsp;and the&nbsp;<strong><em>Alternative Hypothesis<\/em>.&nbsp;<\/strong>The test can have two possible outcomes and based on the statistical results you can either reject the stated hypothesis or accept it. As a rule of thumb, statisticians tend to put the version or formulation of the hypothesis under the Null Hypothesis that<em>&nbsp;<\/em>that needs to be rejected<em>,&nbsp;<\/em>whereas the acceptable and desired version is stated under the Alternative Hypothesis<em>.<\/em><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Statistical significance<\/h2>\n\n\n\n<p>Let\u2019s look at the earlier mentioned example where the Linear Regression model was used to investigating whether a penguins\u2019&nbsp;<em>Flipper Length<\/em>, the independent variable, has an impact on&nbsp;<em>Body Mass,&nbsp;<\/em>the dependent variable. We can formulate this model with the following statistical expression:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_37.png\" width=\"100%\"><\/p>\n\n\n\n<p>Then, once the OLS estimates of the coefficients are estimated, we can formulate the following Null and Alternative Hypothesis to test whether the Flipper Length has a<strong>&nbsp;<em>statistically significant<\/em>&nbsp;<\/strong>impact on the Body Mass:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_38.png\" width=\"100%\"><\/p>\n\n\n\n<p>where H0 and H1 represent Null Hypothesis and Alternative Hypothesis, respectively. Rejecting the Null Hypothesis would mean that a one-unit increase in&nbsp;<em>Flipper Length<\/em>&nbsp;has a direct impact on the&nbsp;<em>Body Mass<\/em>. Given that the parameter estimate of ?1 is describing this impact of the independent variable,&nbsp;<em>Flipper Length<\/em>, on the dependent variable,&nbsp;<em>Body Mass.<\/em>&nbsp;This hypothesis can be reformulated as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_39.png\" width=\"100%\"><\/p>\n\n\n\n<p>where H0 states that the parameter estimate of ?1 is equal to 0, that is<em>&nbsp;Flipper Length<\/em>&nbsp;effect on&nbsp;<em>Body Mass&nbsp;<\/em>is&nbsp;<strong><em>statistically insignificant<\/em><\/strong>&nbsp;whereas<em>&nbsp;<\/em>H0 states that the parameter estimate of ?1 is not equal to 0 suggesting that&nbsp;<em>Flipper Length<\/em>&nbsp;effect on&nbsp;<em>Body Mass<\/em>&nbsp;is&nbsp;<strong><em>statistically significant<\/em><\/strong><em>.<\/em><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Type I and Type II Errors<\/h2>\n\n\n\n<p>When performing Statistical Hypothesis Testing one needs to consider two conceptual types of errors: Type I error and Type II error. The Type I error occurs when the Null is wrongly rejected whereas the Type II error occurs when the Null Hypothesis is wrongly not rejected. A confusion<a href=\"https:\/\/www.dataschool.io\/simple-guide-to-confusion-matrix-terminology\/\" rel=\"noreferrer noopener\" target=\"_blank\">&nbsp;matrix<\/a>&nbsp;can help to clearly visualize the severity of these two types of errors.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>As a rule of thumb, statisticians tend to put the version the hypothesis under the&nbsp;<em>Null Hypothesis&nbsp;<\/em>that<em>&nbsp;<\/em>that needs to be rejected<em>,&nbsp;<\/em>whereas the acceptable and desired version is stated under the&nbsp;<em>Alternative Hypothesis.<\/em><\/p>\n<\/blockquote>\n\n\n\n<h1 class=\"wp-block-heading\">Statistical Tests<\/h1>\n\n\n\n<p>Once the Null and the Alternative Hypotheses are stated and the test assumptions are defined, the next step is to determine which statistical test is appropriate and to calculate the<em>&nbsp;<\/em><strong><em>test statistic<\/em><\/strong>. Whether or not to reject or not reject the Null can be determined by comparing the test statistic with the<strong>&nbsp;<em>critical value<\/em>.&nbsp;<\/strong>This comparison shows whether or not the observed test statistic is more extreme than the defined critical value and it can have two possible results:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>The test statistic is more extreme than the critical value ? the null hypothesis can be rejected<\/li>\n\n\n\n<li>The test statistic is not as extreme as the critical value ? the null hypothesis cannot be rejected<\/li>\n<\/ul>\n\n\n\n<p>The critical value is based on a prespecified&nbsp;<strong><em>significance level<\/em>&nbsp;?<\/strong>&nbsp;(usually chosen to be equal to 5%) and the type of probability distribution the test statistic follows. The critical value divides the area under this probability distribution curve into the&nbsp;<strong><em>rejection region(s)<\/em><\/strong>&nbsp;and&nbsp;<strong><em>non-rejection region<\/em><\/strong>. There are numerous statistical tests used to test various hypotheses. Examples of Statistical tests are&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Student%27s_t-test\" rel=\"noreferrer noopener\" target=\"_blank\">Student\u2019s t-test<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/F-test\" rel=\"noreferrer noopener\" target=\"_blank\">F-test<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Chi-squared_test\" rel=\"noreferrer noopener\" target=\"_blank\">Chi-squared test<\/a>,&nbsp;<a href=\"https:\/\/www.stata.com\/support\/faqs\/statistics\/durbin-wu-hausman-test\/\" rel=\"noreferrer noopener\" target=\"_blank\">Durbin-Hausman-Wu Endogeneity test<\/a>, W<a href=\"https:\/\/en.wikipedia.org\/wiki\/White_test#:~:text=In%20statistics%2C%20the%20White%20test,by%20Halbert%20White%20in%201980.\" rel=\"noreferrer noopener\" target=\"_blank\">hite Heteroskedasticity test<\/a>. In this article, we will look at two of these statistical tests.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>The Type I error occurs when the Null is wrongly rejected whereas the Type II error occurs when the Null Hypothesis is wrongly not rejected.<\/p>\n<\/blockquote>\n\n\n\n<h2 class=\"wp-block-heading\">Student\u2019s t-test<\/h2>\n\n\n\n<p>One of the simplest and most popular statistical tests is the Student\u2019s t-test. which can be used for testing various hypotheses especially when dealing with a hypothesis where the main area of interest is to find evidence for the statistically significant effect of a&nbsp;<strong><em>single variable<\/em><\/strong><em>.&nbsp;<\/em>The<strong>&nbsp;<\/strong>test statistics of the t-test follows&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Student%27s_t-distribution\" rel=\"noreferrer noopener\" target=\"_blank\"><strong><em>Student\u2019s t distribution<\/em><\/strong><\/a>&nbsp;and can be determined as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_40.png\" width=\"100%\"><\/p>\n\n\n\n<p>where h0 in the nominator is the value against which the parameter estimate is being tested. So, the t-test statistics are equal to the parameter estimate minus the hypothesized value divided by the standard error of the coefficient estimate. In the earlier stated hypothesis, where we wanted to test whether Flipper Length has a statistically significant impact on Body Mass or not. This test can be performed using a t-test and the h0 is in that case equal to the 0 since the slope coefficient estimate is tested against value 0.<\/p>\n\n\n\n<p>There are two versions of the t-test: a&nbsp;<strong><em>two-sided t-test<\/em>&nbsp;<\/strong>and a&nbsp;<strong><em>one-sided t-test<\/em><\/strong>. Whether you need the former or the latter version of the test depends entirely on the hypothesis that you want to test.<\/p>\n\n\n\n<p>The two-sided<strong>&nbsp;<\/strong>or<strong>&nbsp;<em>two-tailed t-test<\/em>&nbsp;<\/strong>can be used when the hypothesis is testing&nbsp;<em>equal<\/em>&nbsp;versus&nbsp;<em>not equal<\/em>&nbsp;relationship under the Null and Alternative Hypotheses that is similar to the following example:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_41.png\" width=\"100%\"><\/p>\n\n\n\n<p>The two-sided t-test has<strong>&nbsp;<em>two rejection regions<\/em><\/strong>&nbsp;as visualized in the figure below:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_42.png\" width=\"100%\"><br>Image Source:&nbsp;<a href=\"https:\/\/www.geo.fu-berlin.de\/en\/v\/soga\/Basics-of-statistics\/Hypothesis-Tests\/Introduction-to-Hypothesis-Testing\/Critical-Value-and-the-p-Value-Approach\/index.html\" rel=\"noreferrer noopener\" target=\"_blank\"><em>Hartmann, K., Krois, J., Waske, B. (2018): E-Learning Project SOGA: Statistics and Geospatial Data Analysis. Department of Earth Sciences, Freie Universitaet Berlin<\/em><\/a><\/p>\n\n\n\n<p>In this version of the t-test, the Null is rejected if the calculated t-statistics is either too small or too large.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_43.png\" width=\"100%\"><\/p>\n\n\n\n<p>Here, the test statistics are compared to the critical values based on the sample size and the chosen significance level. To determine the exact value of the cutoff point,<a href=\"https:\/\/www.google.com\/search?q=t-table+two+sided&amp;client=safari&amp;rls=en&amp;sxsrf=ALeKk01KSlU3EEtBeMcXPuh13ud42kRCWw%3A1618592162824&amp;tbm=isch&amp;ictx=1&amp;fir=ZGAb8l8KaBNJiM%252CZaqfSsY36WrUvM%252C_&amp;vet=1&amp;usg=AI4_-kSaUb_tv_3EBZQRhYaQVYYaJ1uBHQ&amp;sa=X&amp;ved=2ahUKEwjBtZrXnYPwAhWHgv0HHQPmASUQ9QF6BAgSEAE&amp;biw=1981&amp;bih=1044#imgrc=ZGAb8l8KaBNJiM\" rel=\"noreferrer noopener\" target=\"_blank\">&nbsp;the two-sided t-distribution table<\/a>&nbsp;can be used.<\/p>\n\n\n\n<p>The one-sided or&nbsp;<strong><em>one-tailed t-test&nbsp;<\/em><\/strong>can be used when the hypothesis is testing&nbsp;<em>positive\/negative<\/em>&nbsp;versus&nbsp;<em>negative\/positive&nbsp;<\/em>relationship under the Null and Alternative Hypotheses that is similar to the following examples:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_44.png\" width=\"100%\"><\/p>\n\n\n\n<p>One-sided t-test has a&nbsp;<strong><em>single<\/em><\/strong><em>&nbsp;<\/em><strong><em>rejection region<\/em>&nbsp;<\/strong>and depending<strong>&nbsp;<\/strong>on the hypothesis side the rejection region is either on the left-hand side or the right-hand side as visualized in the figure below:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_45.png\" width=\"100%\"><br>Image Source:&nbsp;<a href=\"https:\/\/www.geo.fu-berlin.de\/en\/v\/soga\/Basics-of-statistics\/Hypothesis-Tests\/Introduction-to-Hypothesis-Testing\/Critical-Value-and-the-p-Value-Approach\/index.html\" rel=\"noreferrer noopener\" target=\"_blank\"><em>Hartmann, K., Krois, J., Waske, B. (2018): E-Learning Project SOGA: Statistics and Geospatial Data Analysis. Department of Earth Sciences, Freie Universitaet Berlin<\/em><\/a><\/p>\n\n\n\n<p>In this version of the t-test, the Null is rejected if the calculated t-statistics is smaller\/larger than the critical value.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_46.png\" width=\"100%\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">F-test<\/h2>\n\n\n\n<p>F-test is another very popular statistical test often used to test hypotheses testing&nbsp;<em>a&nbsp;<\/em><strong><em>joint statistical significance of multiple variables<\/em><\/strong><em>.&nbsp;<\/em>This is the case when you want to test whether multiple independent variables have a statistically significant impact on a dependent variable. Following is an example of a statistical hypothesis that can be tested using the F-test:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_47.png\" width=\"100%\"><\/p>\n\n\n\n<p>where the Null states that the three variables corresponding to these coefficients are jointly statistically insignificant and the Alternative states that these three variables are jointly statistically significant. The test statistics of the F-test follows&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/F-distribution\" rel=\"noreferrer noopener\" target=\"_blank\">F distribution<\/a>&nbsp;and can be determined as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_48.png\" width=\"100%\"><\/p>\n\n\n\n<p>where the SSRrestricted is&nbsp;<em>the<\/em><strong><em>&nbsp;sum of squared residuals<\/em>&nbsp;<\/strong>of the&nbsp;<strong><em>restricted<\/em>&nbsp;<em>model&nbsp;<\/em><\/strong>which is the same model excluding from the data the target variables stated as insignificant under the Null<em>,&nbsp;<\/em>the SSRunrestricted is the sum of squared residuals of the&nbsp;<strong><em>unrestricted<\/em>&nbsp;<em>model<\/em><\/strong><em>&nbsp;<\/em>which is the model that includes all variables, the q represents the number of variables that are being jointly tested for the insignificance under the Null, N is the sample size, and the k is the total number of variables in the unrestricted model. SSR values are provided next to the parameter estimates after running the OLS regression and the same holds for the F-statistics as well. Following is an example of MLR model output where the SSR and F-statistics values are marked.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_49.png\" width=\"100%\"><br>Image Source:<a href=\"https:\/\/www.uio.no\/studier\/emner\/sv\/oekonomi\/ECON4150\/v18\/lecture7_ols_multiple_regressors_hypothesis_tests.pdf\" rel=\"noreferrer noopener\" target=\"_blank\">&nbsp;Stock and Whatson<\/a><\/p>\n\n\n\n<p>F-test has&nbsp;<strong>a single rejection region&nbsp;<\/strong>as visualized below:<img loading=\"lazy\" decoding=\"async\" width=\"401\" height=\"223\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_50.jpeg\"><br>Image Source:&nbsp;<a href=\"https:\/\/www.statisticshowto.com\/probability-and-statistics\/f-statistic-value-test\/\" rel=\"noreferrer noopener\" target=\"_blank\"><em>U of Michigan<\/em><\/a><\/p>\n\n\n\n<p>If the calculated F-statistics is bigger than the critical value, then the Null can be rejected which suggests that the independent variables are jointly statistically significant. The rejection rule can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_51.png\" width=\"100%\"><\/p>\n\n\n\n<h1 class=\"wp-block-heading\">P-Values<\/h1>\n\n\n\n<p>Another quick way to determine whether to reject or to support the Null Hypothesis is by using&nbsp;<strong><em>p-values<\/em><\/strong>. The p-value is the probability of the condition under the Null occurring. Stated differently, the p-value is the probability, assuming the null hypothesis is true, of observing a result at least as extreme as the test statistic. The smaller the p-value, the stronger is the evidence against the Null Hypothesis, suggesting that it can be rejected.<\/p>\n\n\n\n<p>The interpretation of a&nbsp;<em>p<\/em>-value is dependent on the chosen significance level. Most often, 1%, 5%, or 10% significance levels are used to interpret the p-value. So, instead of using the t-test and the F-test, p-values of these test statistics can be used to test the same hypotheses.<\/p>\n\n\n\n<p>The following figure shows a sample output of an OLS regression with two independent variables. In this table, the p-value of the t-test, testing the statistical significance of&nbsp;<em>class_size<\/em>&nbsp;variable\u2019s parameter estimate, and the p-value of the F-test, testing the joint statistical significance of the&nbsp;<em>class_size,<\/em>&nbsp;and&nbsp;<em>el_pct&nbsp;<\/em>variables parameter estimates, are underlined.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_52.png\" width=\"100%\"><br>Image Source:<a href=\"https:\/\/www.uio.no\/studier\/emner\/sv\/oekonomi\/ECON4150\/v18\/lecture7_ols_multiple_regressors_hypothesis_tests.pdf\" rel=\"noreferrer noopener\" target=\"_blank\">&nbsp;Stock and Whatson<\/a><\/p>\n\n\n\n<p>The p-value corresponding to the&nbsp;<em>class_size<\/em>&nbsp;variable is 0.011 and when comparing this value to the significance levels 1% or 0.01 , 5% or 0.05, 10% or 0.1, then the following conclusions can be made:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>0.011 &gt; 0.01 ? Null of the t-test can\u2019t be rejected at 1% significance level<\/li>\n\n\n\n<li>0.011 &lt; 0.05 ? Null of the t-test can be rejected at 5% significance level<\/li>\n\n\n\n<li>0.011 &lt; 0.10 ?Null of the t-test can be rejected at 10% significance level<\/li>\n<\/ul>\n\n\n\n<p>So, this p-value suggests that the coefficient of the&nbsp;<em>class_size<\/em>&nbsp;variable is statistically significant at 5% and 10% significance levels. The p-value corresponding to the F-test<em>&nbsp;<\/em>is 0.0000 and since 0 is smaller than all three cutoff values; 0.01, 0.05, 0.10, we can conclude that the Null of the F-test can be rejected in all three cases. This suggests that the coefficients of&nbsp;<em>class_size<\/em>&nbsp;and&nbsp;<em>el_pct<\/em>&nbsp;variables are jointly statistically significant at 1%, 5%, and 10% significance levels.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Limitation of p-values<\/h2>\n\n\n\n<p>Although, using p-values has many benefits but it has also limitations<strong>.&nbsp;<\/strong>Namely, the p-value depends on both the magnitude of association and the sample size. If the magnitude of the effect is small and statistically insignificant, the p-value might still show a&nbsp;<strong><em>significant impact<\/em><\/strong><em>&nbsp;<\/em>because the large sample size is large. The opposite can occur as well, an effect can be large, but fail to meet the p&lt;0.01, 0.05, or 0.10 criteria if the sample size is small.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\">Inferential Statistics<\/h1>\n\n\n\n<p>Inferential statistics uses sample data to make reasonable judgments about the population from which the sample data originated. It\u2019s used to investigate the relationships between variables within a sample and make predictions about how these variables will relate to a larger population.<\/p>\n\n\n\n<p>Both&nbsp;<strong><em>Law of Large Numbers (LLN)<\/em><\/strong>&nbsp;and&nbsp;<strong><em>Central Limit Theorem (CLM)<\/em><\/strong>&nbsp;have a significant role in Inferential statistics because they show that the experimental results hold regardless of what shape the original population distribution was when the data is large enough. The more data is gathered, the more accurate the statistical inferences become, hence, the more accurate parameter estimates are generated.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Law of Large Numbers (LLN)<\/h2>\n\n\n\n<p>Suppose&nbsp;<strong>X1, X2, . . . , Xn<\/strong>&nbsp;are all independent random variables with the same underlying distribution, also called independent identically-distributed or i.i.d, where all X\u2019s have the same mean&nbsp;<strong>?<\/strong>&nbsp;and standard deviation&nbsp;<strong>?<\/strong>. As the sample size grows, the probability that the average of all X\u2019s is equal to the mean ? is equal to 1. The Law of Large Numbers can be summarized as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_53.png\" width=\"100%\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Central Limit Theorem (CLM)<\/h2>\n\n\n\n<p>Suppose&nbsp;<strong>X1, X2, . . . , Xn<\/strong>&nbsp;are all independent random variables with the same underlying distribution, also called independent identically-distributed or i.i.d, where all X\u2019s have the same mean&nbsp;<strong>?<\/strong>&nbsp;and standard deviation&nbsp;<strong>?<\/strong>. As the sample size grows, the probability distribution of X&nbsp;<strong><em>converges in the distribution<\/em><\/strong>&nbsp;in Normal distribution with mean&nbsp;<strong>?&nbsp;<\/strong>and variance&nbsp;<strong>?-<\/strong>squared. The Central Limit Theorem can be summarized as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_54.png\" width=\"100%\"><\/p>\n\n\n\n<p>Stated differently, when you have a population with mean ? and standard deviation ? and you take sufficiently large random samples from that population with replacement, then the distribution of the sample means will be approximately normally distributed.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\">Dimensionality Reduction Techniques<\/h1>\n\n\n\n<p>Dimensionality reduction is the transformation of data from a&nbsp;<strong><em>high-dimensional space<\/em><\/strong>&nbsp;into a&nbsp;<strong><em>low-dimensional space&nbsp;<\/em><\/strong>such that this low-dimensional representation of the data still contains the meaningful properties of the original data as much as possible.<\/p>\n\n\n\n<p>With the increase in popularity in Big Data, the demand for these dimensionality reduction techniques, reducing the amount of unnecessary data and features, increased as well. Examples of popular dimensionality reduction techniques are&nbsp;<a href=\"https:\/\/builtin.com\/data-science\/step-step-explanation-principal-component-analysis\" rel=\"noreferrer noopener\" target=\"_blank\">Principle Component Analysis<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Factor_analysis\" rel=\"noreferrer noopener\" target=\"_blank\">Factor Analysis<\/a>,&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Canonical_correlation\" rel=\"noreferrer noopener\" target=\"_blank\">Canonical Correlation<\/a>,&nbsp;<a href=\"https:\/\/towardsdatascience.com\/understanding-random-forest-58381e0602d2\" rel=\"noreferrer noopener\" target=\"_blank\">Random Forest<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Principle Component Analysis (PCA)<\/h2>\n\n\n\n<p>Principal Component Analysis or PCA is a dimensionality reduction technique that is very often used to reduce the dimensionality of large data sets, by transforming a large set of variables into a smaller set that still contains most of the information or the variation in the original large dataset.<\/p>\n\n\n\n<p>Let\u2019s assume we have a data X with p variables; X1, X2, \u2026., Xp with&nbsp;<strong><em>eigenvectors<\/em><\/strong>&nbsp;e1, \u2026, ep, and&nbsp;<strong><em>eigenvalues<\/em><\/strong>&nbsp;?1,\u2026, ?p. Eigenvalues show the variance explained by a particular data field out of the total variance. The idea behind PCA is to create new (independent) variables, called Principal Components, that are a linear combination of the existing variable. The i<em>th<\/em>&nbsp;principal component can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_55.png\" width=\"100%\"><\/p>\n\n\n\n<p>Then using&nbsp;<strong>Elbow Rule<\/strong>&nbsp;or&nbsp;<a href=\"https:\/\/docs.displayr.com\/wiki\/Kaiser_Rule\" rel=\"noreferrer noopener\" target=\"_blank\"><strong>Kaiser Rule<\/strong><\/a>, you can determine the number of principal components that optimally summarize the data without losing too much information. It is also important to look at&nbsp;<strong><em>the proportion of total variation (PRTV)<\/em>&nbsp;<\/strong>that is explained by each principal component to decide whether it is beneficial to include or to exclude it. PRTV for the i<em>th<\/em>&nbsp;principal component can be calculated using eigenvalues as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_56.png\" width=\"100%\"><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Elbow Rule<\/h2>\n\n\n\n<p>The elbow rule or the elbow method is a heuristic approach that is used to determine the number of optimal principal components from the PCA results. The idea behind this method is to plot&nbsp;<em>the explained variation&nbsp;<\/em>as a function of the number of components and pick the elbow of the curve as the number of optimal principal components. Following is an example of such a scatter plot where the PRTV (Y-axis) is plotted on the number of principal components (X-axis). The elbow corresponds to the X-axis value 2, which suggests that the number of optimal principal components is 2.<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_57.png\" width=\"100%\"><br>Image Source:&nbsp;<a href=\"https:\/\/raw.githubusercontent.com\/TatevKaren\/Multivariate-Statistics\/main\/Elbow_rule_%25varc_explained.png\" rel=\"noreferrer noopener\" target=\"_blank\">Multivariate Statistics Github<\/a><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Factor Analysis (FA)<\/h2>\n\n\n\n<p>Factor analysis or FA is another statistical method for dimensionality reduction. It is one of the most commonly used inter-dependency techniques and is used when the relevant set of variables shows a systematic inter-dependence and the objective is to find out the latent factors that create a commonality. Let\u2019s assume we have a data X with p variables; X1, X2, \u2026., Xp. FA model can be expressed as follows:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_58.png\" width=\"100%\"><\/p>\n\n\n\n<p>where X is a [p x N] matrix of p variables and N observations, \u00b5 is [p x N] population mean matrix, A is [p x k] common&nbsp;<strong><em>factor loadings matrix<\/em><\/strong>, F [k x N] is the matrix of common factors and u [pxN] is the matrix of specific factors. So, put it differently, a factor model is as a series of multiple regressions, predicting each of the variables Xi from the values of the unobservable common factors fi:<img decoding=\"async\" alt=\"Fundamentals Of Statistics For Data Scientists and Analysts\" src=\"https:\/\/www.kdnuggets.com\/wp-content\/uploads\/fundamentals-statistics-data-scientists-analysts_59.png\" width=\"100%\"><\/p>\n\n\n\n<p>Each variable has k of its own common factors, and these are related to the observations via factor loading matrix for a single observation as follows: In factor analysis, the&nbsp;<strong><em>factors<\/em><\/strong>&nbsp;are calculated to&nbsp;<strong><em>maximize<\/em>&nbsp;<em>between-group variance<\/em><\/strong>&nbsp;while&nbsp;<strong><em>minimizing in-group varianc<\/em><\/strong><em>e<\/em>. They are factors because they group the underlying variables. Unlike the PCA, in FA the data needs to be normalized, given that FA assumes that the dataset follows Normal Distribution.<\/p>\n\n\n\n<p>&nbsp;<br>&nbsp;<br><strong><a href=\"https:\/\/www.linkedin.com\/in\/tatev-karen-aslanyan\/\" target=\"_blank\" rel=\"noreferrer noopener\">Tatev Karen Aslanyan<\/a><\/strong>&nbsp;is an experienced full-stack data scientist with a focus on Machine Learning and AI. She is also the co-founder of&nbsp;<a href=\"https:\/\/lunartech.ai\/course-overview\/\" target=\"_blank\" rel=\"noreferrer noopener\">LunarTech<\/a>, an online tech educational platform, and the creator of The Ultimate Data Science Bootcamp.Tatev Karen, with Bachelor and Masters in Econometrics and Management Science, has grown in the field of Machine Learning and AI, focusing on Recommender Systems and NLP, supported by her scientific research and published papers. Following five years of teaching, Tatev is now channeling her passion into LunarTech, helping shape the future of data science.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Key statistical concepts for your data science or data analysis journey. As Karl Pearson, a British mathematician has once stated,&nbsp;Statistics&nbsp;is the grammar of science and this holds especially for Computer&#8230; <a class=\"read-more-link\" href=\"https:\/\/tbekk.com\/devstream\/2023\/08\/08\/fundamentals-of-statistics-for-data-scientists-and-analysts\/\">Read more &raquo;<\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[51,15],"tags":[151,16,41],"class_list":["post-808","post","type-post","status-publish","format-standard","hentry","category-article","category-statistics","tag-data-analysis","tag-data-science","tag-statistics"],"_links":{"self":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/808","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/comments?post=808"}],"version-history":[{"count":1,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/808\/revisions"}],"predecessor-version":[{"id":809,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/808\/revisions\/809"}],"wp:attachment":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/media?parent=808"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/categories?post=808"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/tags?post=808"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}