{"id":933,"date":"2024-06-23T07:03:36","date_gmt":"2024-06-23T07:03:36","guid":{"rendered":"https:\/\/tbekk.com\/devstream\/?p=933"},"modified":"2024-06-23T07:03:36","modified_gmt":"2024-06-23T07:03:36","slug":"the-mathematics-of-neural-networks-a-complete-example","status":"publish","type":"post","link":"https:\/\/tbekk.com\/devstream\/2024\/06\/23\/the-mathematics-of-neural-networks-a-complete-example\/","title":{"rendered":"The Mathematics of Neural Networks &#8211; A complete example"},"content":{"rendered":"\n<hr class=\"wp-block-separator has-text-color has-medium-gray-color has-alpha-channel-opacity has-medium-gray-background-color has-background is-style-wide\"\/>\n\n\n\n<ul class=\"wp-block-list\">\n<li><em><strong>Link:<\/strong><\/em> <a href=\"https:\/\/medium.com\/@SSiddhant\/the-mathematics-of-neural-networks-a-complete-example-65f2b12cdea2\"><em>Medium<\/em><\/a><\/li>\n\n\n\n<li><em><strong>Author:<\/strong><\/em> <a href=\"https:\/\/medium.com\/@SSiddhant?source=post_page-----65f2b12cdea2--------------------------------\"><em>Siddhant<\/em><\/a><\/li>\n\n\n\n<li><em><strong>Publication date:<\/strong><\/em> <em>June, 2024<\/em><\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-text-color has-medium-gray-color has-alpha-channel-opacity has-medium-gray-background-color has-background is-style-wide\"\/>\n\n\n\n<p id=\"22d4\">Neural Networks are a method of artificial intelligence in which computers are taught to process data in a way similar to the human brain. Neural networks learn through being fed multiple instances of data as input, predicting an output, finding the error from the actual answer to the machine\u2019s answer, and then fine-tuning it\u2019s weights to reduce this error.<\/p>\n\n\n\n<p id=\"1208\">Whilst a neural network might seem very complex, it is actually a clever utilisation of linear algebra and multivariate calculus. This article aims to go through a full iteration of the maths that undermines a neural network.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"b7dc\">Assumptions &amp; Pre-Knowledge<\/h1>\n\n\n\n<p id=\"e0f7\">Neural networks require a solid understanding of a college-level of&nbsp;<a href=\"https:\/\/www.khanacademy.org\/math\/multivariable-calculus\" rel=\"noreferrer noopener\" target=\"_blank\">calculus<\/a>&nbsp;and&nbsp;<a href=\"https:\/\/www.khanacademy.org\/math\/linear-algebra\" rel=\"noreferrer noopener\" target=\"_blank\">linear algebra<\/a>. Great refreshers can be found on Khan Academy website (linked in the previous sentence). An algorithm that is imperative for this example is&nbsp;<strong>gradient descent<\/strong>,&nbsp;<a href=\"https:\/\/www.youtube.com\/watch?v=sDv4f4s2SB8&amp;ab_channel=StatQuestwithJoshStarmer\" rel=\"noreferrer noopener\" target=\"_blank\">which is explained well in this video.<\/a><\/p>\n\n\n\n<p id=\"7f91\">For a course that is more relevant on neural networks,&nbsp;<a href=\"https:\/\/www.youtube.com\/watch?v=Ixl3nykKG9M&amp;t=199s&amp;ab_channel=AdamDhalla\" rel=\"noreferrer noopener\" target=\"_blank\">this video by Adam Dhalla<\/a>&nbsp;teaches you only the necessary areas of calculus and linear algebra that are needed for this example.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"11f1\">Neural Network Basics<\/h1>\n\n\n\n<p id=\"382a\">The example we will be using is:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*CAqB1h-zZm3ujuNjCFiiag.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"d50f\">Typically, the&nbsp;<strong>input layer<\/strong>&nbsp;(green) are the&nbsp;<strong>input variables<\/strong>&nbsp;from a dataset, and the&nbsp;<strong>output layer<\/strong>&nbsp;(red) is the neural network\u2019s<strong>&nbsp;prediction value<\/strong>. Within the hidden and output layers, a&nbsp;<strong>weighted sum&nbsp;<\/strong>(denoted by&nbsp;<em>s<\/em>) is taken for the each node, followed by the application of an&nbsp;<strong>activation function&nbsp;<\/strong>(denoted by&nbsp;<em>a<\/em>) which normalises the value according to a desired&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/Activation_function\" rel=\"noreferrer noopener\" target=\"_blank\">activation function<\/a>.<\/p>\n\n\n\n<p id=\"33dc\">The process of feeding data through a network from input to output is called&nbsp;<strong>forward propagation.&nbsp;<\/strong>The process of observing the error rate of the forward propagation and feeding the error backward to fine-tune the weights of the neural network is called&nbsp;<strong>back propagation<\/strong>. We forward propagate before we back propagate.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"0fc7\">Forward Propagation<\/h1>\n\n\n\n<p id=\"5efe\">Note: I am using the&nbsp;<strong>sigmoid<\/strong>&nbsp;function as the activation function for this example (activation functions are used as mappings for inputs to be within certain ranges \u2014 in the case of sigmoid, the range is (0, 1)).<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*yHPMOt_eFOXAMeHUn5dZCg.png\" alt=\"\"\/><\/figure>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"fdfe\"><strong>Hidden Layer<\/strong><\/h1>\n\n\n\n<p id=\"7c6a\">Hidden Layer 1:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*TVvv4bp1fjQq3tdavzqVTg.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*n2N2PHuOxVnMendz8NZwKQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"ef4e\">Hidden Layer 2:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*vLTNtiui6e4A0WTXv3-hZw.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*NBBPzeS4JxVd55L5GtTBCQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"c9b3\">Hidden Layer 3:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*LFgejZ4XTBkCJrMVmAJnXA.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*eXLhnUUqHdg53s1oGh3MmA.png\" alt=\"\"\/><\/figure>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"fa15\">Output Layer<\/h1>\n\n\n\n<p id=\"89bb\">Output Layer 1:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*XgNyKpsbvq6K966Jpb73RA.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*bkW-H5Is8jTbKwab9kLuyA.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"f7e0\">Output Layer 2:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*DXVi55F1gQyP6AWGU9mX6g.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*pdAMfdsZyyHSFdrh_J0TNQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*zakTvZhHBUoidD7tEfCY2Q.png\" alt=\"\"\/><\/figure>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"bec2\">Mean Squared Error (MSE) Calculation<\/h1>\n\n\n\n<p id=\"404d\">The mean squared error is a measure of the difference between the expected and actual outputs. We are looking for a low MSE score, which indicates a better fit of the model to the data. We will be using&nbsp;<strong>gradient descent&nbsp;<\/strong>as a way to decrease this value.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*6yTL90jgzoPZAx_iiThCyA.png\" alt=\"\"\/><\/figure>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"a851\">Back Propagation<\/h1>\n\n\n\n<p id=\"4795\">Now that the predicted value is calculated, the neural network needs to adjust its weights based on the prediction error. This is done through back propagation.<\/p>\n\n\n\n<p id=\"f482\">For this example, consider a learning rate of 0.1<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*3LRFE_vbRW0S5GluUmHW9A.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"d2b2\">The general mathematical idea behind back-propagation is to apply the chain rule to find the change in the error function over the change in a weight. Consider weight 7 for this example:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*QgEDy_WR36ajuol0RyX_4w.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"a7b9\">All three partial equations can be derived from our work.<\/p>\n\n\n\n<p id=\"5d2b\">Firstly,<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*bHSDnrbPpfBr4QrgkqN2HQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"17b1\">Secondly,<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*Tx4PeDlFA_lhyu8NkCdwrA.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"4257\">And lastly,<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*IPU57RA78WJAo-LqX0H65w.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"83aa\">Hence, putting all three terms together,<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*dC-XRh17kTT8PlYwky9EAQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"65f1\">This formula can be done for all weights connecting the hidden layer to the output layer.<\/p>\n\n\n\n<p id=\"fa70\"><em>note: often authors may write the equation using a delta: \u03b4\u2080\u2081= (a\u2080\u2081\u2212expected\u2081) \u00d7 a\u2080\u2081 \u00d7 (1\u2212a\u2080\u2081), so the equation can be written as \u2202E\u2080\u2081 \/ \u2202w\u2087 = \u03b4\u2080\u2081 \u00d7 a\u2095\u2081<\/em><\/p>\n\n\n\n<p id=\"c532\">Now we have the gradient of the error function.<\/p>\n\n\n\n<p id=\"c9c6\">We want to apply gradient descent to get a new value of weight w\u2087. The new w\u2087 (we can symbolise this w\u2087\u2019) can be obtained by subtracting the learning rate multiplied by the gradient from w\u2087.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*81OHeDisAJE7NcXemSuCrg.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"7683\">So in general, for an output neuron:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*JklQcGyl8hVJx4BZxNfnvg.png\" alt=\"\"\/><\/figure>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"3db7\">Output Layer<\/h1>\n\n\n\n<p id=\"a376\">Now, applying real numbers from the example to find new values of w\u2087 through w\u2081\u2082<\/p>\n\n\n\n<p id=\"0251\">Output Layer 1:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*9I6TSPX_-nSRsx5btyEjNg.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*-VLSAIv1VoayXSUsw02png.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*xIlk2hMS5lny3t-tqm-ztQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*VQy24lxuCRUFvOqaZytfNg.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"63b4\">Output Layer 2:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*N-2BswG_5qUCXEBe6_fQew.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*H__VDUISgLWhzYlddloDWg.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*_zFOPaZNMhM9Go4sYbGlsA.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*yF-FG_n4d0VPFVt90eEpZQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"1f09\">The Hidden Layer (Derivation)<\/h1>\n\n\n\n<p id=\"da2a\">Finding a way to optimise the weights for the hidden layer has a much much larger derivation \u2014 none of this section is relevant to the calculations, so feel free to skip this part if need be.<\/p>\n\n\n\n<p id=\"a161\">Consider updating the weights for w\u2081 \u2014 in principle, updating any weight will have the same style of formula in terms of revolving around partial differentiations.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*MmPow-gPI3AYpp8v3TicZA.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"2e43\">However this time around, we are a lot further away from the output neurons \u2014 hence, to find the values of the individual components of the RHS of this equation, there is going to be a lot more \u201cchaining\u201d\u2026<\/p>\n\n\n\n<p id=\"a279\">For the first derivative:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*UG7nbOafOeF8mkL5eVlXfQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"39f2\">Where:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*sYKcyaJ1_rcJycHnoICAyQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*ydAi8Ynh26G1FHf8_UTh2Q.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"acb9\">Now, since we have calcuated \u03b4\u2080\u2081 and \u03b4\u2080\u2082 previously (see the calculations made in the output layer section of this article), we can substitute in the values of these deltas into the equation.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*Zulj8q38t10RL_I4hF9r7A.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*Pv4QmLWinBGWXqMi8CvGug.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"56c1\">Hence, the derivative of the weighted sum with respect to the previous layer\u2019s neurons is essentially just the corresponding weight.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*CGcQOMhlGO2p53_ZeJ7A6g.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*EvJtZSRIP9Lzy_0Y4ZSmRg.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"691e\">Now, substituting these values in for the partial error term:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*pkf2FKEP1hl4c6GUTa_OJA.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*xljJzVPH1i-IphLQF-peNg.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"cbdd\">The value of \u2202a\u2095\u2081 \/ \u2202s\u2095\u2081 would just be the derivative of the sigmoid function<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*KHuW8gkn_LgVtI5lobwsyA.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"18e3\">And the value of \u2202s\u2095\u2081 \/ \u2202w\u2081 is the output of the previous layer neuron (which in this case, is the input layer neuron since there is only one hidden layer)<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*w5OksFQRk_mlm6g9FPg4Rw.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"f0db\">So putting it all together:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*hmWwZNYLjvI3yeqZLHSjwA.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"bf86\">I hope you can see what has occurred in these steps \u2014 a similar working process can be done to find the formula of all of the weights (which I won\u2019t show).<\/p>\n\n\n\n<p id=\"c2bb\">But in essence, to find the value of an updated weight, first calculate delta of the weight\u2019s output neuron, and then subtract the old weight from the delta, multiplied by the delta, multiplied by the previous value of the weight\u2019s input neuron.<\/p>\n\n\n\n<p id=\"fe75\">If that is difficult to understand, then the calculations below may help you see what is occurring numerically.<\/p>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"deae\">The Hidden Layer (Calculation)<\/h1>\n\n\n\n<p id=\"a254\">Previous delta values calculated:<br>\u03b4\u2080\u2081 = -0.0984<br>\u03b4\u2080\u2082 = 0.1479<\/p>\n\n\n\n<p id=\"31a1\">Hidden layer 1:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*404nY92ibTdCk0gNs_npQQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*7In0NSoR57bgXrgnr0jCKg.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*kdX-QqQGeZ66llxpyRMKLg.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"b9a1\">Hidden layer 2:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*9S8sP3Qpc3e0eFzjwxZyaA.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*hAdKKLBrtIJdInnI6p12dg.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*88cf_2AgGeShZu7oE13Mxg.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"e201\">Hidden layer 3:<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*_PD9C-aStjc6eIgmLQYCVA.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*ZJniTOEZE9hMqNc6-jQsFg.png\" alt=\"\"\/><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*avnCa1D3AvQb115uswV0UQ.png\" alt=\"\"\/><\/figure>\n\n\n\n<p id=\"a684\">And we\u2019re done!<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/miro.medium.com\/v2\/resize:fit:1260\/1*bdW3Q7xGJyURB5h0NFxA6Q.png\" alt=\"\"\/><figcaption class=\"wp-element-caption\">Neural network with updated weights<\/figcaption><\/figure>\n\n\n\n<h1 class=\"wp-block-heading\" id=\"27af\">Closing Thoughts<\/h1>\n\n\n\n<p id=\"11e7\">The following was a complete example of a forward and back propagation for a neural network with 3 layers.<\/p>\n\n\n\n<p id=\"c4fb\">Typically neural networks are trained on multiple instances of data and can also be trained for multiple iterations (we call these&nbsp;<strong>epochs)<\/strong>. Doing this will gradually increase\/decrease the weights depending on the instance, until a neural network is optimised for a set of instances.<\/p>\n\n\n\n<p id=\"7b43\">This process was&nbsp;<em>very&nbsp;<\/em>laborious and math heavy \u2014 thankfully that is why we have computers simulate all of this work. Libraries like&nbsp;<a href=\"https:\/\/pytorch.org\/tutorials\/beginner\/basics\/buildmodel_tutorial.html\" rel=\"noreferrer noopener\" target=\"_blank\">PyTorch<\/a>&nbsp;abstract many of the mathematical complexities and should definitely be used for any sort of model training.<\/p>\n\n\n\n<p id=\"18e4\">Nonetheless, a complete walkthrough of the maths would definitely help reinforce the understanding needed when implementing this model.<a href=\"https:\/\/medium.com\/tag\/python?source=post_page-----65f2b12cdea2---------------python-----------------\"><\/a><\/p>\n\n\n\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Neural Networks are a method of artificial intelligence in which computers are taught to process data in a way similar to the human brain. Neural networks learn through being fed&#8230; <a class=\"read-more-link\" href=\"https:\/\/tbekk.com\/devstream\/2024\/06\/23\/the-mathematics-of-neural-networks-a-complete-example\/\">Read more &raquo;<\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[132],"tags":[17,320],"class_list":["post-933","post","type-post","status-publish","format-standard","hentry","category-nn","tag-machine-learning","tag-math"],"_links":{"self":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/933","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/comments?post=933"}],"version-history":[{"count":1,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/933\/revisions"}],"predecessor-version":[{"id":934,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/posts\/933\/revisions\/934"}],"wp:attachment":[{"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/media?parent=933"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/categories?post=933"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tbekk.com\/devstream\/wp-json\/wp\/v2\/tags?post=933"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}