Last updated: 2026-10-07

U
Undergraduate level
FDN
Foundational — Knowledge that endures for decades — core principles

What Is an Artificial Neuron?

The name invites a comparison the mechanism doesn't really support. An artificial neuron is a small, specific arithmetic operation — multiply each input by a learned number, add them up, add one more learned number, then pass the result through a fixed function — and most of what follows in this section builds on that operation directly, without needing any claim about brains to make it work. McCulloch and Pitts, writing in 1943, showed that networks of simple threshold units could in principle compute any logical function a biological neuron plausibly could1 — the connection to biology is real, historically, as a source of the idea, and limited, practically, to that: a loose motivating analogy for a mathematical object whose actual behaviour is fully specified by the equations below.

The Weighted Sum FoundationalKnowledge that endures for decades — core principles

Given inputs \(x_1, x_2, \ldots, x_n\), a neuron computes:

\[ z = w_1x_1 + w_2x_2 + \cdots + w_nx_n + b \]

Each weight \(w_i\) says how much — and in which direction, since a weight can be negative — its matching input should push the result. The bias \(b\) is a constant added regardless of the inputs, shifting the whole threshold up or down; without it, a neuron would be forced to treat the all-zero input as a fixed, unmovable boundary case, which is rarely what's wanted. The weighted sum \(z\) alone is just a linear function of the inputs, no more expressive than ordinary linear regression — what makes it a neuron rather than a regression line is what happens next.

The Activation Function FoundationalKnowledge that endures for decades — core principles

The weighted sum \(z\) is passed through an activation function \(f\) to produce the neuron's actual output:

\[ y = f(z) \]

The simplest choice is a threshold (or step) function: output 1 if \(z\) clears some threshold, 0 otherwise. It's the easiest to reason about by hand, and the worked example below uses exactly this choice — but it has a serious practical drawback for training: its flat regions have a slope of exactly zero almost everywhere, which gives a training process nothing to push against. Modern networks mostly use smoother alternatives instead — the sigmoid, which squashes any real number into the range (0, 1), or the ReLU (rectified linear unit), which outputs zero for any negative input and the input unchanged otherwise — chosen specifically because they stay differentiable, which matters once training means following a gradient rather than applying a fixed rule.a step function is the oldest choice, not the only one

Worked Example: Two Inputs and a Threshold FoundationalKnowledge that endures for decades — core principles

Take a neuron with two inputs, weights \(w_1 = 1\), \(w_2 = 1\), bias \(b = -1.5\), and a threshold activation that outputs 1 when \(z \geq 0\):

graph LR X1["x1"] -->|"w1 = 1"| S["weighted sum z"] X2["x2"] -->|"w2 = 1"| S B["bias = -1.5"] --> S S --> F["activation f(z)"] F --> Y["output y"]
x1x2z = x1 + x2 − 1.5y = f(z)
00−1.50
10−0.50
01−0.50
110.51

This table is exactly the logical AND function: the neuron outputs 1 only when both inputs are 1. That isn't a coincidence specific to these particular numbers — it's a consequence of the equation \(z=0\) being, geometrically, a straight line (or, with more inputs, a flat plane) cutting the input space in two, with the neuron answering 1 on one side of it and 0 on the other. Changing the weights and bias moves and rotates that line; which functions a single neuron can and can't represent this way is exactly where the next page picks up.

Note well. A single neuron's decision boundary is always a straight line (or flat plane, or hyperplane) — the weights set its orientation, the bias sets its offset from the origin. Every function a lone neuron can learn has to be separable by one straight cut.

References


  1. McCulloch, W. S., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5, 115–133. https://doi.org/10.1007/BF02478259 ↩