Neural networks are machine learning models that learn from examples instead of following rules written in advance. A network is made of many simple computing units — artificial neurons arranged in layers — and everything it learns is stored as numbers: the weights of the connections between those neurons.
Below: how a neural network is built and how it learns (no formulas), how fully connected, convolutional and recurrent networks and transformers differ, how ChatGPT and other large language models fit in, where you use neural networks every day and where they fall short. At the end, you’ll find a short history, a glossary and answers to common questions.
Neural networks in plain English
A person writes an ordinary program by setting rules like “if the photo shows whiskers and pointy ears, it’s a cat.” On real photos, such rules break: the cat may face away, sit in the dark or hide half behind the couch. A neural network is trained differently: it sees thousands of photos labeled “cat” and “not a cat” and finds the features that tell them apart on its own.
The name is a nod to biology: the idea was inspired by the brain, where neurons are linked by synapses that can grow stronger or weaker. In an artificial network, weights play the role of synapses. But an artificial neuron isn’t a cell — it’s a simple math operation, and the network “thinks” nothing like a person does.
Neural networks are one method of machine learning, and machine learning is part of artificial intelligence. A network with more than one hidden layer is called deep, hence the term “deep learning.” The large language models behind ChatGPT, Gemini and Claude are deep neural networks too — just very big ones.
How neural networks work
The neuron: weigh, add up, decide
Every artificial neuron does three things. It multiplies each input number by its own weight and adds up the results. It adds a bias — one more number that shifts the firing threshold. Then it passes the sum through an activation function: for example, the popular ReLU function turns negative values into zero and leaves positive ones as they are.
Think of weights as volume knobs: they decide how strongly each input signal affects the result. The activation function lets the network find complex, nonlinear patterns. Without it, any number of layers would collapse into one simple formula.
Layers: from pixels to an answer
Neurons are grouped into layers: the input layer receives data, hidden layers transform it and the output layer gives the answer. The classic textbook example is recognizing handwritten digits from the MNIST dataset. A 28×28-pixel image becomes 784 input numbers, one for the brightness of each pixel, and the output layer has 10 neurons — one for each digit from 0 to 9.
The deeper the layer, the more complex the features it responds to. In image networks, the first layers usually react to edges and blotches of color, while deeper ones respond to whole patterns and parts of objects. Nobody teaches the network these features directly: they emerge on their own during training.
Training: error and backpropagation
A new network starts with random weights, so its answers are meaningless. Training is the same loop repeated over thousands or millions of examples with known correct answers.
Prediction
The network gets an example and gives an answer
Error
The loss function compares the answer with the correct one
Backpropagation
An algorithm works out how each weight affected the error
Weight update
The weights move a small step toward a smaller error
The loss function measures the error: the further the answer is from the correct one, the bigger the number. Backpropagation runs from the network’s output back to its input and, for each weight, works out how it affected the error and which way to nudge it. That nudge is a step of gradient descent: like walking down a mountain in fog, always stepping where the slope goes down.
One full pass through all the training examples is called an epoch. A trained network is tested on examples it didn’t see during training. If it almost never errs on familiar data but often errs on new data, the network has overfit: it memorized the answers instead of learning the pattern.
Main types of neural networks and where they’re used
There are many neural network architectures, but four main ones are enough to understand how familiar services work.
| Network type | How it works | Where it’s used |
|---|---|---|
| Fully connected | Each neuron feeds the whole next layer | Tables, forecasts, parts of models |
| Convolutional (CNN) | Filters scan the image for patterns | Photos, video, medical scans |
| Recurrent (RNN) | Reads step by step, remembers the past | Speech, time series, sensors |
| Transformer | Attention links all parts of the input | Chatbots, translation, search, photos |
Fully connected networks
The simplest design is the multilayer perceptron: every neuron in one layer is connected to every neuron in the next. Such networks suit tabular data — for example, estimating an apartment’s price from its size, neighborhood and floor — and serve as building blocks inside more complex models, including transformers.
For images, a fully connected design is wasteful. For a 200×200-pixel color image, each neuron in the first layer would get 120,000 weights, and you need a lot of such neurons.
Convolutional networks (CNNs)
A convolutional network doesn’t look at the whole picture at once. Small filters “slide” across the image looking for repeating patterns — edges, textures, shapes — and the same filter works in any part of the frame. That’s why the network recognizes a cat both in the center of a photo and in the corner, and it needs far fewer weights than a fully connected one.
Convolutional networks recognize faces and objects in photos and help analyze medical images. In 2012, the convolutional network AlexNet won the ImageNet image recognition contest: the correct answer wasn’t among its top five guesses 15.3% of the time, versus 26.2% for the runner-up.
Recurrent networks (RNNs)
A recurrent network processes a sequence one element at a time — word by word or reading by reading — and keeps a “memory” of what came before. Plain RNNs quickly forget the beginning of a long sequence. To deal with this, LSTM networks (long short-term memory) were described in 1997.
LSTM powered the neural machine translation system that Google launched in Google Translate in 2016. Since then, recurrent networks have largely given way to transformers in text and speech tasks, but they’re still used for sequential data such as time series.
Transformers
Google researchers proposed the transformer in 2017 in the paper “Attention Is All You Need.” Its key idea is the attention mechanism: the model looks at all the words in a sentence at once and, for each one, decides which other words matter for its meaning. In “I arrived at the bank after crossing the river,” attention links “bank” to “river,” so the model reads it as a riverbank.
Unlike a recurrent network, a transformer processes the whole sequence in parallel, so it’s faster to train on huge amounts of data. The GPT, Gemini and Llama language models are built on transformers, and since 2020 transformers have been applied to images too: a picture is cut into small squares that are processed like a sequence of “words.”
There are other architectures as well. For example, generative adversarial networks (GANs) and diffusion models can create new images, not just recognize existing ones.
How large language models relate to neural networks
A large language model (LLM) is a deep neural network — most often a transformer — with billions of parameters, trained on massive amounts of text. Its core skill is simple: given the beginning of a text, predict the next token, a word or part of a word. By repeating this step again and again, the model writes answers, emails and code.
Training happens in several stages. First comes pretraining: the model reads a huge corpus of unlabeled text and learns to continue it. Such a “raw” model is still bad at following instructions, so it’s fine-tuned on example dialogues and with RLHF — reinforcement learning from human feedback: people compare candidate answers, and the model learns to produce the ones rated higher.
Parameter counts show how fast models have grown.
60 million
AlexNet, 2012
parameters, image recognition
175 billion
GPT-3, 2020
parameters, OpenAI language model
405 billion
Llama 3.1 405B, 2024
parameters, Meta’s open model
Source: Developer papers and blogs: NeurIPS 2012, OpenAI, Meta
Llama 3.1 405B was trained on more than 15 trillion tokens using more than 16,000 Nvidia H100 GPUs. Its context window is 128,000 tokens: that’s how much text (your prompt, files, earlier messages) the model takes into account at once. Anything that doesn’t fit in the window, it simply doesn’t see.
ChatGPT, Gemini and Claude are products that run on such models. Besides the model itself, these chatbots include web search, file handling and other tools. We covered how the assistants differ in practice and which one to choose in our ChatGPT vs Gemini vs Claude comparison.
The relationship is simple: every LLM is a neural network, but far from every neural network is an LLM. The network that recognizes your face in Face ID doesn’t write text, and a text-only model can’t see images. That said, modern models are increasingly multimodal: Gemini, for example, was trained from the start to work with text, images, audio and video.
Where neural networks work every day in 2026
In 2026, neural networks are built into familiar services even when the word “AI” appears nowhere. Here are a few examples, with figures from the companies themselves and a regulator:
- Face unlock. In Face ID, the Neural Engine turns a depth map of your face into a mathematical representation, and separate neural networks guard against attempts to unlock an iPhone with a mask. According to Apple, the chance that a random person could unlock your phone is less than 1 in 1,000,000.
- Translation. Google Translate began moving to neural machine translation in 2016: according to Google, errors dropped by 55–85% on several major language pairs. Voice translators run on neural networks too.
- Spam-free email. According to Google, Gmail’s filters block more than 99.9% of spam, phishing and malware. In 2019, the company said TensorFlow-based models were blocking around 100 million additional spam messages a day.
- Traffic and travel time. Since 2020, Google Maps has predicted travel times using DeepMind’s graph neural networks. According to the company, forecasts became up to 50% more accurate in Berlin, Tokyo, Sydney and several other cities.
- Recommendations. Back in 2016, Google described how YouTube recommendations are built by two deep neural networks: one picks candidate videos from a huge catalog, the other ranks them.
- Search and chatbots. Since May 2025, AI Overviews in Google Search have been available in more than 200 countries and territories and in more than 40 languages. ChatGPT, Gemini and Claude answer questions and write text and code.
- Images and music. Generative neural networks create images and tracks from a text description — for example, music in Suno, Udio and Adobe Firefly.
- Medicine. The public list kept by the US regulator FDA includes more than 1,600 AI-enabled medical devices and programs authorized for marketing in the US (data through the end of June 2026). About three quarters of them fall under radiology, meaning they work with medical images.
Limits: errors, data and energy
Errors and hallucinations
A neural network gives the most likely answer, not a verified one. In a recognition system, this looks like an ordinary mistake: a dog taken for a cat. In language models, it looks like a hallucination: a confident, polished, but made-up fact, quote or link.
OpenAI researchers explained this in a 2025 paper: in training and in tests, models are rewarded more for a lucky guess than for an honest “I don’t know,” so guessing pays off. By their account, hallucinations persist even in state-of-the-art systems.
Mistakes can be costly. In June 2023, a federal court in New York fined lawyers $5,000 for filing a brief that cited nonexistent court decisions made up by ChatGPT (Mata v. Avianca).
Data and bias
A neural network knows only what was in its training data, skews included. In the Gender Shades study (MIT, 2018), three commercial systems that classify gender from photos had error rates of no more than 0.8% for lighter-skinned men but 20.8–34.7% for darker-skinned women.
If data is scarce or outdated, a model makes more mistakes. A language model’s knowledge is also limited to the point when text collection for its training ended: it learns about later events only if it gets them through web search or from your documents.
Finally, a neural network is hard to inspect from the inside. Anthropic acknowledges that models are mostly treated as a “black box”: a prompt goes in, a response comes out, and it’s unclear why the model answered the way it did. A separate field, interpretability, studies how they work inside.
Energy
Training and running large models takes a lot of electricity. According to the International Energy Agency, data centers used about 415 TWh in 2024 — around 1.5% of the world’s electricity. By 2030, the agency expects their consumption to more than double to about 945 TWh, and AI is the main driver of that growth.
A single prompt takes little energy. In August 2025, Google reported that the median Gemini text prompt uses 0.24 Wh — about the same as watching TV for less than nine seconds — and that this figure fell 33-fold over 12 months.
A brief history of neural networks
Neural networks are more than 80 years old: the idea appeared long before smartphones and chatbots.
1943
A model of the neuron
Warren McCulloch and Walter Pitts described an artificial neuron mathematically
July 1958
The perceptron
Frank Rosenblatt’s network on an IBM 704 computer learned to tell cards marked on the left from cards marked on the right after 50 trials
1986
Backpropagation
David Rumelhart, Geoffrey Hinton and Ronald Williams showed in Nature how to train multilayer networks
1997
LSTM
Recurrent networks with long short-term memory were described
2012
AlexNet
A convolutional network trained on graphics cards won the ImageNet contest
June 2017
The transformer
Google researchers published “Attention Is All You Need”
November 30, 2022
ChatGPT
OpenAI launched a chatbot based on a GPT-3.5 series model
October 8, 2024
Nobel Prize
John Hopfield and Geoffrey Hinton were awarded the Nobel Prize in Physics for discoveries that enable machine learning with neural networks
Glossary: key terms
- Neuron — a unit of the network: it adds up inputs multiplied by weights and passes the sum through an activation function.
- Weight — a number that sets the strength of a connection between neurons. The weights hold everything the network has learned.
- Activation function — a nonlinear transformation, such as ReLU. Without it, the network can’t learn complex patterns.
- Parameters — all of a model’s weights and biases. Their count is usually what people mean by a model’s size.
- Loss function — a formula that shows how far the network’s answer is from the correct one.
- Backpropagation — a way to calculate how each weight contributed to the error.
- Gradient descent — step-by-step adjustment of the weights toward a smaller error.
- Epoch — one full pass through all the training examples.
- Overfitting — the network has memorized the training examples but performs poorly on new ones.
- Token — a piece of text (a word, part of a word or a character) that a language model works with.
- Context window — how many tokens a model takes into account at once, its “working memory.”
- Hallucination — a plausible but false answer from a model.
Sources and methodology
Analysis of public sourcesChecked 29 September 2026
- Google Machine Learning Crash Course: Neural networks developers.google.com
- Anthropic: Glossary platform.claude.com
- Vaswani et al., “Attention Is All You Need” (2017) arxiv.org
- OpenAI: Why language models hallucinate (2025) openai.com
- IEA: Energy and AI (2025) iea.org
FAQ
Are neural networks and artificial intelligence the same thing?
No. Artificial intelligence is a broad field, machine learning is part of it, and neural networks are one machine learning method. ChatGPT, Gemini and Claude all run on neural networks.
Does a neural network think like a human?
No. The idea of neural networks was inspired by the brain, but an artificial neuron is a simple math operation. The network finds statistical patterns in data and doesn’t check on its own whether its answer is true.
How is a neural network different from a regular program?
A regular program is written by a person who sets the rules. A neural network gets examples and tunes its weights on its own — large models have billions of them — so it can handle tasks where rules are impossible to spell out, such as recognizing faces and speech.
Why does ChatGPT sometimes make up facts?
A language model predicts a plausible continuation of text; it doesn’t look for the truth. According to OpenAI, training and tests reward models more for guessing than for an honest “I don’t know,” so even new models hallucinate.
Can I train a neural network myself?
Yes. A small network, for example one that recognizes MNIST handwritten digits, can be trained on an ordinary computer with the free PyTorch or TensorFlow libraries. You can learn the basics in Google’s free Machine Learning Crash Course.















Leave a Reply