Neural networks: what they are and how they work

What a neural network is, how it learns from examples, how convolutional and recurrent networks and transformers differ, and how ChatGPT and LLMs fit in.

A neural network diagram: an input layer, two hidden layers and an output layer, with a signal path from an image to a text answer

At a glance

  • A neural network is a model of layered “neurons” that learns from examples by tuning its weights
  • Training is a loop: prediction, measuring the error, backpropagation and a weight update
  • Main types: fully connected, convolutional and recurrent networks, and transformers behind ChatGPT
  • Neural networks make mistakes and “hallucinate,” so check important facts in their answers
Content9

Neural networks are machine learning models that learn from examples instead of following rules written in advance. A network is made of many simple computing units — artificial neurons arranged in layers — and everything it learns is stored as numbers: the weights of the connections between those neurons.

Below: how a neural network is built and how it learns (no formulas), how fully connected, convolutional and recurrent networks and transformers differ, how ChatGPT and other large language models fit in, where you use neural networks every day and where they fall short. At the end, you’ll find a short history, a glossary and answers to common questions.

Neural networks in plain English

A person writes an ordinary program by setting rules like “if the photo shows whiskers and pointy ears, it’s a cat.” On real photos, such rules break: the cat may face away, sit in the dark or hide half behind the couch. A neural network is trained differently: it sees thousands of photos labeled “cat” and “not a cat” and finds the features that tell them apart on its own.

The name is a nod to biology: the idea was inspired by the brain, where neurons are linked by synapses that can grow stronger or weaker. In an artificial network, weights play the role of synapses. But an artificial neuron isn’t a cell — it’s a simple math operation, and the network “thinks” nothing like a person does.

Neural networks are one method of machine learning, and machine learning is part of artificial intelligence. A network with more than one hidden layer is called deep, hence the term “deep learning.” The large language models behind ChatGPT, Gemini and Claude are deep neural networks too — just very big ones.

How neural networks work

The neuron: weigh, add up, decide

Every artificial neuron does three things. It multiplies each input number by its own weight and adds up the results. It adds a bias — one more number that shifts the firing threshold. Then it passes the sum through an activation function: for example, the popular ReLU function turns negative values into zero and leaves positive ones as they are.

Think of weights as volume knobs: they decide how strongly each input signal affects the result. The activation function lets the network find complex, nonlinear patterns. Without it, any number of layers would collapse into one simple formula.

Layers: from pixels to an answer

Neurons are grouped into layers: the input layer receives data, hidden layers transform it and the output layer gives the answer. The classic textbook example is recognizing handwritten digits from the MNIST dataset. A 28×28-pixel image becomes 784 input numbers, one for the brightness of each pixel, and the output layer has 10 neurons — one for each digit from 0 to 9.

The deeper the layer, the more complex the features it responds to. In image networks, the first layers usually react to edges and blotches of color, while deeper ones respond to whole patterns and parts of objects. Nobody teaches the network these features directly: they emerge on their own during training.

Training: error and backpropagation

A new network starts with random weights, so its answers are meaningless. Training is the same loop repeated over thousands or millions of examples with known correct answers.

How a neural network learns from one example
  1. Prediction

    The network gets an example and gives an answer

  2. Error

    The loss function compares the answer with the correct one

  3. Backpropagation

    An algorithm works out how each weight affected the error

  4. Weight update

    The weights move a small step toward a smaller error

The loss function measures the error: the further the answer is from the correct one, the bigger the number. Backpropagation runs from the network’s output back to its input and, for each weight, works out how it affected the error and which way to nudge it. That nudge is a step of gradient descent: like walking down a mountain in fog, always stepping where the slope goes down.

One full pass through all the training examples is called an epoch. A trained network is tested on examples it didn’t see during training. If it almost never errs on familiar data but often errs on new data, the network has overfit: it memorized the answers instead of learning the pattern.

Main types of neural networks and where they’re used

There are many neural network architectures, but four main ones are enough to understand how familiar services work.

Network typeHow it worksWhere it’s used
Fully connectedEach neuron feeds the whole next layerTables, forecasts, parts of models
Convolutional (CNN)Filters scan the image for patternsPhotos, video, medical scans
Recurrent (RNN)Reads step by step, remembers the pastSpeech, time series, sensors
TransformerAttention links all parts of the inputChatbots, translation, search, photos

Fully connected networks

The simplest design is the multilayer perceptron: every neuron in one layer is connected to every neuron in the next. Such networks suit tabular data — for example, estimating an apartment’s price from its size, neighborhood and floor — and serve as building blocks inside more complex models, including transformers.

For images, a fully connected design is wasteful. For a 200×200-pixel color image, each neuron in the first layer would get 120,000 weights, and you need a lot of such neurons.

Convolutional networks (CNNs)

A convolutional network doesn’t look at the whole picture at once. Small filters “slide” across the image looking for repeating patterns — edges, textures, shapes — and the same filter works in any part of the frame. That’s why the network recognizes a cat both in the center of a photo and in the corner, and it needs far fewer weights than a fully connected one.

Convolutional networks recognize faces and objects in photos and help analyze medical images. In 2012, the convolutional network AlexNet won the ImageNet image recognition contest: the correct answer wasn’t among its top five guesses 15.3% of the time, versus 26.2% for the runner-up.

Recurrent networks (RNNs)

A recurrent network processes a sequence one element at a time — word by word or reading by reading — and keeps a “memory” of what came before. Plain RNNs quickly forget the beginning of a long sequence. To deal with this, LSTM networks (long short-term memory) were described in 1997.

LSTM powered the neural machine translation system that Google launched in Google Translate in 2016. Since then, recurrent networks have largely given way to transformers in text and speech tasks, but they’re still used for sequential data such as time series.

Transformers

Google researchers proposed the transformer in 2017 in the paper “Attention Is All You Need.” Its key idea is the attention mechanism: the model looks at all the words in a sentence at once and, for each one, decides which other words matter for its meaning. In “I arrived at the bank after crossing the river,” attention links “bank” to “river,” so the model reads it as a riverbank.

Unlike a recurrent network, a transformer processes the whole sequence in parallel, so it’s faster to train on huge amounts of data. The GPT, Gemini and Llama language models are built on transformers, and since 2020 transformers have been applied to images too: a picture is cut into small squares that are processed like a sequence of “words.”

There are other architectures as well. For example, generative adversarial networks (GANs) and diffusion models can create new images, not just recognize existing ones.

How large language models relate to neural networks

A large language model (LLM) is a deep neural network — most often a transformer — with billions of parameters, trained on massive amounts of text. Its core skill is simple: given the beginning of a text, predict the next token, a word or part of a word. By repeating this step again and again, the model writes answers, emails and code.

Training happens in several stages. First comes pretraining: the model reads a huge corpus of unlabeled text and learns to continue it. Such a “raw” model is still bad at following instructions, so it’s fine-tuned on example dialogues and with RLHF — reinforcement learning from human feedback: people compare candidate answers, and the model learns to produce the ones rated higher.

Parameter counts show how fast models have grown.

60 million

AlexNet, 2012

parameters, image recognition

175 billion

GPT-3, 2020

parameters, OpenAI language model

405 billion

Llama 3.1 405B, 2024

parameters, Meta’s open model

Source: Developer papers and blogs: NeurIPS 2012, OpenAI, Meta

Llama 3.1 405B was trained on more than 15 trillion tokens using more than 16,000 Nvidia H100 GPUs. Its context window is 128,000 tokens: that’s how much text (your prompt, files, earlier messages) the model takes into account at once. Anything that doesn’t fit in the window, it simply doesn’t see.

ChatGPT, Gemini and Claude are products that run on such models. Besides the model itself, these chatbots include web search, file handling and other tools. We covered how the assistants differ in practice and which one to choose in our ChatGPT vs Gemini vs Claude comparison.

The relationship is simple: every LLM is a neural network, but far from every neural network is an LLM. The network that recognizes your face in Face ID doesn’t write text, and a text-only model can’t see images. That said, modern models are increasingly multimodal: Gemini, for example, was trained from the start to work with text, images, audio and video.

Where neural networks work every day in 2026

In 2026, neural networks are built into familiar services even when the word “AI” appears nowhere. Here are a few examples, with figures from the companies themselves and a regulator:

  • Face unlock. In Face ID, the Neural Engine turns a depth map of your face into a mathematical representation, and separate neural networks guard against attempts to unlock an iPhone with a mask. According to Apple, the chance that a random person could unlock your phone is less than 1 in 1,000,000.
  • Translation. Google Translate began moving to neural machine translation in 2016: according to Google, errors dropped by 55–85% on several major language pairs. Voice translators run on neural networks too.
  • Spam-free email. According to Google, Gmail’s filters block more than 99.9% of spam, phishing and malware. In 2019, the company said TensorFlow-based models were blocking around 100 million additional spam messages a day.
  • Traffic and travel time. Since 2020, Google Maps has predicted travel times using DeepMind’s graph neural networks. According to the company, forecasts became up to 50% more accurate in Berlin, Tokyo, Sydney and several other cities.
  • Recommendations. Back in 2016, Google described how YouTube recommendations are built by two deep neural networks: one picks candidate videos from a huge catalog, the other ranks them.
  • Search and chatbots. Since May 2025, AI Overviews in Google Search have been available in more than 200 countries and territories and in more than 40 languages. ChatGPT, Gemini and Claude answer questions and write text and code.
  • Images and music. Generative neural networks create images and tracks from a text description — for example, music in Suno, Udio and Adobe Firefly.
  • Medicine. The public list kept by the US regulator FDA includes more than 1,600 AI-enabled medical devices and programs authorized for marketing in the US (data through the end of June 2026). About three quarters of them fall under radiology, meaning they work with medical images.

Limits: errors, data and energy

Errors and hallucinations

A neural network gives the most likely answer, not a verified one. In a recognition system, this looks like an ordinary mistake: a dog taken for a cat. In language models, it looks like a hallucination: a confident, polished, but made-up fact, quote or link.

OpenAI researchers explained this in a 2025 paper: in training and in tests, models are rewarded more for a lucky guess than for an honest “I don’t know,” so guessing pays off. By their account, hallucinations persist even in state-of-the-art systems.

Mistakes can be costly. In June 2023, a federal court in New York fined lawyers $5,000 for filing a brief that cited nonexistent court decisions made up by ChatGPT (Mata v. Avianca).

Data and bias

A neural network knows only what was in its training data, skews included. In the Gender Shades study (MIT, 2018), three commercial systems that classify gender from photos had error rates of no more than 0.8% for lighter-skinned men but 20.8–34.7% for darker-skinned women.

If data is scarce or outdated, a model makes more mistakes. A language model’s knowledge is also limited to the point when text collection for its training ended: it learns about later events only if it gets them through web search or from your documents.

Finally, a neural network is hard to inspect from the inside. Anthropic acknowledges that models are mostly treated as a “black box”: a prompt goes in, a response comes out, and it’s unclear why the model answered the way it did. A separate field, interpretability, studies how they work inside.

Energy

Training and running large models takes a lot of electricity. According to the International Energy Agency, data centers used about 415 TWh in 2024 — around 1.5% of the world’s electricity. By 2030, the agency expects their consumption to more than double to about 945 TWh, and AI is the main driver of that growth.

A single prompt takes little energy. In August 2025, Google reported that the median Gemini text prompt uses 0.24 Wh — about the same as watching TV for less than nine seconds — and that this figure fell 33-fold over 12 months.

A brief history of neural networks

Neural networks are more than 80 years old: the idea appeared long before smartphones and chatbots.

Key dates
  1. 1943

    A model of the neuron

    Warren McCulloch and Walter Pitts described an artificial neuron mathematically

  2. July 1958

    The perceptron

    Frank Rosenblatt’s network on an IBM 704 computer learned to tell cards marked on the left from cards marked on the right after 50 trials

  3. 1986

    Backpropagation

    David Rumelhart, Geoffrey Hinton and Ronald Williams showed in Nature how to train multilayer networks

  4. 1997

    LSTM

    Recurrent networks with long short-term memory were described

  5. 2012

    AlexNet

    A convolutional network trained on graphics cards won the ImageNet contest

  6. June 2017

    The transformer

    Google researchers published “Attention Is All You Need”

  7. November 30, 2022

    ChatGPT

    OpenAI launched a chatbot based on a GPT-3.5 series model

  8. October 8, 2024

    Nobel Prize

    John Hopfield and Geoffrey Hinton were awarded the Nobel Prize in Physics for discoveries that enable machine learning with neural networks

Glossary: key terms

  • Neuron — a unit of the network: it adds up inputs multiplied by weights and passes the sum through an activation function.
  • Weight — a number that sets the strength of a connection between neurons. The weights hold everything the network has learned.
  • Activation function — a nonlinear transformation, such as ReLU. Without it, the network can’t learn complex patterns.
  • Parameters — all of a model’s weights and biases. Their count is usually what people mean by a model’s size.
  • Loss function — a formula that shows how far the network’s answer is from the correct one.
  • Backpropagation — a way to calculate how each weight contributed to the error.
  • Gradient descent — step-by-step adjustment of the weights toward a smaller error.
  • Epoch — one full pass through all the training examples.
  • Overfitting — the network has memorized the training examples but performs poorly on new ones.
  • Token — a piece of text (a word, part of a word or a character) that a language model works with.
  • Context window — how many tokens a model takes into account at once, its “working memory.”
  • Hallucination — a plausible but false answer from a model.

Sources and methodology

Analysis of public sourcesChecked 29 September 2026

  1. Google Machine Learning Crash Course: Neural networks developers.google.com
  2. Anthropic: Glossary platform.claude.com
  3. Vaswani et al., “Attention Is All You Need” (2017) arxiv.org
  4. OpenAI: Why language models hallucinate (2025) openai.com
  5. IEA: Energy and AI (2025) iea.org

FAQ

Are neural networks and artificial intelligence the same thing?

No. Artificial intelligence is a broad field, machine learning is part of it, and neural networks are one machine learning method. ChatGPT, Gemini and Claude all run on neural networks.

Does a neural network think like a human?

No. The idea of neural networks was inspired by the brain, but an artificial neuron is a simple math operation. The network finds statistical patterns in data and doesn’t check on its own whether its answer is true.

How is a neural network different from a regular program?

A regular program is written by a person who sets the rules. A neural network gets examples and tunes its weights on its own — large models have billions of them — so it can handle tasks where rules are impossible to spell out, such as recognizing faces and speech.

Why does ChatGPT sometimes make up facts?

A language model predicts a plausible continuation of text; it doesn’t look for the truth. According to OpenAI, training and tests reward models more for guessing than for an honest “I don’t know,” so even new models hallucinate.

Can I train a neural network myself?

Yes. A small network, for example one that recognizes MNIST handwritten digits, can be trained on an ordinary computer with the free PyTorch or TensorFlow libraries. You can learn the basics in Google’s free Machine Learning Crash Course.

Was this article helpful?

Be the first to rate

Author

Vasyl Vasyliev

Has been working on AppMaxx since 2019, writing reviews of smartphones, laptops and other gadgets, game and movie roundups, and how-tos on Windows and apps.

All articles by this author

Leave a Reply

Your email address will not be published. Required fields are marked *