# YOLO (You Only Look Once)

YOLO (You Only Look Once) is a real-time object detection system based on convolutional neural networks, introduced in 2015 by Joseph Redmon et al. It requires only one forward pass to predict bounding boxes and class probabilities for the entire image.

You Only Look Once (YOLO) is a series of real-time object detection systems based on convolutional neural networks. First introduced by Joseph Redmon et al. in 2015, YOLO has undergone several iterations and improvements, becoming one of the most popular object detection frameworks. The name "You Only Look Once" refers to the fact that the algorithm requires only one forward propagation pass through the neural network to make predictions, unlike previous region proposal-based techniques like R-CNN that require thousands for a single image.

YOLO frames object detection as a single regression problem, directly predicting bounding box coordinates and class probabilities from image pixels. This unified architecture enables real-time processing speeds while maintaining competitive accuracy, making it suitable for applications such as video surveillance, autonomous driving, and robotics.

## Overview

Compared to previous methods like R-CNN and OverFeat, instead of applying the model to an image at multiple locations and scales, YOLO applies a single neural network to the full image. This network divides the image into regions and predicts bounding boxes and probabilities for each region. These bounding boxes are weighted by the predicted probabilities.

The approach draws on ideas from [deep learning](https://www.wikiprompt.org/wiki/deep-learning) and [neural networks](https://www.wikiprompt.org/wiki/neural-network), particularly convolutional architectures that have proven effective in computer vision tasks. YOLO's design prioritizes speed and simplicity, trading some localization accuracy for significant computational efficiency.

### OverFeat

OverFeat was an early influential model for simultaneous object classification and localization. Its architecture is as follows:

1. Train a neural network for image classification only ("classification-trained network"), such as AlexNet.
2. Remove the last layer of the trained network, and for every possible object class, initialize a network module at the last layer ("regression network"). The base network has its parameters frozen. The regression network is trained to predict the (x, y) coordinates of two corners of the object's bounding box.
3. During inference, run the classification-trained network over the same image at many different zoom levels and croppings. For each, it outputs a class label and a probability. Each output is then processed by the regression network of the corresponding class, resulting in thousands of bounding boxes with class labels and probabilities. These boxes are merged until only one single box with a single class label remains.

OverFeat demonstrated the feasibility of end-to-end learning for detection but remained computationally expensive due to its multi-scale inference procedure. YOLO simplified this by performing detection in a single pass.

## Versions

There are two parts to the YOLO series. The original part contained YOLOv1, v2, and v3, all released on a website maintained by Joseph Redmon. Later versions (v4 and beyond) were developed by other researchers and maintained separately.

### YOLOv1

The original YOLO algorithm, introduced in 2015, divides the image into an S x S grid of cells. If the center of an object's bounding box falls into a grid cell, that cell is said to "contain" that object. Each grid cell predicts B bounding boxes and confidence scores for those boxes. These confidence scores reflect how confident the model is that the box contains an object and how accurate it thinks the box is that it predicts.

In more detail, the network performs the same convolutional operation over each of the S^2 patches. The output of the network on each patch is a tuple (p_1, ..., p_C, c_1, x_1, y_1, w_1, h_1, ..., c_B, x_B, y_B, w_B, h_B), where p_i is the conditional probability that the cell contains an object of class i, conditional on the cell containing at least one object. The terms x_j, y_j, w_j, h_j are the center coordinates, width, and height of the j-th predicted bounding box centered in the cell. Multiple bounding boxes are predicted to allow each prediction to specialize in one kind of bounding box. The term c_j is the predicted intersection over union (IoU) of each bounding box with its corresponding ground truth.

The network architecture has 24 convolutional layers followed by 2 fully connected layers. During training, for each cell that contains a ground truth bounding box, only the predicted bounding box with the highest IoU with the ground truth is used for gradient descent. The selected box's coordinates are trained to approach the ground truth, the class probability for the correct class is trained towards 1, and other class probabilities are trained towards zero. Cells without ground truth objects contribute only to confidence score loss, pushing confidence towards zero.

### YOLOv2 and YOLOv3

YOLOv2, released in 2016, introduced batch normalization, anchor boxes, and a multi-scale training strategy, improving recall and localization accuracy. It also used a custom darknet-19 architecture and could run at up to 67 frames per second on a GPU. YOLOv3, released in 2018, added a feature pyramid network with three scales of detection, allowing better handling of small objects. It used a deeper darknet-53 backbone and logistic regression for class prediction.

### Later Versions

YOLOv4, released in 2020 by Alexey Bochkovskiy and colleagues, incorporated CSPDarknet53, PANet, and mosaic data augmentation, achieving state-of-the-art speed and accuracy trade-offs. YOLOv5, YOLOv6, YOLOv7, and YOLOv8 followed, with improvements in efficiency, deployment on edge devices, and support for tasks beyond detection, such as instance segmentation and pose estimation. These later versions are maintained by different groups, including Ultralytics for YOLOv5 and YOLOv8.

## Applications and Impact

YOLO has been widely adopted in industry and academia due to its real-time capabilities. It is used in [Tesla Autopilot](https://www.wikiprompt.org/wiki/tesla-autopilot) for object detection in autonomous driving, in [Waymo](https://www.wikiprompt.org/wiki/waymo)'s self-driving vehicles, and in [Cruise](https://www.wikiprompt.org/wiki/cruise)'s robotaxi fleet. It also powers applications in [Amazon Web Services](https://www.wikiprompt.org/wiki/amazon-web-services) Rekognition and [Google Cloud](https://www.wikiprompt.org/wiki/google-cloud) Vision for image analysis. In [artificial intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence) research, YOLO serves as a baseline for comparing new detection methods and has inspired numerous variants and extensions.

The framework's efficiency has made it popular for embedded and mobile deployments, including on devices from [Apple](https://www.wikiprompt.org/wiki/apple), [Samsung Electronics](https://www.wikiprompt.org/wiki/samsung-electronics), and [Qualcomm](https://www.wikiprompt.org/wiki/qualcomm). Its open-source implementations, particularly darknet and PyTorch versions, have contributed to its widespread use in [machine learning](https://www.wikiprompt.org/wiki/machine-learning) education and prototyping.

## Limitations and Future Directions

Despite its strengths, YOLO has limitations. It can struggle with small objects and overlapping instances compared to two-stage detectors like Faster R-CNN. Its single-pass nature also makes it less flexible for tasks requiring iterative refinement. Researchers have addressed these issues through architectural changes, such as attention mechanisms and transformer-based backbones, aligning with trends in [transformers](https://www.wikiprompt.org/wiki/transformer) and [large language models](https://www.wikiprompt.org/wiki/large-language-model) for vision tasks.

Future work continues to focus on improving accuracy-speed trade-offs, robustness to domain shift, and integration with other modalities like depth and lidar. As hardware improves, including specialized accelerators like [AWS Trainium](https://www.wikiprompt.org/wiki/aws-trainium) and [Cerebras](https://www.wikiprompt.org/wiki/cerebras) systems, YOLO-based models are expected to remain relevant for real-time perception in [generative AI](https://www.wikiprompt.org/wiki/generative-ai) pipelines and beyond.

---
Source: https://www.wikiprompt.org/wiki/yolo-paper
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-10-07T16:32:19.091628+00:00
