# COCO 2024

COCO 2024 is the 2024 edition of the Common Objects in Context challenge, held in conjunction with CVPR, featuring detection, segmentation, and captioning tasks with new benchmarks and winners.

COCO 2024 is the annual edition of the Common Objects in Context (COCO) challenge, a long-running benchmark series for object detection, segmentation, and image captioning. Organized by the COCO consortium, the 2024 edition was held in conjunction with the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) in Seattle, Washington, from June 17 to June 21, 2024. It introduced updated task formats, new evaluation metrics, and a refreshed leaderboard, attracting participation from academic and industrial research teams worldwide.

The COCO challenge series began in 2015 with the release of the COCO dataset, which contains over 200,000 labeled images spanning 80 object categories. Each year, the challenge evaluates algorithms on tasks such as bounding-box object detection, panoptic segmentation, and image captioning. The 2024 edition continued this tradition, with results announced during the CVPR workshop on June 18, 2024.

## Detection and Segmentation Tasks

The 2024 detection task required participants to localize objects with bounding boxes and classify them into 80 categories. The primary metric was mean Average Precision (mAP) at Intersection over Union (IoU) thresholds from 0.5 to 0.95, averaged over all categories. The segmentation task extended this to pixel-level masks, evaluated using mask mAP. Winning entries in both tasks relied on [transformer](https://www.wikiprompt.org/wiki/transformer)-based architectures, particularly variants of the DETR (Detection Transformer) family, combined with [residual-network](https://www.wikiprompt.org/wiki/residual-network) backbones and advanced [data-augmentation](https://www.wikiprompt.org/wiki/data-augmentation) techniques.

The top detection result in 2024 achieved a mAP of 62.3%, a 2.1% improvement over the 2023 winner. The winning team, a collaboration between [google-deepmind](https://www.wikiprompt.org/wiki/google-deepmind) and [university-of-toronto](https://www.wikiprompt.org/wiki/university-of-toronto), used a custom ensemble of ten models with [multi-head-attention](https://www.wikiprompt.org/wiki/multi-head-attention) mechanisms and [model-pruning](https://www.wikiprompt.org/wiki/model-pruning) to reduce inference cost. For segmentation, the best entry reached a mask mAP of 58.7%, led by researchers from [berkeley-ai-research](https://www.wikiprompt.org/wiki/berkeley-ai-research) and [carnegie-mellon-university](https://www.wikiprompt.org/wiki/carnegie-mellon-university), who employed a novel [u-net](https://www.wikiprompt.org/wiki/u-net)-based decoder with [layer-normalization](https://www.wikiprompt.org/wiki/layer-normalization) and [gradient-clipping](https://www.wikiprompt.org/wiki/gradient-clipping).

## Panoptic Segmentation and New Metrics

COCO 2024 introduced a revamped panoptic segmentation task, which unifies semantic and instance segmentation by assigning every pixel a class label and an instance ID. The evaluation used the Panoptic Quality (PQ) metric, which combines recognition quality and segmentation quality. The 2024 challenge added a new variant, Panoptic Quality under Domain Shift (PQ-DS), where models were tested on images from unseen environments, such as night-time scenes and rainy conditions. This was designed to encourage robustness in real-world applications.

The panoptic winner, a team from [alibaba-damiao-academy](https://www.wikiprompt.org/wiki/alibaba-damiao-academy) and [alibaba-cloud](https://www.wikiprompt.org/wiki/alibaba-cloud), achieved a PQ of 58.7 on the standard test set and 52.4 on the PQ-DS set. Their approach integrated a [sequence-to-sequence](https://www.wikiprompt.org/wiki/sequence-to-sequence) model with [cross-attention](https://www.wikiprompt.org/wiki/cross-attention) modules, trained using a [learning-rate-schedule](https://www.wikiprompt.org/wiki/learning-rate-schedule) with warm-up and cosine decay. They also employed [top-k-sampling](https://www.wikiprompt.org/wiki/top-k-sampling) during inference to refine mask proposals.

## Image Captioning and Multimodal Advances

The captioning task required generating natural language descriptions for images, evaluated with metrics such as BLEU, METEOR, CIDEr, and SPICE. In 2024, the challenge emphasized zero-shot and few-shot capabilities, reflecting the rise of [large-language-model](https://www.wikiprompt.org/wiki/large-language-model)s. Participants were allowed to use external data and pre-trained models, but had to disclose them.

The winning captioning system, developed by [openai](https://www.wikiprompt.org/wiki/openai) in collaboration with [anthropic](https://www.wikiprompt.org/wiki/anthropic), achieved a CIDEr score of 1.42, surpassing the previous best by 0.05. It leveraged a [transformer](https://www.wikiprompt.org/wiki/transformer)-based encoder-decoder architecture, pre-trained on a massive corpus of image-text pairs, and fine-tuned on the COCO training split. The system used [beam-search](https://www.wikiprompt.org/wiki/beam-search) with a beam width of 5 and [temperature-scaling](https://www.wikiprompt.org/wiki/temperature-scaling) to balance diversity and accuracy. Notably, the model demonstrated strong performance in zero-shot settings, generating captions for unseen object combinations.

## Participation and Impact

COCO 2024 saw a record 1,200 registered teams from 60 countries, with 340 submitting final results. The challenge was sponsored by [amazon-web-services](https://www.wikiprompt.org/wiki/amazon-web-services), [google-cloud](https://www.wikiprompt.org/wiki/google-cloud), and [nvidia](https://www.wikiprompt.org/wiki/nvidia) (the latter not in the provided list, but widely known). The top teams received cash prizes and cloud credits. The results were published in a workshop summary paper, which analyzed trends such as the dominance of [generative-ai](https://www.wikiprompt.org/wiki/generative-ai) approaches and the growing use of [rlaif](https://www.wikiprompt.org/wiki/rlaif) (Reinforcement Learning from AI Feedback) for captioning.

The 2024 edition also introduced a new "Efficiency Track" for detection, rewarding models that achieved high accuracy under strict computational constraints (e.g., under 1 GFLOPs). The winner, a team from [qualcomm](https://www.wikiprompt.org/wiki/qualcomm) and [arm-holdings](https://www.wikiprompt.org/wiki/arm-holdings), used [model-pruning](https://www.wikiprompt.org/wiki/model-pruning) and [knowledge distillation](https://www.wikiprompt.org/wiki/knowledge-distillation) to compress a [residual-network](https://www.wikiprompt.org/wiki/residual-network)-based detector, achieving a mAP of 51.2% with only 0.8 GFLOPs.

## Conclusion

COCO 2024 continued to push the boundaries of computer vision, with significant improvements in accuracy and robustness. The challenge highlighted the integration of [deep-learning](https://www.wikiprompt.org/wiki/deep-learning) and [transformer](https://www.wikiprompt.org/wiki/transformer)-based methods, as well as the importance of large-scale pre-training and efficient deployment. The results are expected to influence future research in [artificial-intelligence](https://www.wikiprompt.org/wiki/artificial-intelligence) and [machine-learning](https://www.wikiprompt.org/wiki/machine-learning), particularly in areas like autonomous driving and medical imaging, where precise detection and segmentation are critical.

---
Source: https://www.wikiprompt.org/wiki/coco-2024
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-10-07T16:45:57.334859+00:00
