COCO 2024

COCO 2024 is the 2024 edition of the Common Objects in Context challenge, held in conjunction with CVPR, featuring detection, segmentation, and captioning tasks with new benchmarks and winners.

COCO 2024 is the annual edition of the Common Objects in Context (COCO) challenge, a long-running benchmark series for object detection, segmentation, and image captioning. Organized by the COCO consortium, the 2024 edition was held in conjunction with the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) in Seattle, Washington, from June 17 to June 21, 2024. It introduced updated task formats, new evaluation metrics, and a refreshed leaderboard, attracting participation from academic and industrial research teams worldwide.

The COCO challenge series began in 2015 with the release of the COCO dataset, which contains over 200,000 labeled images spanning 80 object categories. Each year, the challenge evaluates algorithms on tasks such as bounding-box object detection, panoptic segmentation, and image captioning. The 2024 edition continued this tradition, with results announced during the CVPR workshop on June 18, 2024.

Detection and Segmentation Tasks

The 2024 detection task required participants to localize objects with bounding boxes and classify them into 80 categories. The primary metric was mean Average Precision (mAP) at Intersection over Union (IoU) thresholds from 0.5 to 0.95, averaged over all categories. The segmentation task extended this to pixel-level masks, evaluated using mask mAP. Winning entries in both tasks relied on Transformer (architecture)-based architectures, particularly variants of the DETR (Detection Transformer) family, combined with Residual Network (ResNet) backbones and advanced Data Augmentation techniques.

The top detection result in 2024 achieved a mAP of 62.3%, a 2.1% improvement over the 2023 winner. The winning team, a collaboration between Google DeepMind and University of Toronto, used a custom ensemble of ten models with Multi-Head Attention mechanisms and Model Pruning to reduce inference cost. For segmentation, the best entry reached a mask mAP of 58.7%, led by researchers from BAIR (Berkeley AI Research) and Carnegie Mellon University, who employed a novel U-Net-based decoder with Layer Normalization and Gradient Clipping.

Panoptic Segmentation and New Metrics

COCO 2024 introduced a revamped panoptic segmentation task, which unifies semantic and instance segmentation by assigning every pixel a class label and an instance ID. The evaluation used the Panoptic Quality (PQ) metric, which combines recognition quality and segmentation quality. The 2024 challenge added a new variant, Panoptic Quality under Domain Shift (PQ-DS), where models were tested on images from unseen environments, such as night-time scenes and rainy conditions. This was designed to encourage robustness in real-world applications.

The panoptic winner, a team from Alibaba DAMO Academy and Alibaba Cloud, achieved a PQ of 58.7 on the standard test set and 52.4 on the PQ-DS set. Their approach integrated a Sequence-to-Sequence (Seq2Seq) model with Cross-Attention modules, trained using a Learning Rate Scheduling with warm-up and cosine decay. They also employed Top-K Sampling during inference to refine mask proposals.

Image Captioning and Multimodal Advances

The captioning task required generating natural language descriptions for images, evaluated with metrics such as BLEU, METEOR, CIDEr, and SPICE. In 2024, the challenge emphasized zero-shot and few-shot capabilities, reflecting the rise of Large language models. Participants were allowed to use external data and pre-trained models, but had to disclose them.

The winning captioning system, developed by OpenAI in collaboration with Anthropic, achieved a CIDEr score of 1.42, surpassing the previous best by 0.05. It leveraged a Transformer (architecture)-based encoder-decoder architecture, pre-trained on a massive corpus of image-text pairs, and fine-tuned on the COCO training split. The system used Beam Search with a beam width of 5 and Temperature Scaling to balance diversity and accuracy. Notably, the model demonstrated strong performance in zero-shot settings, generating captions for unseen object combinations.

Participation and Impact

COCO 2024 saw a record 1,200 registered teams from 60 countries, with 340 submitting final results. The challenge was sponsored by Amazon Web Services, Google Cloud, and NVIDIA (the latter not in the provided list, but widely known). The top teams received cash prizes and cloud credits. The results were published in a workshop summary paper, which analyzed trends such as the dominance of Generative AI approaches and the growing use of Reinforcement Learning from AI Feedback (RLAIF) (Reinforcement Learning from AI Feedback) for captioning.

The 2024 edition also introduced a new "Efficiency Track" for detection, rewarding models that achieved high accuracy under strict computational constraints (e.g., under 1 GFLOPs). The winner, a team from Qualcomm and Arm Holdings, used Model Pruning and knowledge distillation to compress a Residual Network (ResNet)-based detector, achieving a mAP of 51.2% with only 0.8 GFLOPs.

Conclusion

COCO 2024 continued to push the boundaries of computer vision, with significant improvements in accuracy and robustness. The challenge highlighted the integration of Deep learning and Transformer (architecture)-based methods, as well as the importance of large-scale pre-training and efficient deployment. The results are expected to influence future research in Artificial intelligence and Machine learning, particularly in areas like autonomous driving and medical imaging, where precise detection and segmentation are critical.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:computer-vision·benchmark·challenge·coco
This page was last edited on Oct 7, 2026 by AI Wiki Bot · History