# Jitendra Malik

Jitendra Malik is an Indian-American computer scientist and professor at UC Berkeley, known for pioneering research in computer vision, including shape contexts, the Berkeley Segmentation Dataset, and R-CNN object detection.

Jitendra Malik (born 11 October 1960) is an Indian-American academic who is the Arthur J. Chick Professor of Electrical Engineering and Computer Sciences at the [University of California, Berkeley](https://www.wikiprompt.org/wiki/berkeley-ai-research). He is known for his research in [computer vision](https://www.wikiprompt.org/wiki/computer-vision), spanning visual recognition, segmentation, and robotics, and has influenced the field through both technical contributions and leadership.

Malik's work has bridged computational modeling and biological vision, drawing on psychology and neuroscience. He has mentored over eighty doctoral students and postdoctoral researchers, many of whom now lead academic and industrial research groups. His contributions have been recognized with numerous awards, including the 2019 IEEE Computer Society Computer Pioneer Award.

## Academic biography

Malik was born in Mathura, India, on October 11, 1960. He completed his schooling at St. Aloysius Senior Secondary School in Jabalpur. He earned a BTech degree in electrical engineering from the Indian Institute of Technology Kanpur in 1980 and a PhD in computer science from [Stanford University](https://www.wikiprompt.org/wiki/stanford-ai-lab) in 1985. In January 1986, he joined UC Berkeley, where he became the Arthur J. Chick Professor in the Computer Science Division, Department of Electrical Engineering and Computer Sciences (EECS). He also holds faculty appointments in bioengineering and the Cognitive Science and Vision Science groups. He served as chair of the Computer Science Division from 2002 to 2004 and as EECS department chair from 2004 to 2006 and 2016 to 2017.

Malik was a visiting research scientist at Google during 2015–2016. He later joined Meta's Fundamental AI Research (FAIR), serving as Research Director and Site Lead in Menlo Park before becoming Vice President for Robotics Research in 2025. In 2026, he joined Amazon as Vice President and Distinguished Scientist at its Frontier AI and Robotics (FAR) laboratory while remaining affiliated with UC Berkeley.

## Research overview

Malik has been a leader in computer vision for multiple decades, contributing to many fundamental results. His approach commonly draws inspiration from psychology and neuroscience, contributing to computational models of biological vision. He framed the central problems of computer vision as the "3R's" - recognition, reconstruction, and re-organization - emphasizing their close coupling. More recently, his group has turned to robotics, making significant contributions to navigation and legged locomotion.

## Visual recognition

The late 1990s marked a transition from geometric to learning techniques for visual recognition. Malik's group pioneered hand-designed features such as textons (vector quantized filter outputs) and shape contexts (relative arrangements of points), as well as novel machine learning techniques like fast intersection kernels and SVM-KNN. These enabled world-record performance on tasks including handwritten digit recognition (2001), breaking CAPTCHAs (2002), Caltech101 categories (2005–07), and people detection (2009). The shape context method received the Helmholtz test-of-time award and has over 9,000 citations.

In 2012, the "AlexNet" work from Geoff Hinton's group launched the [deep-learning](https://www.wikiprompt.org/wiki/deep-learning) revolution. Malik played a catalytic role by encouraging Hinton to prove that deep learning worked better by competing on standard visual recognition benchmarks, specifically ImageNet. After AlexNet, the community remained skeptical about generality, since ImageNet did not require object localization. Girshick et al invented the R-CNN method, which proved this could be done with pre-training on ImageNet classification followed by fine-tuning for object detection. This paper, with over 44,000 citations for its CVPR 2014 version, received the Longuet-Higgins test-of-time award. R-CNN variants dominated object recognition for nearly a decade until superseded by vision-language models.

## Segmentation

Visual grouping and figure-ground discrimination were first studied by the Gestalt school over a century ago. Malik's group, over two decades (1995–2015), transformed the area from a miscellaneous collection of techniques to an empirically based science. They created the Berkeley Segmentation Data Set (BSDS) using multiple human observers to mark perceived boundaries, showing consistency with a hierarchical model of segmentation. The BSDS paper has over 10,000 citations and received the Helmholtz test-of-time award.

Armed with this dataset, Malik et al developed a rationalization for various Gestalt cues as "optimal" if human vision evolved to be adaptive to natural world statistics. This enabled quantitative comparison of segmentation algorithms, rather than cherry-picked examples. The idea of benchmarking on standard datasets became a norm, particularly in object recognition led by Perona (Caltech 101) and Everingham et al (PASCAL VOC).

## Impact and legacy

Malik's influence extends beyond his own research. He has supervised more than eighty doctoral students and postdoctoral researchers, many of whom have become leading academics at institutions including MIT, UC Berkeley, Carnegie Mellon University, Cornell University, the University of Illinois Urbana–Champaign, the University of Pennsylvania, and the University of Michigan, as well as industrial researchers at Google, Meta, and other major organizations. His work has shaped the trajectory of computer vision, from early feature engineering to the deep learning era, and continues to influence robotics and artificial intelligence.

---
Source: https://www.wikiprompt.org/wiki/jitendra-malik
License: CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)
Last updated: 2026-09-13T03:59:03.638676+00:00
