Skip to main content
Omkar Thawakar

PhD Researcher / Computer Vision

Omkar Thawakar

PhD Researcher in Multimodal Artificial Intelligence

Mohamed bin Zayed University of Artificial Intelligence

Research in multimodal reasoning, video understanding and retrieval, large multimodal models, self-evolving AI systems, and real-world deployment.

HF Model downloads700K+
Granted US patents2
ICLR 2025Spotlight · Top 2%
Peer reviews completed200+

Biography

I am a PhD researcher at the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI).

My research spans multimodal reasoning, composed video retrieval, open-world video instance segmentation, large multimodal models, and self-evolving agentic systems. I am fortunate to be advised by Prof. Fahad Khan, Dr. Rao Anwer, and Dr. Salman Khan.

I am particularly interested in connecting foundational research with real-world deployment. My work includes efficient language models, on-device multimodal retrieval, and applied AI platforms. During my PhD, I am also working as a Research Scientist Intern at Adobe Research from June to September 2026. Before beginning my PhD, I worked across both academic research and production machine learning.

Research Scientist Intern · Jun 2026–Sep 2026
Adobe Research
PhD Researcher · Aug 2023–present
Mohamed bin Zayed University of Artificial Intelligence, UAE
Researcher · Nov 2021–Jul 2023
Mohamed bin Zayed University of Artificial Intelligence, UAE
Machine Learning Engineer · Feb 2020–May 2021
Chefling India Pvt. Ltd.
Research Assistant · Jan 2019–Feb 2021
Indian Institute of Technology Ropar, India
B.Tech · Computer Science · 2015–2019
SGGS Institute of Engineering & Technology, India

News

Recent publications, research impact, awards, and technology-transfer milestones.

  1. LLM Post-Training accepted to IEEE TPAMI

    Our survey, “LLM Post-Training: A Deep Dive into Reasoning Large Language Models,” has been accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence.

  2. One paper accepted at ECCV 2026

    New work on visual-token attention in self-evolving large multimodal models.

  3. Three papers accepted at CVPR 2026

    Research spanning self-evolving LMMs, reason-aware composed video retrieval, and fine-grained recognition.

  4. MobiLLaMA receives an ICLR spotlight

    Recognized among the top 2% of submissions at the ICLR SLLM workshop.

  5. All Languages Matter selected as a CVPR highlight

    A multilingual evaluation of large multimodal models across 100 culturally diverse languages.

  6. Winner of the National Grant Program for Culture & Creativity

    Awarded a 100K AED grant by the UAE Ministry of Culture.

  7. Khalifa Fund Entrepreneurship Competition winner

    Won first place and a 250K AED product grant at Abu Dhabi Business Week.

  8. Sandook-Al-Watan Student Pitch Competition winner

    Won first place and received a 30K AED project grant.

  9. Microsoft Founders Hub grant for Nutrigenics.Care

    Awarded $150K in technical support for the clinical nutrition intelligence platform.

  10. First place at Pitch Day with Microsoft

    Won the MBZUAI Incubation & Entrepreneurship Center pitch competition.

  11. Best B.Tech Thesis Award

    Recognized for “Video Super-Resolution using Recurrent GAN” at SGGSIE&T.

  12. Best Research Paper at IC3NS 2018

    Received the award for work on applying machine-learning algorithms to an IPM device.

Research

Developing multimodal systems that reason about complex visual information, adapt through feedback, and operate beyond closed-world assumptions.

  1. Multimodal Learning and Reasoning

    Large multimodal models, grounded visual reasoning, and step-by-step problem solving across images, video, and language.

    LlamaV-o1 · EvoLMM · VISE
  2. Video Understanding and Retrieval

    Composed video retrieval, dense modifications, temporal context modeling, and reason-aware media search.

    CoVR-R · BSE-CoVR · CoVR
  3. Open-World Visual Understanding

    Video instance segmentation and recognition systems that identify unknowns and learn beyond fixed vocabularies.

    OW-VISFormer · MSSTS
  4. Efficient and Transparent AI

    Compact language and multimodal models designed for reproducibility, deployment, and resource-limited environments.

    MobiLLaMA · On-device AI
  5. Agentic and Self-Evolving Systems

    Models and agents that improve through continuous rewards, tool use, retrieval, and interaction with real environments.

    Self-evolution · RAG · Agents

Selected publications

A selection of recent peer-reviewed work.

Browse all publications →
2026 ECCV

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Shahbaz Khan

2026 CVPR

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

Omkar Thawakar, Shravan Venkatraman, Ritesh Thawkar, Abdelrahman M. Shaker, Hisham Cholakkal, Rao Muhammad Anwer, Salman H. Khan, Fahad Shahbaz Khan

2026 CVPR

CoVR-R: Reason-Aware Composed Video Retrieval

Omkar Thawakar, Dmitry Demidov, Vaishnav Potlapalli, Bogireddy Sai Prasanna Teja, Viswanatha Reddy Gajjala, Alaa Mostafa Lasheen, Rao Muhammad Anwer, Fahad Shahbaz Khan

2025 ACL

LlamaV-o1: Rethinking Step-by-Step Visual Reasoning in LLMs

Omkar Thawakar, Dinura Dissanayake, Ketan Pravin More, Ritesh Thawkar, Ahmed Heakl, Noor Ahsan, Yuhao Li, Mohammed Zumri, Jean Lahoud, Rao Muhammad Anwer, Hisham Cholakkal, Ivan Laptev, Mubarak Shah, Fahad Shahbaz Khan, Salman H. Khan

2025 ICLR

MobiLLaMA: Towards Accurate & Lightweight Fully Transparent GPT

Omkar Thawakar, Ashmal Vayani, Salman H. Khan, Hisham Cholakkal, Rao Muhammad Anwer, Michael Felsberg, Tim Baldwin, Eric P. Xing, Fahad Shahbaz Khan · Spotlight, Top 2%

Selected projects

Representative systems spanning multimodal research, on-device intelligence, and applied AI.

CVPR 2026 · Composed retrieval

CoVR-R

Reason-aware composed video retrieval that explains how reference media and edit instructions determine a match.

On-device multimodal AI · iOS

VisQ

An iPhone application for natural-language and reference-guided image and video retrieval with on-device reasoning.

Personalized health AI

Nutrigenics.Care

A privacy-conscious nutrition and lifestyle platform combining conversational AI, health data, and evidence-grounded guidance.

Efficient language models · ICLR 2025

MobiLLaMA

An accurate, lightweight, and fully transparent language model family designed for efficient research and deployment.

People

Collaboration, mentorship & supervision

Students and researchers I have collaborated with, mentored, or supervised across multimodal AI, video understanding, efficient learning, and applied research.

PhD students

  • Dmitry Demidov Dmitry DemidovPhD student
  • Sara Ghaboura Sara GhabouraPhD student
  • Wafa Alghallabi Wafa AlghallabiPhD student

Master’s students

  • Shravan Venkatraman Shravan VenkatramanMaster’s student
  • Ritesh Thawkar Ritesh ThawkarMaster’s student
  • Ketan More Ketan MoreMaster’s student

Researchers

  • Ashmal Vayani Ashmal VayaniMaster’s at UCF
  • Noor AhsanResearcher
  • Ketan More Ketan MoreResearcher
  • Ritesh Thawkar Ritesh ThawkarResearcher

Contact

Email omkar.thawakar@mbzuai.ac.ae thawakar.omkar@gmail.com
Location MBZUAI, Masdar City
Abu Dhabi, United Arab Emirates Download curriculum vitae ↓