Skip to main content
Omkar Thawakar

PhD Researcher / Computer Vision

Omkar Thawakar

PhD Researcher in Multimodal Artificial Intelligence

Mohamed bin Zayed University of Artificial Intelligence

Research in multimodal reasoning, video understanding and retrieval, large multimodal models, self-evolving AI systems, and real-world deployment.

HF Model downloads700K+
Granted US patents2
ICLR 2025Spotlight · Top 2%
Peer reviews completed200+

Biography

I am a PhD researcher at the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI).

My research spans multimodal reasoning, composed video retrieval, open-world video instance segmentation, large multimodal models, and self-evolving agentic systems. I am fortunate to be advised by Prof. Fahad Khan, Dr. Rao Anwer, and Dr. Salman Khan.

I am particularly interested in connecting foundational research with real-world deployment. My work includes efficient language models, on-device multimodal retrieval, and applied AI platforms. During my PhD, I am also working as a Research Scientist Intern at Adobe Research from June to September 2026. Before beginning my PhD, I worked across both academic research and production machine learning.

Research Scientist Intern · Jun 2026–Sep 2026
Adobe Research
PhD Researcher · Aug 2023–present
Mohamed bin Zayed University of Artificial Intelligence, UAE
Researcher · Nov 2021–Jul 2023
Mohamed bin Zayed University of Artificial Intelligence, UAE
Machine Learning Engineer · Feb 2020–May 2021
Chefling India Pvt. Ltd.
Research Assistant · Jan 2019–Feb 2021
Indian Institute of Technology Ropar, India
B.Tech · Computer Science · 2015–2019
SGGS Institute of Engineering & Technology, India

News

Recent publications, research impact, awards, and technology-transfer milestones.

  1. One paper accepted at ECCV 2026

    New work on visual-token attention in self-evolving large multimodal models.

  2. Three papers accepted at CVPR 2026

    Research spanning self-evolving LMMs, reason-aware composed video retrieval, and fine-grained recognition.

  3. MobiLLaMA receives an ICLR spotlight

    Recognized among the top 2% of submissions at the ICLR SLLM workshop.

  4. All Languages Matter selected as a CVPR highlight

    A multilingual evaluation of large multimodal models across 100 culturally diverse languages.

  5. Khalifa Fund Entrepreneurship Competition winner

    Received a 250K AED award supporting applied AI innovation and deployment.

Research

Developing multimodal systems that reason about complex visual information, adapt through feedback, and operate beyond closed-world assumptions.

  1. Multimodal Learning and Reasoning

    Large multimodal models, grounded visual reasoning, and step-by-step problem solving across images, video, and language.

    LlamaV-o1 · EvoLMM · VISE
  2. Video Understanding and Retrieval

    Composed video retrieval, dense modifications, temporal context modeling, and reason-aware media search.

    CoVR-R · BSE-CoVR · CoVR
  3. Open-World Visual Understanding

    Video instance segmentation and recognition systems that identify unknowns and learn beyond fixed vocabularies.

    OW-VISFormer · MSSTS
  4. Efficient and Transparent AI

    Compact language and multimodal models designed for reproducibility, deployment, and resource-limited environments.

    MobiLLaMA · On-device AI
  5. Agentic and Self-Evolving Systems

    Models and agents that improve through continuous rewards, tool use, retrieval, and interaction with real environments.

    Self-evolution · RAG · Agents

Selected publications

A selection of recent peer-reviewed work.

Browse all publications →
2026 ECCV

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, et al.

2026 CVPR

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

Omkar Thawakar, et al.

2026 CVPR

CoVR-R: Reason-Aware Composed Video Retrieval

Omkar Thawakar, et al.

2025 ACL

LlamaV-o1: Rethinking Step-by-Step Visual Reasoning in LLMs

Omkar Thawakar, et al.

2025 ICLR

MobiLLaMA: Towards Accurate & Lightweight Fully Transparent GPT

Omkar Thawakar, et al. · Spotlight, Top 2%

Selected projects

Representative systems spanning multimodal research, on-device intelligence, and applied AI.

CVPR 2026 · Composed retrieval

CoVR-R

Reason-aware composed video retrieval that explains how reference media and edit instructions determine a match.

On-device multimodal AI · iOS

VisQ

An iPhone application for natural-language and reference-guided image and video retrieval with on-device reasoning.

Personalized health AI

Nutrigenics.Care

A privacy-conscious nutrition and lifestyle platform combining conversational AI, health data, and evidence-grounded guidance.

Efficient language models · ICLR 2025

MobiLLaMA

An accurate, lightweight, and fully transparent language model family designed for efficient research and deployment.

People

Collaboration, mentorship & supervision

Students and researchers I have collaborated with, mentored, or supervised across multimodal AI, video understanding, efficient learning, and applied research.

PhD students

  • Dmitry Demidov Dmitry DemidovPhD student
  • Sara Ghaboura Sara GhabouraPhD student
  • Wafa Alghallabi Wafa AlghallabiPhD student

Master’s students

  • Shravan Venkatraman Shravan VenkatramanMaster’s student
  • Ritesh Thawkar Ritesh ThawkarMaster’s student
  • Ketan More Ketan MoreMaster’s student

Researchers

  • Ashmal Vayani Ashmal VayaniMaster’s at UCF
  • Noor AhsanResearcher
  • Ketan More Ketan MoreResearcher
  • Ritesh Thawkar Ritesh ThawkarResearcher

Contact

Email omkar.thawakar@mbzuai.ac.ae thawakar.omkar@gmail.com
Location MBZUAI, Masdar City
Abu Dhabi, United Arab Emirates Download curriculum vitae ↓