Muhammad Maaz (mmaaz60)

mmaaz60

Geek Repo

Company:@mbzuai

Location:Abu Dhabi, UAE

Home Page:https://www.muhammadmaaz.com

Github PK Tool:Github PK Tool


Organizations
mbzuai-oryx

Muhammad Maaz's starred repositories

MiniGPT-4

Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)

Language:PythonLicense:BSD-3-ClauseStargazers:25141Issues:220Issues:450

ConvNeXt-V2

Code release for ConvNeXt V2 model

Language:PythonLicense:NOASSERTIONStargazers:1403Issues:8Issues:68

Video-ChatGPT

[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

Language:PythonLicense:CC-BY-4.0Stargazers:1053Issues:14Issues:107

LLaVA-pp

🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)

groundingLMM

[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.

MobiLlama

MobiLlama : Small Language Model tailored for edge devices

Language:PythonLicense:Apache-2.0Stargazers:563Issues:13Issues:12

bubogpt

BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Language:PythonLicense:BSD-3-ClauseStargazers:488Issues:10Issues:19

MovieChat

[CVPR 2024] 🎬💭 chat with over 10K frames of video!

Language:PythonLicense:BSD-3-ClauseStargazers:457Issues:10Issues:68

XrayGPT

[BIONLP@ACL 2024] XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models.

GeoChat

[CVPR 2024 🔥] GeoChat, the first grounded Large Vision Language Model for Remote Sensing

prompt-pretraining

Official implementation for the paper "Prompt Pre-Training with Over Twenty-Thousand Classes for Open-Vocabulary Visual Recognition"

Language:PythonLicense:Apache-2.0Stargazers:249Issues:5Issues:13

Video-LLaVA

PG-Video-LLaVA: Pixel Grounding in Large Multimodal Video Models

SwiftFormer

[ICCV'23] Official repository of paper SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision Applications

PromptSRC

[ICCV'23 Main Track, WECIA'23 Oral] Official repository of paper titled "Self-regulating Prompts: Foundational Model Adaptation without Forgetting".

Language:PythonLicense:MITStargazers:200Issues:5Issues:15

MultiPLY

Code for MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World

PromptAlign

[NeurIPS 2023] Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization

ST-LLM

[ECCV 2024🔥] Official implementation of the paper "ST-LLM: Large Language Models Are Effective Temporal Learners"

Language:PythonLicense:Apache-2.0Stargazers:80Issues:8Issues:17

PALO

Vision-language conversation in 10 languages including English, Chinese, French, Spanish, Russian, Japanese, Arabic, Hindi, Bengali and Urdu.

Language:PythonLicense:Apache-2.0Stargazers:73Issues:7Issues:4

ClimateGPT

[EMNLP'23] ClimateGPT: a specialized LLM for conversations related to Climate Change and Sustainability topics in both English and Arabic languages.

Multimodality-Representation-Learning

This repository provides a comprehensive collection of research papers focused on multimodal representation learning, all of which have been cited and discussed in the survey just accepted https://dl.acm.org/doi/abs/10.1145/3617833 .

vlm-evaluation

VLM Evaluation: Benchmark for VLMs, spanning text generation tasks from VQA to Captioning

Language:PythonLicense:NOASSERTIONStargazers:63Issues:5Issues:8

llmblueprint

[ICLR 2024] Official code for the paper "LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts"

Language:Jupyter NotebookStargazers:60Issues:3Issues:4
Language:Jupyter NotebookStargazers:50Issues:0Issues:0

vafa

[MICCAI 2023] Official code repository of paper titled "Frequency Domain Adversarial Training for Robust Volumetric Medical Segmentation" accepted in MICCAI 2023 conference.

Language:PythonLicense:MITStargazers:46Issues:2Issues:1

MAVOS

Efficient Video Object Segmentation via Modulated Cross-Attention Memory

License:BSD-3-ClauseStargazers:44Issues:5Issues:2

hateclipper

Hate-CLIPper: Multimodal Hateful Meme Classification with Explicit Cross-modal Interaction of CLIP features - Accepted at EMNLP 2022 Workshop

Language:Jupyter NotebookStargazers:38Issues:2Issues:8
Language:PythonLicense:Apache-2.0Stargazers:37Issues:2Issues:3

composed-video-retrieval

Composed Video Retrieval

Language:PythonLicense:Apache-2.0Stargazers:29Issues:2Issues:4

sm-vit

Official repository for the paper "Salient Mask-Guided Vision Transformer for Fine-Grained Classification" (VISIGRAPP '23)

Language:PythonLicense:MITStargazers:17Issues:3Issues:0