mllm

34 projects share this GitHub topic

mllm — unilm ★22.2kmllmAgent-S — ★12.1kMobileAgent — ★9kSpatialLM — ★4.6kAwesome-LLM-Reasoning — ★3.7kNExT-GPT — ★3.6kEagle — ★3.3kInternLM-XComposer — ★2.9kmPLUG-DocOwl — ★2.4kcambrian — ★2kawesome-yolo-object-detection — ★1.8kSa2VA — ★1.6kvllm-mlx — ★1.5kOpenEMMA — ★946NEO — ★878JarvisArt — ★8534KAgent — ★816FSDrive — ★814awesome-llm-and-aigc — ★811MPP-LLaVA — ★685Woodpecker — ★649Groma — ★585Awesome-LLMs-meet-Multimodal-Generation — ★551Awesome-Multimodal-Modeling — ★508Awesome-GUI-Agents — ★446JarvisEvo — ★420LLMGA — ★395Awesome_Multimodel_LLM — ★377EVE — ★376mega-data-factory — ★370R1-VL — ★352Awesome-LVLM-Hallucination — ★325Youku-mPLUG — ★307OmniAgent — ★60Agent-S★ 12.1kMobileAgent★ 9kSpatialLM★ 4.6kAwesome-LLM-Reasoning★ 3.7kNExT-GPT★ 3.6kEagle★ 3.3kInternLM-XComposer★ 2.9kmPLUG-DocOwl★ 2.4kcambrian★ 2kawesome-yolo-object-dete…★ 1.8kSa2VA★ 1.6kvllm-mlx★ 1.5kOpenEMMA★ 946NEO★ 878JarvisArt★ 8534KAgent★ 816FSDrive★ 814awesome-llm-and-aigc★ 811MPP-LLaVA★ 685Woodpecker★ 649Groma★ 585Awesome-LLMs-meet-Multim…★ 551Awesome-Multimodal-Model…★ 508Awesome-GUI-Agents★ 446JarvisEvo★ 420LLMGA★ 395Awesome_Multimodel_LLM★ 377EVE★ 376mega-data-factory★ 370R1-VL★ 352Awesome-LVLM-Hallucinati…★ 325Youku-mPLUG★ 307OmniAgent★ 60 · GitHub ↗

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
★ 22.2k
Agent-S
Agent S: an open agentic framework that uses computers like a human
★ 12.1k
MobileAgent
Mobile-Agent: The Powerful GUI Agent Family
★ 9k
SpatialLM
[NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
★ 4.6k
Awesome-LLM-Reasoning
From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
★ 3.7k
NExT-GPT
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3.6k
Eagle
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
★ 3.3k
InternLM-XComposer
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio…
★ 2.9k
mPLUG-DocOwl
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
★ 2.4k
cambrian
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
★ 2k
awesome-yolo-object-detection
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object…
★ 1.8k
Sa2VA
Official Repo For Pixel-LLM Codebase: Sa2VA (PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st…
★ 1.6k
vllm-mlx
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama,…
★ 1.5k
OpenEMMA
OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.
★ 946
NEO
NEO Series: Native Vision-Language Models from First Principles
★ 878
JarvisArt
[NeurIPS' 2025] JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
★ 853
4KAgent
[NeurIPS 2025] 4KAgent: Agentic Any Image to 4K Super-Resolution. An intelligent computer vision agent that…
★ 816
FSDrive
[NeurIPS 2025 spotlight] Official implementation for "FutureSightDrive: Thinking Visually with…
★ 814
awesome-llm-and-aigc
🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language…
★ 811
MPP-LLaVA
Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support…
★ 685
Woodpecker
✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
★ 649
Groma
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
★ 585
Awesome-LLMs-meet-Multimodal-Generation
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
★ 551
Awesome-Multimodal-Modeling
Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
★ 508
Awesome-GUI-Agents
A curated collection of resources, tools, and frameworks for developing GUI Agents.
★ 446
JarvisEvo
[CVPR' 2026] JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator…
★ 420
LLMGA
This project is the official implementation of 'LLMGA: Multimodal Large Language Model based Generation…
★ 395
Awesome_Multimodel_LLM
Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for…
★ 377
EVE
EVE Series: Encoder-Free Vision-Language Models from BAAI
★ 376
mega-data-factory
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
★ 370
R1-VL
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy…
★ 352
Awesome-LVLM-Hallucination
up-to-date curated list of state-of-the-art Large vision language models hallucinations research work,…
★ 325
Youku-mPLUG
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks
★ 307
OmniAgent
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that…
★ 60 · GitHub ↗
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.