vision-language

18 projects share this GitHub topic

vision-language — Chinese-CLIP ★6kvision-languageOFA — ★2.6kQwen-VL-Series-Finetune — ★1.9kAdvancedLiterateMachinery — ★1.8kVideo-ChatGPT — ★1.5kawesome-japanese-llm — ★1.4kDriveLM — ★1.3kONE-PEACE — ★1.1kTinyLLaVA_Factory — ★995calvin — ★964daclip-uir — ★815SEED — ★642cliport — ★547Visual-Chinese-LLaMA-Alpaca — ★461awesome-vla-for-ad — ★453World-Simulator — ★383lmms-finetune — ★373ViP-LLaVA — ★339OFA★ 2.6kQwen-VL-Series-Finetune★ 1.9kAdvancedLiterateMachiner…★ 1.8kVideo-ChatGPT★ 1.5kawesome-japanese-llm★ 1.4kDriveLM★ 1.3kONE-PEACE★ 1.1kTinyLLaVA_Factory★ 995calvin★ 964daclip-uir★ 815SEED★ 642cliport★ 547Visual-Chinese-LLaMA-Alp…★ 461awesome-vla-for-ad★ 453World-Simulator★ 383lmms-finetune★ 373ViP-LLaVA★ 339

Lines connect members that are measurably related to each other. Dot size reflects stars.

🧬 Members
Chinese-CLIP
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
★ 6k
OFA
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a…
★ 2.6k
Qwen-VL-Series-Finetune
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1.9k
AdvancedLiterateMachinery
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project…
★ 1.8k
Video-ChatGPT
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation…
★ 1.5k
awesome-japanese-llm
日本語LLMまとめ - Overview of Japanese LLMs
★ 1.4k
DriveLM
[ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering
★ 1.3k
ONE-PEACE
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One…
★ 1.1k
TinyLLaVA_Factory
A Framework of Small-scale Large Multimodal Models
★ 995
calvin
CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
★ 964
daclip-uir
[ICLR 2024] Controlling Vision-Language Models for Universal Image Restoration. 5th place in the NTIRE 2024…
★ 815
SEED
Official implementation of SEED-LLaMA (ICLR 2024).
★ 642
cliport
CLIPort: What and Where Pathways for Robotic Manipulation
★ 547
Visual-Chinese-LLaMA-Alpaca
多模态中文LLaMA&Alpaca大语言模型(VisualCLA)
★ 461
awesome-vla-for-ad
🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
★ 453
World-Simulator
[IEEE TPAMI 2026] Simulating the Real World: Survey & Resources, which contains our survey "Simulating the…
★ 383
lmms-finetune
A minimal codebase for finetuning large multimodal models, supporting llava-1.5/1.6, llava-interleave,…
★ 373
ViP-LLaVA
[CVPR2024] ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
★ 339
🔗 Related families

Measured from GitHub topics shared by both projects, weighted by how rare each topic is.