A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
-
Updated
Jul 14, 2026 - Python
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
[CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
🎩 An Alfred 5 Workflow for using OpenAI Chat API to interact with GPT models 🤖💬 It also allows image generation/editing/understanding 🖼️, speech-to-text conversion 🎤, and text-to-speech synthesis 🔈
[ICML2026 Spotlight] UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
[CVPR'26 Demo] Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
A Unified Framework for Image-to-Graph Generation. Paper accepted @ ECCV22.
Code release for Ming-UniVision: Joint Image Understanding and Geneation with a Continuous Unified Tokenizer
多模型视觉理解 MCP 服务器,为不支持图片理解的 AI 编码模型提供视觉能力:分析截图、报错、UI 与文档,可接入多家主流视觉大模型。Multi-model vision MCP server that adds image understanding to AI coding models without native vision — analyze screenshots, errors, UI and documents via major vision LLM providers.
Official implementation of "UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis" - A unified medical vision-language model that integrates multimodal understanding and generation capabilities.
WACV 2024 Papers: Discover cutting-edge research from WACV 2024, the leading computer vision conference. Stay updated on the latest in computer vision and deep learning, with code included. ⭐ support visual intelligence development!
🍑 relsim: Relational Visual Similarity | pip install relsim 🌍 (CVPR 2026)
This is the implement of the paper "DynamicVis: An Efficient and General Visual Foundation Model for Remote Sensing Image Understanding"
Latest Advances on (RL based) Multimodal Reasoning and Generation in Multimodal LLMs
A deep learning project to tell a story with an image or a video.
Collection of open datasets in computer vision.
This GitHub repository shows how to integrate openai GPT-3 language model and ChatGPT API into a Unity project. It can be a useful way to add natural language processing capabilities to your application.
A large-scale curated dataset of Visual.ly infographics with metadata and additional crowdsourced annotations for research applications in computer vision and natural language processing.
📘 [Teaching] Class CVIU78101: Introduction to Computer Vision for Image Understanding Course
HumanVLM (LLaVA-based): Foundation for Human-Scene Vision-Language Model (Journal of Information Fusion 2025)
Add a description, image, and links to the image-understanding topic page so that developers can more easily learn about it.
To associate your repository with the image-understanding topic, visit your repo's landing page and select "manage topics."