LLaVA-OneVision 2.0
A full-frame-rate vision-language model that unifies image and video understanding.
Glint-VL-Video-8B is the first large video-understanding model to use codec streams as its visual unit. Built on our independently developed and open-source LLaVA-OneVision stack, it identifies key events, locates time ranges, and extracts evidence from long videos for review, retrieval, and process analysis.
Open-source foundation:LLaVA-OneVision 2.0
Research spans VLMs, visual foundations, multimodal embeddings, multimodal RAG, face recognition, and 3D vision. Choose a direction to view its code, papers, and project pages.
A full-frame-rate vision-language model that unifies image and video understanding.
A vision-language model training framework with fully open training and evaluation datasets, training code, and model weights.
A visual-token reconstruction approach that improves multimodal model performance on fine-grained visual understanding, chart interpretation, and document parsing.
An LLM-based streaming referring video segmentation method for real-time, continuous segmentation of referred objects in long videos, suitable for video analytics and robotic perception.
The first codec-stream-based vision encoder to unify image and video understanding tasks.
A region-level supervised visual representation learning method that delivers clear gains in segmentation, dense detection, and visual perception for multimodal large language models.
A general-purpose visual foundation model trained with a multi-label signal margin loss for visual representation learning.
A universal and compact representation learning framework for image retrieval and the first work to introduce margin-based learning to ViTs.
A vision encoder that unifies image and video understanding tasks, advancing toward a unified ViT foundation model.
A universal multimodal embedding model that uses an MLLM-as-a-Judge mechanism to mine hard negatives with soft matching scores and achieve finer-grained cross-modal semantic alignment.
A universal embedding learning framework for multimodal large language models that combines two-stage text discriminative knowledge distillation and contrastive learning while mitigating false negatives.
A multimodal interleaved document transformation paradigm that combines real and synthetic text to build large-scale interleaved image-text training datasets.
An alignment framework that improves compositional understanding in vision-language pretraining through self-distillation and a grounded image-text contrastive loss that mitigates catastrophic forgetting.
A vision-language pretraining method that adaptively combines original and generated captions to improve the robustness and generalization of CLIP-like models on noisy data and zero-shot retrieval tasks.
A vision-language representation model built on the post-Transformer RWKV architecture, combining the parallel training efficiency of Transformers with the efficient inference of RNNs.
An efficient CLIP distillation method that applies image-semantic balancing to reduce transfer bias and transfer knowledge from large vision-language models to smaller models.
A cross-modal alignment framework for text-based person retrieval, applicable to security monitoring and cross-camera person search.
A comprehensive benchmark for document retrieval-augmented generation (RAG) systems, with realistic test sets and evaluation protocols for assessing document-level retrieval and generation quality.
An efficient training method for large-scale face recognition that can be applied directly to face recognition model training with hundreds of millions of identities.
A validation-free metric for estimating the intrinsic potential of face recognition datasets. It predicts whether a dataset can produce a high-performance model without full training, reducing the cost of large-scale face data cleaning and selection.
A world-grounded human motion recovery method based on gravity-view coordinates. It robustly recovers 3D human motion with absolute ground position from monocular video for motion capture, AR/VR, and related applications.
A streaming motion generation framework based on a diffusion autoregressive model for real-time motion generation in games, animation, and digital humans.
A unified framework that uses scene priors from video diffusion models for radiance field reconstruction, improving robust 3D reconstruction from sparse views for AR/VR content creation and digital twins.
A method that uses 2D motion data to improve 3D motion generation, producing diverse 3D motions from 2D-only annotations while reducing the cost of collecting 3D motion data.
A framework that uses 2D data pretraining for monocular multi-view 3D motion recovery, reducing reliance on expensive motion-capture equipment.
The LLaVA CommunityThe LLaVA community is a major force for innovation in open-source multimodal AI, advancing the deep integration of visual understanding and large language models. Its key contribution is the development of Visual Instruction Tuning, enabling models not only to recognize image content but also to conduct multi-turn question answering, logical analysis, and task execution across images, documents, and video. By opening model code, training methods, data-construction practices, and evaluation systems, LLaVA has substantially lowered the barrier to multimodal research and application development while establishing an important foundation for continued progress in image understanding, OCR, chart analysis, long-video understanding, and spatial reasoning. Through an open, reproducible, and extensible technical ecosystem, the LLaVA community is accelerating multimodal AI innovation in customer service, content creation, education, industrial inspection, robotics, and other real-world settings. provides a full model-optimization stack spanning foundational training, parameter-efficient fine-tuning, and knowledge distillation, closed by objective evaluation and production deployment. Data stays on site while capability accumulates in the customer model.
Build datasets, training recipes, and version baselines around domain data and target tasks, with support for 4B and 8B VLM training.
Use SFT, LoRA, and QLoRA to adapt general models to customer tasks, domain knowledge, and output requirements.
Transfer knowledge and business reasoning from larger models into smaller ones while balancing quality, latency, and private deployment cost.
Model optimization delivery flow
City governance orchestrates video, image, and multimodal capabilities through algorithms. Finance reshapes operational control with an integrated cloud-edge-device architecture while processing documents locally.
Combines CV and multimodal models around a perception—understanding—analysis—alert—governance workflow, turning distributed visual resources into searchable, analyzable, and actionable business information.
Combines CV and multimodal models with edge computing nodes in a cloud-edge-device architecture, providing unified capabilities for operations, security, internal control, and retail banking while improving efficiency and accelerating digital transformation.
Tell us about your business scenario, data conditions, and deployment requirements. Our original model team will plan your model fine-tuning, evaluation, and private deployment path.