Glint AI Studio

Glint-VL-Video-8BA vision-language model for long-video understanding

Glint-VL-Video-8B is the first large video-understanding model to use codec streams as its visual unit. Built on our independently developed and open-source LLaVA-OneVision stack, it identifies key events, locates time ranges, and extracts evidence from long videos for review, retrieval, and process analysis.

Open-source foundationLLaVA-OneVision 2.0

Open research and proprietary models form one foundation

Research spans VLMs, visual foundations, multimodal embeddings, multimodal RAG, face recognition, and 3D vision. Choose a direction to view its code, papers, and project pages.

LLaVA-OneVision-1.5

A vision-language model training framework with fully open training and evaluation datasets, training code, and model weights.

ViCToR

A visual-token reconstruction approach that improves multimodal model performance on fine-grained visual understanding, chart interpretation, and document parsing.

StreamingRVOS

An LLM-based streaming referring video segmentation method for real-time, continuous segmentation of referred objects in long videos, suitable for video analytics and robotic perception.

Full-stack model fine-tuning services, tailored to your business

The LLaVA CommunityThe LLaVA community is a major force for innovation in open-source multimodal AI, advancing the deep integration of visual understanding and large language models. Its key contribution is the development of Visual Instruction Tuning, enabling models not only to recognize image content but also to conduct multi-turn question answering, logical analysis, and task execution across images, documents, and video. By opening model code, training methods, data-construction practices, and evaluation systems, LLaVA has substantially lowered the barrier to multimodal research and application development while establishing an important foundation for continued progress in image understanding, OCR, chart analysis, long-video understanding, and spatial reasoning. Through an open, reproducible, and extensible technical ecosystem, the LLaVA community is accelerating multimodal AI innovation in customer service, content creation, education, industrial inspection, robotics, and other real-world settings. provides a full model-optimization stack spanning foundational training, parameter-efficient fine-tuning, and knowledge distillation, closed by objective evaluation and production deployment. Data stays on site while capability accumulates in the customer model.

Model Training

Build datasets, training recipes, and version baselines around domain data and target tasks, with support for 4B and 8B VLM training.

  • Data governance and multimodal annotation
  • Training monitoring and checkpoint recovery
  • Model versions and experiment records

Model Fine-Tuning

Use SFT, LoRA, and QLoRA to adapt general models to customer tasks, domain knowledge, and output requirements.

  • Cost-efficient parameter adaptation
  • Private blind tests and A/B comparison
  • Controlled quality, cost, and delivery cycle

Model Distillation

Transfer knowledge and business reasoning from larger models into smaller ones while balancing quality, latency, and private deployment cost.

  • Distillation for LLMs below 32B
  • Quantization, compilation, and inference tuning
  • Cloud, private, and edge deployment

Model optimization delivery flow

  1. 01
    Data Preparation
  2. 02
    Model OptimizationTrain / Tune / Distill
  3. 03
    Evaluate and Align
  4. 04
    Deploy and Deliver

Deliver model capability into operations

City governance orchestrates video, image, and multimodal capabilities through algorithms. Finance reshapes operational control with an integrated cloud-edge-device architecture while processing documents locally.

Glint-VL-Video-8BGlint-VL-SE-8B

Visual Intelligence for City Management

Combines CV and multimodal models around a perception—understanding—analysis—alert—governance workflow, turning distributed visual resources into searchable, analyzable, and actionable business information.

  • Video-8B understands long-video events, locates time ranges, and extracts evidence
  • SE-8B strengthens critical-event, person, vehicle, and fine-grained attribute recognition
  • Cloud-edge collaboration closes the loop from spatiotemporal analysis to proactive risk discovery
Glint-VL-BK-8B

Intelligent Finance

Combines CV and multimodal models with edge computing nodes in a cloud-edge-device architecture, providing unified capabilities for operations, security, internal control, and retail banking while improving efficiency and accelerating digital transformation.

  • BK-8B continuously understands branches, vaults, and office buildings
  • Connects human behavior, critical events, locations, and supporting evidence
  • Completes delivery with a local 32B LLM, OCR, and an AI gateway

Contact us now to build your enterprise model

Tell us about your business scenario, data conditions, and deployment requirements. Our original model team will plan your model fine-tuning, evaluation, and private deployment path.