Z Ai: Z.ai: GLM 5V Turbo

by Z Ai multimodal coding agentic image-input video-input

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,... Native multimodal vision-language understanding Strong vision-based coding from screenshots and diagrams Long-horizon planning with visual context Agentic task execution across modalities Higher price than text-only GLM models Vision features add latency Vision-based coding and UI development Multimodal agent workflows Screenshot-to-code tasks

Choose a model to compare against Z.ai: GLM 5V Turbo

Specifications

Specifications for Z.ai: GLM 5V Turbo
LabZ Ai
Context window202,752
Input price $1.20/1M
Output price $4.00/1M
Release2026-04-01 00:00:00

Strengths

  • Native multimodal vision-language understanding
  • Strong vision-based coding from screenshots and diagrams
  • Long-horizon planning with visual context
  • Agentic task execution across modalities

Weaknesses

  • Higher price than text-only GLM models
  • Vision features add latency

Best for

  • Vision-based coding and UI development
  • Multimodal agent workflows
  • Screenshot-to-code tasks

In Depth: Z.ai: GLM 5V Turbo

Summary

Z.ai: GLM 5V Turbo is an AI model from Z Ai.

Released 2026-04-01 00:00:00. It currently appears in the Overall category on LMRank and 1 other category. It supports Image, text, video input and produces Text output, with a context window of 202.8K tokens. Input pricing is $1.20 per 1M tokens and output is $4.00 per 1M tokens on OpenRouter.

Sources & Further Reading

Related Models