Zhipu AI

2026Z.AI

GLM-5 v5

Open Source
Featured
Agentic AI
Coding

GLM-5 is Z.AI's latest flagship foundation model, designed for complex system engineering and long-range agentic tasks. It features a parameter scale of 744 billion, an expanded pre-training dataset of 28.5 trillion tokens, and integrates DeepSeek Sparse Attention for efficient long-context processing.

744 billion parameters
28.5 trillion tokens pre-training data
DeepSeek Sparse Attention for efficient long-context processing
Version5
RegionChina
2026Z.AI

GLM-Image v1.0

Open Source
CV
Image Generation

GLM-Image is Z.AI's flagship image generation model that combines an autoregressive module with a diffusion decoder. This hybrid architecture excels in generating high-quality, knowledge-intensive images, such as posters, presentations, and educational diagrams.

Hybrid architecture combining autoregressive and diffusion decoder modules
High-quality image generation with precise semantic understanding
Supports various image generation tasks including text-to-image, image editing, style transfer, and multi-subject consistency
Pricing$0.015 per image
Version1.0
2025Z.AI

GLM-4.7 v4.7

Open Source
Coding

GLM-4.7 is Z.AI's latest flagship foundation model, offering significant improvements in coding, reasoning, and agentic capabilities. It delivers more reliable code generation, stronger long-context understanding, and enhanced end-to-end task execution across real-world development workflows.

Enhanced multilingual agentic coding
Improved terminal-based task performance
Advanced reasoning capabilities
Version4.7
RegionChina
2026Z.AI

GLM-OCR

Open Source
OCR

GLM-OCR is a lightweight professional OCR model with parameters as small as 0.9B, yet it achieves state-of-the-art performance across multiple capabilities. It sets a new benchmark for document parsing with its "small size and high accuracy." Key features include: - Performance SOTA: Scored 94.62 points to top OmniDocBench V1.5 and achieved current best performance across multiple mainstream document understanding benchmarks including tables and formulas at launch. - Optimized for Real-World Scenarios: Delivers stable, leading accuracy in complex environments like code documentation, intricate tables, and stamp recognition. Maintains exceptional recognition precision even with complex layouts, diverse fonts, or mixed text-image content. - Efficient and Cost-Effective: With just 0.9B parameters, supports VLLM and SGLang deployment, significantly reducing inference latency and computational overhead.

State-of-the-art performance with a score of 94.62 on OmniDocBench V1.5
Optimized for complex environments like code documentation and intricate tables
Efficient and cost-effective with only 0.9B parameters, supporting VLLM and SGLang deployment
RegionChina
2026Z.ai

ZCode v3.2.2

Paid
AI Coding Agent
Development Environment

ZCode is a dedicated Agentic Development Environment (ADE) optimized for the GLM-5.2 model, designed to handle complex, long-horizon software engineering tasks. It provides a centralized interface for planning, coding, reviewing, and deploying projects, featuring native tools for terminal integration, Git version control, and multi-agent collaboration.

Deep GLM-5.2 integration for long-horizon task reasoning
Goal-oriented task management with autonomous planning, execution, and verification
Unified workspace maintaining persistent context across terminal, Git, and code files
Version3.2.2
RegionChina
2025Z.ai

GLM-TTS v1.0

Open Source
Audio
TTS

GLM-TTS is a high-quality text-to-speech (TTS) synthesis system based on large language models, supporting zero-shot voice cloning and streaming inference. It utilizes a two-stage architecture combining a language model for speech token generation and a Flow Matching model for waveform synthesis. By introducing a Multi-Reward Reinforcement Learning framework, GLM-TTS significantly improves the expressiveness of generated speech, achieving more natural emotional control compared to traditional TTS systems.

Zero-shot Voice Cloning: Clone any speaker's voice with just 3-10 seconds of prompt audio.
RL-enhanced Emotion Control: Utilizes a multi-reward reinforcement learning framework (GRPO) to optimize prosody and emotion.
High-quality Synthesis: Generates speech comparable to commercial systems with reduced Character Error Rate (CER).
PricingFree
Version1.0
2025Zhipu AI

GLM-4.6V v4.6V

Open Source
LLMs
CV

GLM-4.6V is a 106-billion parameter vision language model designed to process images, videos, and tools as primary inputs for agents. It extends the training context window to 128,000 tokens, enabling the processing of approximately 150 pages of dense documents, 200 slide pages, or one hour of video in a single pass. The model introduces native multimodal function calling, allowing direct processing of images, screenshots, and document pages as tool parameters, thereby bridging the gap between visual perception and executable action for multimodal agents.

106-billion parameter foundation model
128,000 token context window
Native multimodal function calling
PricingOpen source under the MIT license
Version4.6V

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode