Z.ai

GLM 5.3

Z.ai’s latest flagship text model for complex software engineering and long-horizon agents supports up to a 1M-token context, 128K-token output, three reasoning-effort levels, tool calling, structured output, and context caching, with stronger performance and token efficiency across coding, terminal, and multi-step professional tasks.

Modalities
Chat
Starting price
From $0.234 / call
Context
1.0M context

Z.ai

README

GLM-5.3 is Z.ai's flagship language model released on August 14, 2026 for complex software engineering and long-horizon agent tasks. It retains the GLM-5.2 base model and derives its improvements from expanded post-training. The main model accepts and produces text, provides a one-million-token context window with up to 128,000 output tokens, and supports streaming, function calling, context caching, and structured output.

The release focuses on executable tasks that resemble real professional work: understanding an environment, managing multi-step dependencies, using tools, implementing changes, and validating results. Z.ai reports a 50% gain over GLM-5.2 on its internal Code Bench, an increase from 4.6 to 28.3 on Terminal-Bench 3.0, and from 46.2 to 66.9 on DeepSWE v1.1. Reasoning is always enabled and can be configured at low, high, or max effort.

Key Capabilities

  • Stronger Complex Coding: Improves cross-module analysis, implementation, debugging, validation, and delivery in production-oriented projects.
  • Long-Horizon Agent Execution: Sustains planning and execution across environmental state, tool results, and multi-stage dependencies.
  • Always-On Reasoning: Reasons on every request and cannot disable thinking; low, high, and max balance speed, token use, and task quality.
  • Improved Token Efficiency: Achieves higher completion rates with fewer output tokens than GLM-5.2 on complex agentic coding work.
  • Function Calling: Uses external tools, terminals, services, and data sources to complete executable multi-step professional tasks.
  • Structured Output: Produces formats such as JSON so programs can continue processing and validating model results.
  • Context Caching: Reduces repeated-input cost in continuous development, code review, and multi-turn agent workflows.

Technical Strengths

FeatureBenefit
Post-Training Upgrade on the Same BaseImproves difficult real-world tasks without changing the GLM-5.2 base model.
Production-Environment Task TrainingUses tasks closer to real units of engineering and research work, including diagnosis, modification, experimentation, and validation.
Three Reasoning-Effort Levelslow, high, and max provide explicit control for different latency, cost, and difficulty requirements.
Better Performance and Token EfficiencyRaises completion rates while reducing average output-token use in complex coding-agent evaluations.
Up to 1M Context and 128K OutputOrganizes large projects, long-lived state, tool results, and substantial deliverables in one task.
Tool and Structured InterfacesFunction calling, streaming, and JSON output support observable and verifiable production agents.
Cache-Friendly Persistent WorkflowsCached-input pricing lowers the ongoing cost of repeated project background and stable instructions.

Frequently Asked Questions

What is the main difference between GLM-5.3 and GLM-5.2?

They use the same base model, while GLM-5.3’s gains come primarily from post-training. GLM-5.3 emphasizes complex software engineering, terminal operation, and realistic long-horizon agent work, delivering higher completion rates and better token efficiency in those settings. GLM-5.2 retains the flexibility to disable reasoning.

Can reasoning be disabled on GLM-5.3?

No. GLM-5.3 always reasons, and the official thinking.type value is enabled only. Existing requests that disable thinking must remove or change that setting before migrating, or the request may fail.

How should low, high, and max be selected?

Use low for relatively direct, latency-sensitive tasks; high for work needing stronger analysis with controlled cost; and max for complex coding, long-horizon agents, and difficult reasoning. max is the official default, but production selection should follow quality, latency, and token-use tests.

Can the 1M context and 128K output limits always be reached together?

The official model provides up to a 1M-token context and 128K-token output, but individual requests remain subject to total-length constraints, AIPing’s downstream route, timeouts, and service quotas. The page should not promise that every route can reach both limits simultaneously without production validation.

What security-analysis scenarios are appropriate?

Appropriate uses include authorized code auditing, vulnerability analysis, defensive testing, and remediation guidance. Marketplace examples should focus on defensive use, avoid unauthorized intrusion or destructive exploitation, and comply with LinkModel, AIPing, and applicable legal requirements.

Pricing

Token TypeLinkAI PriceOfficial Price
input$1.260000 / 1M tokens$1.400000 / 1M tokens
output$3.960000 / 1M tokens$4.400000 / 1M tokens
cache_read$0.234000 / 1M tokens$0.260000 / 1M tokens

More from Z.ai