Z.ai Releases GLM-5.3 API, Holds Pricing Flat at $1.40 Per Million Input Tokens
Chinese AI startup Z.ai has launched its GLM-5.3 model via API at unchanged rates of $1.40 per million input tokens and $4.40 per million output tokens, while benchmark scores climb seven points over its predecessor.

Chinese AI startup Z.ai has made its frontier language model, GLM-5.3, available via API. The rollout gives developers direct access for use in software agents, coding tools, and enterprise workflows, at pricing identical to the previous GLM-5.2 generation.
According to VentureBeat, developers will pay $1.40 per million input tokens and $4.40 per million output tokens. Cached input is priced at $0.26 per million tokens, with cached-input storage offered at no charge for a limited window. The API currently supports the OpenAI Chat Completions-compatible protocol for developers on existing coding subscriptions. Open model weights are expected at an unannounced future date.
On a raw per-token basis, GLM-5.3 sits well below several leading proprietary models. A combined cost of one million input tokens and one million output tokens totals $5.80 on GLM-5.3. By comparison, xAI's Grok 4.6 costs $8.00 at lower context tiers, Moonshot AI's Kimi K3 runs $18.00, Anthropic's Claude Opus 5 reaches $30.00, and OpenAI's GPT-5.6 Sol Standard commands $35.00. Lighter utility models remain cheaper still: OpenAI's GPT-5.6 Luna comes in at $1.40 combined, and Google's promotional Gemini 3.7 Flash rates undercut GLM-5.3 on price.
Independent evaluator Artificial Analysis scored GLM-5.3 at 60 on its Intelligence Index, tying Kimi K3 as the top-ranked open-weights model on the platform. That result is seven points above GLM-5.2. The evaluation also flagged a significant operational variable: GLM-5.3 is more verbose than its predecessor, pushing the estimated cost per standard benchmark task from $0.44 to $0.68, even with the token price held flat.
The Cognarah Angle
The GLM-5.3 release exposes a growing blind spot in how developers and procurement teams evaluate inference costs. Headline token prices are no longer a reliable proxy for what a model actually costs to run. Z.ai's decision to hold its per-token rate steady generates clean marketing copy, but the model's increased verbosity quietly shifts the financial burden back onto developers. When a model demands roughly fifty percent more output tokens to complete the same reasoning or coding task, a frozen token price is an effective cost increase dressed as stability.
The broader competitive picture is harder to ignore. Chinese open-weight architectures are now posting benchmark scores that match, or approach, top-tier Western proprietary models at a fraction of the price. When the capability gap narrows to near zero, the premium commanded by legacy enterprise API providers becomes very difficult to justify on technical grounds alone.
If open-weight frontier models already match proprietary performance on coding and reasoning benchmarks, what exactly are engineering teams paying six times more for?
Reporting sourced from VentureBeat. Analysis and Cognarah Angle are Cognarah's own.
Written by
StacyAI-assisted news curation. Every story is reviewed by our editors before publication.



