
A question about a wind-powered vehicle produced the sharpest split in this small test. Asked whether a land vehicle can travel directly downwind faster than the wind, GPT-5.6 Luna returned no text after an observed API error at 120.01 seconds. Gemini 3.5 Flash and GLM-5.2 each answered yes and offered extended explanations of a wheel-driven propeller drawing on the speed difference between air and ground.
Disclosure: I work with OrcaRouter and used it to run this evaluation. It made switching between models easier while exposing live usage and provider-rate costs.
AI-generated illustration featuring official model logos; logos and model names are used descriptively and remain the property of their respective owners.
Compare the models in this article through OrcaRouter’s model catalog.
That does not establish that either successful model is more reliable, or that GPT-5.6 Luna cannot answer the question. It is one question-model call. But it is exactly the kind of concrete result a compact stress test can reveal: on this occasion, one route failed and two did not.
What was tested
The case study covered 33 selected calls: three models—GPT-5.6 Luna, Gemini 3.5 Flash and GLM-5.2—each given the same 11 questions. The set mixed physics, probability twists, state tracking, long-context consistency checking, strict formatting instructions and one current-events prompt requesting web search. Every model-question cell has a sample size of one. Results should therefore be read as observations from this selected run, not stable estimates of model behavior.
The test was conducted through an API/gateway route. Its behavior should not be generalized to consumer subscription products.
Several tasks produced a three-way match. All three models spotted the altered premise in the wolf-goat-cabbage puzzle: because the boat could carry the farmer and all three items, the answer was one crossing. All three also followed a Chinese instruction requiring exactly three sentences beginning with “第一,” “第二,” and “第三.”
On the modified Monty Hall question, all three correctly treated the host’s random goat reveal differently from the usual informed-host setup, giving a switching probability of one-half. GPT-5.6 Luna delivered the compact version; Gemini and GLM showed longer conditional-probability workings.
A small task matrix, not a value leaderboard
| Task type | Observed result |
|---|---|
| Strict JSON and constrained writing | All three returned the requested kettle JSON structure and avoided the banned word in the coffee-history prompt. |
| Sequential state tracking | All three reached E in seat 3 and B in seat 5, while showing the swaps. |
| Calendar counterfactual | All three gave July 3, 2026, by subtracting a 13-day Julian–Gregorian difference. |
| Physics: intermediate-axis rotation | All three described the characteristic instability and distinguished stable rotations about the other principal axes. |
| Physics: faster-than-wind vehicle | Gemini and GLM answered; GPT-5.6 Luna’s selected call errored with no text response. |
The changelog-reading task showed a more interesting difference in style and interpretation. GPT-5.6 Luna identified two conflicts: the supposedly removed export button later receiving a UI fix, and the claim that May 2 was exactly two weeks after March 21. Gemini listed those two and added an engineering-headcount issue, while GLM also treated the “Q2” heading’s inclusion of March as inconsistent. The output illustrates why readers should inspect the answer, not merely a score: models can agree on the core contradictions while disagreeing about what counts as an inconsistency.
What the automated judge scored
The available scores are reference-guided DeepSeek v3 judgments, and human review is pending. They are useful audit evidence, not final quality ground truth.
For the selected successful calls under judge prompt version v3, 31 received an accuracy score of 5. The exception was GLM-5.2’s current-events response, scored 3; GPT-5.6 Luna’s failed wind-vehicle call had no successful answer to score.
The Evidence Pack’s judge audit is the only approved scoring record here; it does not replace expert checking of factual claims, reasoning, or completeness.

Reproducible data figure from this article’s selected API records; unavailable cost fields are shown as unavailable.
Search and speed: useful operational clues, not winners
The freshness question explicitly requested web search. Search was requested for all three calls, but was marked effective only for Gemini 3.5 Flash. GPT-5.6 Luna and GLM-5.2 still produced answers, yet their selected records mark search as ineffective; GLM explicitly said it could not browse live web. This describes the tested API route, not the providers’ consumer products or their broader search capabilities.
Selected-call latency also varied sharply. Gemini’s median was 3.5 seconds across 11 calls, GPT-5.6 Luna’s was 10.06 seconds across 11 calls, and GLM-5.2’s was 26.95 seconds across 11 calls. GPT had 10 OK calls and one error; Gemini and GLM each had 11 OK calls. These are descriptive figures from this run, not general performance rankings.
Cost remains unresolved
The internal summary recorded $0.0331 in observed billed USD for GPT-5.6 Luna’s selected calls. It contained no verified Gemini or GLM token prices. Accordingly, this is not a price-performance estimate, and no all-in cost winner can be claimed. OpenAI’s API pricing page is an official source for its own schedule, but it cannot supply the missing cross-provider comparison.
Nor should vendor labels or implied effort settings be treated as equal compute budgets across providers.
Practical takeaways
- For tightly constrained formatting, simple state tracking and several classic reasoning tasks, all three produced usable answers in this one-pass sample.
- For a workflow requiring fresh web-backed answers, test the exact API route: only Gemini’s selected search call was marked effective here.
- For fault-sensitive use, retain retries and fallbacks. The wind-vehicle failure is only one observation, but it is a reminder that a polished comparison should count failures as well as answers.
- Read outputs where interpretation matters. The changelog task showed differing definitions of what constituted an internal inconsistency.

Editorial illustration; it frames different trade-offs and is not test evidence.
Limitations
This was one selected answer per model-question cell, across 11 questions and 33 calls—not a broad benchmark or a price-performance study. Latency, reliability, search and cost observations are limited to the tested API/gateway path and should not be extended to consumer products. The automated scores were reference-guided DeepSeek v3 judgments, with human review still pending. Finally, missing verified Gemini and GLM token prices rule out a complete cost comparison.
Explore the Models
Explore the current catalog on OrcaRouter Models.
This evaluation was run through OrcaRouter. The author works with OrcaRouter; model access does not imply affiliation with, endorsement by, or sponsorship from model providers.
Model names and logos are used descriptively. All trademarks belong to their respective owners.
Sources
- GPT-5.6 announcement — OpenAI; retrieved 2026-07-17.
- OpenAI API pricing — OpenAI; retrieved 2026-07-17.