# Frozen model selection and price snapshot

**Verified:** July 30, 2026 at 18:52:24 UTC  
**Experiment variable:** model only

| Promptfoo provider ID | Exact OpenRouter model | Type | Input / 1M | Output / 1M |
|---|---|---|---:|---:|
| `openrouter:deepseek/deepseek-v4-pro` | DeepSeek V4 Pro | Open weights | $0.435 | $0.87 |
| `openrouter:z-ai/glm-5.2` | GLM 5.2 | Open weights | $0.60 | $1.25 |
| `openrouter:x-ai/grok-4.5` | Grok 4.5 | Closed | $2.00 | $6.00 |

Primary model cards:

- [DeepSeek V4 Pro](https://openrouter.ai/deepseek/deepseek-v4-pro)
- [GLM 5.2](https://openrouter.ai/z-ai/glm-5.2)
- [Grok 4.5](https://openrouter.ai/x-ai/grok-4.5-20260708)

All three accept text and support the configured `max_tokens` and high-reasoning
settings. The two open-weights labels come from their model cards, not from
price or free-tier status. Grok 4.5 is delivered through xAI's hosted API and
serves as the closed model. It replaced Claude Opus 5 before any model request
because the account's billing region cannot access Anthropic models.

Kimi K3 was the originally selected second open-weights model. Its original
score was invalidated by a shared reasoning/output ceiling, and repeated
corrected runs remained API-incomplete because of OpenRouter upstream-provider
failures. On July 30, 2026, GLM 5.2 formally replaced that row under the
precommitted replacement option. The frozen policy, cases, prompt, answer key,
rubric, and inference settings were not changed. Kimi's attempts remain in the
rerun evidence as a disclosed infrastructure confound and do not enter the
final quality or cost comparison.
