Model profiles
Open a model. Follow the evidence.
Each profile starts with a plain-language summary, then lets you move from a scale to a question and the model’s exact recorded answer.
Anthropic · Phase 2 + 3
Claude Sonnet 5
99.8% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →DeepSeek · Phase 2 + 3
DeepSeek V4 Pro
95.8% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →Google · Phase 2 + 3
Gemini 2.5 Flash Lite
95.9% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →Meta · Phase 2 + 3
Llama 3.3 70B
100.0% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →Mistral · Phase 2 + 3
Mistral Medium 3.1
100.0% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →OpenAI · Phase 2 + 3
GPT-5.4 mini
100.0% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →Qwen / Alibaba · Phase 2 + 3
Qwen 3.8 27B
100.0% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →xAI · Phase 2 + 3
Grok 4.20
95.4% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →Z.ai · Phase 2 + 3
GLM 5.2
98.6% of its main-battery outputs parsed cleanly. Open its response profile and raw-answer explorer.
View profile →