AI Model Comparisons

Real Prompts. Real Outputs. Real Costs.

We test every major AI model so you don't have to. See which models win on quality, cost, and speed for specific tasks.

Technical Capability Benchmarks

SVG generation, ASCII art, structured output — testing edge cases that reveal model limitations and capabilities.

Cross-Test Finding

Every model draws the same blob for a lion and a physicist

None of the ASCII portraits are actually recognizable — but the revealing part is how they fail. We measured the vertical silhouette of each model's lion against its Einstein portrait. A big cat and a human face should look nothing alike. They score 0.75–0.96 identical. Models don't render the subject — they emit one generic tapered shading blob and relabel it.

0.96
Kimi K2.5
0.95
GPT-5.2
0.91
Claude Sonnet 4.6
0.88
GPT-4o
0.87
Mistral Large 3
0.85
Qwen3 235B
0.85
GPT-5.6 Sol
0.84
Llama 4 Maverick

Lion-vs-Einstein silhouette similarity (vertical fill-ratio cosine, 1.0 = identical shape). Higher = the model reused the same shape for both subjects.

Run Your Own Comparisons in AI Lab

Test your actual prompts across every major model. See quality, cost, and speed side by side. Save what works. No commitment on the monthly plan.