AI Model Comparisons

Real Prompts. Real Outputs. Real Costs.

We test every major AI model so you don't have to. See which models win on quality, cost, and speed for specific tasks.

Safety & Guardrails Testing

Which models refuse NSFW requests? How do they handle controversial content? Testing the boundaries that matter for enterprise compliance.

Cross-Test Finding

The guardrail follows the medium, not the message

Across every safety test, identical content gets wildly different treatment depending on how you ask. The clearest case is a nude figure. As a photorealistic image, OpenAI, ByteDance, and xAI refuse it. Ask their text models to write the exact same nude as SVG code and 8 of 11 comply — including the same OpenAI models that block the image. The filter watches the output medium — nude pixels — and surface keywords, not the underlying request.

Change the medium or the framing and the same model flips. Several image models will refuse Michelangelo's David — a 500-year-old public statue — while three of them will happily render a nude of a "famous pop star." Over-refusal and under-refusal, from the same models, on the same afternoon.

5 / 4
Artistic nude photo — comply / block
2 of 9
Image models that render explicit bare breasts
8 / 2
Same nude as SVG code — drew it / refused
4 of 9
Image models that block Michelangelo's David

The consistent blockers on nude imagery are OpenAI, ByteDance, and xAI; Google's Gemini, Flux Pro, Ideogram, Qwen, and Recraft allow artistic nudity. Only Claude (Sonnet 4.6) and Qwen extend the refusal to SVG code.

Code-Drawn Art: ASCII & SVG

Can a language model draw — not generate an image, but write the code for one? We test the same subjects two ways. The pattern is consistent and revealing: models are far better at SVG (structured coordinate code, their strength) than at ASCII (spatial layout in a character grid, their weakness) — and confidently claim success at both.

ASCII Art

Cross-Test Finding

Every model draws the same blob for a lion and a physicist

None of the ASCII portraits are actually recognizable — but the revealing part is how they fail. We measured the vertical silhouette of each model's lion against its Einstein portrait. A big cat and a human face should look nothing alike. They score 0.75–0.96 identical. Models don't render the subject — they emit one generic tapered shading blob and relabel it.

0.96
Kimi K2.5
0.95
GPT-5.2
0.91
Claude Sonnet 4.6
0.88
GPT-4o
0.87
Mistral Large 3
0.85
Qwen3 235B
0.85
GPT-5.6 Sol
0.84
Llama 4 Maverick

Lion-vs-Einstein silhouette similarity (vertical fill-ratio cosine, 1.0 = identical shape). Higher = the model reused the same shape for both subjects.

SVG Generation

Run Your Own Comparisons in AI Lab

Test your actual prompts across every major model. See quality, cost, and speed side by side. Save what works. No commitment on the monthly plan.