source: simon willison: are ai labs pelicanmaxxing?

level: technical

dylan castillo ran a controlled experiment to check if ai labs deliberately train models to excel at drawing pelicans riding bicycles. he tested 8 animals and 6 vehicles, creating 48 prompt combinations. each prompt was run three times through seven models: gpt-5.6 terra, claude sonnet 5, gemini 3.5 flash, grok 4.5, qwen3.7-max, glm-5.2, and deepseek v4 pro. gpt-5.6 luna and gemini 3.1 flash-lite helped evaluate the outputs.

the results showed no evidence of pelicanmaxxing. pelicans on bicycles did not look better than other animal-vehicle pairs. labs were not better at drawing pelicans, bicycles, or the combination, even after adjusting for difficulty. the pelican-bicycle scenes did not appear memorized. pelicans were not drawn better than other animals, and bicycles were not drawn better than other vehicles. no lab drew the combination better than its individual pelican and bicycle performance would predict.

glm-5.2 had a small boost on the exact pelican-bicycle cell, and one sample stood out, but the effect was tiny and not statistically significant. the study provides a rigorous, data-driven answer to a common speculation in the ai community. it shows that the popular benchmark remains a useful, unbiased test of image generation capabilities.

why it matters: it confirms that a widely used informal benchmark is not being gamed, so it remains a reliable quick check for model image generation quality.


source: simon willison: are ai labs pelicanmaxxing?