Brief IA

Qwen3.6-35B-A3B Outperforms Claude Opus 4.7 in a Humorous Test

🎨 Creative AI·Tom Levy·

Qwen3.6-35B-A3B Outperforms Claude Opus 4.7 in a Humorous Test

Qwen3.6-35B-A3B Outperforms Claude Opus 4.7 in a Humorous Test
Key Takeaways
1Alibaba's Qwen3.6-35B-A3B has surpassed Anthropic's Claude Opus 4.7 in a pelican image generation test.
2The Qwen model, using a 20.9 GB file from Unsloth, operated via LM Studio on a MacBook Pro M5.
3Despite the humorous nature of the test, Qwen also excelled in creating an SVG of a flamingo on a bicycle.
💡Why it mattersQwen's performance highlights the rapid evolution of AI model capabilities, even in unconventional tasks.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Qwen3.6-35B-A3B: A Model Ahead

In a humorous image generation test, Alibaba's Qwen3.6-35B-A3B model outperformed Anthropic's Claude Opus 4.7. This test, dubbed the "bicycle pelican benchmark," was used to compare the capabilities of the two recent models.

The Qwen3.6-35B-A3B model, utilizing the 20.9 GB quantized file Qwen3.6-35B-A3B-UD-Q4_K_S.gguf from Unsloth, operated on a MacBook Pro M5 via LM Studio and the llm-lmstudio plugin. It produced an image of a pelican deemed superior to that generated by Claude Opus 4.7, which notably ruined the bike's frame. Even when adjusting the "thinking_level" parameter to the maximum, Claude Opus failed to match Qwen's quality.

A Test Beyond the Joke

Although the pelican benchmark was designed as a joke, it revealed a correlation between the quality of the produced images and the overall utility of the models. Many people are convinced that labs are training for this absurd benchmark, although this remains unconfirmed. However, this link between the quality of pelicans and the general utility of the models has been broken.

I have immense respect for Qwen, but I strongly doubt that a 21 GB quantized version of their latest model is more powerful or useful than Anthropic's latest proprietary version. Nevertheless, for specific tasks like creating an SVG of a bicycle pelican, Qwen3.6-35B-A3B seems to be the best choice currently.

An Unexpected Result

In addition to the main test, another challenge was posed to the models: generating an SVG of a flamingo on a bicycle. Once again, Qwen took the lead, particularly thanks to a humorous comment on the produced SVG: "Sunglasses on flamingo!".

This unexpected result shows that, despite the absurd nature of the test, Qwen3.6-35B-A3B can deliver impressive performance in creative and visual tasks.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.