SkateBench is a playful benchmark that evaluates large language models on their ability to correctly name skateboarding tricks across 210 test cases. The current ranking lists the top-performing models such‑state, including o3-pro, Gemini-3-Flash-High, and GPT-5.1-Default, with success rates ranging from 0% to 100%.
Highlights
Ranks 44 LLMs on a niche but measurable skill: naming skateboarding tricks.
Based on 210 distinct trick‑naming tests, providing granular performance data.
Top models (e.g., o3-pro, Gemini-3-Flash-High, GPT-5.1-Default) achieve near‑perfect scores.
Visualizes success rates on a 0%–100% scale for quick comparison.
auto-generated
via Ranking Models By Skateboarding Knowledge
Context
Audience
AI researchers, machine learning engineers, and data scientists interested in novel LLM evaluation methods or domain‑specific language understanding