SkateBench is a playful benchmark that evaluates large language models on their ability to correctly name skateboarding tricks across 210 test cases. The current ranking lists the top-performing models such‑state, including o3-pro, Gemini-3-Flash-High, and GPT-5.1-Default, with success rates ranging from 0% to 100%.
Highlights
Ranks 44 LLMs on a niche but measurable skill: naming skateboarding tricks.
Based on 210 distinct trick‑naming tests, providing granular performance data.
Top models (e.g., o3-pro, Gemini-3-Flash-High, GPT-5.1-Default) achieve near‑perfect scores.
Visualizes success rates on a 0%–100% scale for quick comparison.
auto-generated
via Ranking Models By Skateboarding Knowledge
Context
Audience
AI researchers, machine learning engineers, and data scientists interested in novel LLM evaluation methods or domain‑specific language understanding
LLM evaluation frameworksOpenAI API documentationHugging Face model hubPrompt engineering guidesBenchmark leaderboards
Discover Similar Content
tbench.ai
Terminal-Bench
A benchmark for terminal agents
github.com
GitHub - vectara/hallucination-leaderboard: Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents
Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents - vectara/hallucination-leaderboard
z.ai
GLM-5: From Vibe Coding to Agentic Engineering
GLM-5 is a 744B-parameter MoE model (40B active) from Zhipu AI, scaled up from GLM-4.5's 355B with 28.5T pre-training tokens and DeepSeek Sparse Atten...