Skip to main content

Best models

Best AI model rankings

Browse BenchLM ranking surfaces by benchmark category, workflow, provider, license, and value.

Core leaderboards

Use cases

Best Long Context AI Models

AI models ranked on sourced long-context and memory benchmarks including LongBench v2, MRCRv2, AI-Needle, Graphwalks, and MMLongBench-Doc.

Best Tool Use & Function Calling Models

AI models ranked on sourced tool-use benchmarks including BFCL v4, MCP Atlas, Toolathlon, Tau Bench, and related tool-calling evaluations.

Best AI Models for Web Research

AI models ranked on sourced browsing and research benchmarks including BrowseComp, WebArena, and WebVoyager.

Best Computer Use AI Models

AI models ranked on sourced computer-use and GUI benchmarks including OSWorld-Verified, ScreenSpot Pro, WebArena, WebVoyager, and Vision2Web.

Best Document AI Models

AI models ranked on sourced document AI and OCR benchmarks including OfficeQA Pro, OmniDocBench 1.5, CC-OCR, and MMLongBench-Doc.

Best Image Understanding Models

AI models ranked on sourced image-understanding benchmarks including MMMU-Pro, RealWorldQA, AI2D, CountBench, RefCOCO, and related grounding evaluations.

Best Frontend & App Dev Models

AI models ranked on sourced frontend and app-development benchmarks including React Native Evals, Design2Code, Vision2Web, and related web-task evaluations.

Best Factuality AI Models

AI models ranked on sourced factuality and hallucination-adjacent benchmarks including SimpleQA, HLE without tools, and Facts-VLM.

Best LLMs for Research

AI models ranked for research work — hard knowledge, agentic web research, and deep search benchmarks.

Best LLMs for Roleplay

AI models ranked for roleplay and persona work — instruction adherence, persona consistency, and creative response benchmarks.

Best LLMs for Data Analysis

AI models ranked for data analysis — quantitative reasoning, discrete reasoning over text, and analysis-code generation.

Model groups

Best Open Source LLMs

Top open weight AI models you can download and run locally, ranked by benchmark performance.

Best Proprietary LLMs

Top proprietary/closed-source AI models ranked by benchmark performance.

Best Reasoning AI Models

Top AI models with dedicated reasoning capabilities, ranked by benchmark performance.

Best OpenAI Models

All OpenAI models ranked by benchmark performance — GPT-5, GPT-4o, o1, o3, and more.

Best Anthropic Models

All Anthropic Claude models ranked by benchmark performance.

Best Google AI Models

All Google Gemini and Gemma models ranked by benchmark performance.

Best Meta AI Models

All Meta Llama models ranked by benchmark performance.

Best DeepSeek Models

All DeepSeek models ranked by benchmark performance.

Best AI Models Overall

The top AI models ranked by overall benchmark performance across all categories.

Best Large Context Window LLMs

AI models with the largest context windows (200K+ tokens), ranked by benchmark performance.

Best Chinese AI Models

A live ranking of AI models from Chinese labs, using the same current public ranking contract as the overall leaderboard.

European AI Models

European AI models from Mistral, H Company, LightOn, and Aleph Alpha — ranked models first, then tracked sparse rows.

Best Non-Reasoning LLMs

Top standard AI models (no chain-of-thought reasoning) ranked by benchmark performance. Faster and cheaper than reasoning models.

Best Mistral Models

All Mistral AI models ranked by benchmark performance — Mistral Large, Mixtral, and more.

Best xAI Grok Models

All xAI Grok models ranked by benchmark performance.

Best Alibaba Qwen Models

All Alibaba Qwen models ranked by benchmark performance.

Best LLMs for AI Agents

AI models ranked for agentic work — tool use, browsing, computer use, and long-horizon task execution.

Best LLMs for Writing

AI models ranked for writing quality — instruction following, tone control, and human preference scores.

Best LLMs for Translation

AI models ranked for translation and multilingual work, from BenchLM's multilingual benchmark category.

Best Multimodal LLMs

AI models ranked for multimodal understanding — images, documents, charts, and grounded visual reasoning.

Value rankings