Model profile
Gemma 4 E4B
Evidence coverage
14 of 323 tracked benchmarks are published. 0 are verified and 14 provisional. 6 of 8 categories are measured.
- Published / tracked
- 14 / 323
- Verified
- 0
- Provisional
- 14
- Categories with evidence
- 6 / 8
Evidence by category
- Agentic1 benchmarkReported
- Coding1 benchmarkReported
- Reasoning2 benchmarksReported
- Knowledge8 benchmarksReported
- Math0 benchmarksNot measured
- Multilingual0 benchmarksNot measured
- Multimodal1 benchmarkReported
- Inst. Following1 benchmarkReported
Gemma 4 E4B ranks #155 out of 200 models on the public leaderboard with an overall score of 43.2/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.
Gemma 4 E4B is a open weight model with a 128K token context window. It uses explicit chain-of-thought reasoning, which typically improves performance on math and complex reasoning tasks at the cost of higher latency and token usage.
Gemma 4 E4B sits inside the Gemma 4 family alongside Gemma 4 31B, Gemma 4 26B A4B, Gemma 4 12B, Gemma 4 E2B. This profile currently has 14 of 323 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.
Peer position
Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.
Range 42.58–43.87
- Ling 2.6 FlashInclusionAICompare#15443.87Ling 2.6 Flash is #154 with a score of 43.87.
- Gemma 4 E4BCurrent modelGoogle#15543.2Gemma 4 E4B is #155 with a score of 43.2.
- Mistral Medium 3MistralCompare#15643.2Mistral Medium 3 is #156 with a score of 43.2.
- Sarvam 105BSarvamCompare#15742.97Sarvam 105B is #157 with a score of 42.97.
- Claude 4 SonnetAnthropicCompare#15842.79Claude 4 Sonnet is #158 with a score of 42.79.
- GPT-OSS 20BOpenAICompare#15942.74GPT-OSS 20B is #159 with a score of 42.74.
- DeepSeek R1 Distill Qwen 32BDeepSeekCompare#16042.58DeepSeek R1 Distill Qwen 32B is #160 with a score of 42.58.
Category percentile
More
Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.
Category evidence
Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.
| Category | Score | Rank | Percentile | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|---|
| AgenticWeight 22%1 benchmarkReported | Score pending | Not ranked | Not available | 22% | 1 benchmark | Reported |
| CodingRank Not rankedWeight 20%1 benchmarkReported | 52.0 | Not ranked | Not available | 20% | 1 benchmark | Reported |
| ReasoningRank Not rankedWeight 17%2 benchmarksReported | 25.4 | Not ranked | Not available | 17% | 2 benchmarks | Reported |
| KnowledgeRank Not rankedWeight 12%8 benchmarksReported | 67.4 | Not ranked | Not available | 12% | 8 benchmarks | Reported |
| MathWeight 5%0 benchmarksNot measured | Not measured | Not ranked | Not available | 5% | 0 benchmarks | Not measured |
| MultilingualWeight 7%0 benchmarksNot measured | Not measured | Not ranked | Not available | 7% | 0 benchmarks | Not measured |
| MultimodalRank Not rankedWeight 12%1 benchmarkReported | 52.6 | Not ranked | Not available | 12% | 1 benchmark | Reported |
| Inst. FollowingWeight 5%1 benchmarkReported | Score pending | Not ranked | Not available | 5% | 1 benchmark | Reported |
Benchmark Details
Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.
Agentic1 benchmark
τ²-Bench Tool-Agent-User Evaluation
Coding1 benchmark
Artificial Analysis SciCode
Reasoning2 benchmarks
Artificial Analysis Long Context Reasoning
Critical Physics Tasks
Knowledge8 benchmarks
Massive Multitask Language Understanding Professional
Graduate-Level Google-Proof Q&A
Artificial Analysis GPQA Diamond
Artificial Analysis Humanity's Last Exam
Artificial Analysis Omniscience Index
Artificial Analysis Omniscience Accuracy
Artificial Analysis Omniscience Hallucination Rate
Multimodal1 benchmark
Artificial Analysis MMMU-Pro
Inst. Following1 benchmark
Artificial Analysis IFBench
Frequently Asked Questions
How does Gemma 4 E4B perform overall in AI benchmarks?
Gemma 4 E4B has 14 published benchmark scores on BenchLM, but it does not yet have enough non-generated coverage to receive a global overall rank.
Is Gemma 4 E4B good for knowledge and understanding?
Gemma 4 E4B has visible benchmark coverage in knowledge and understanding, but BenchLM does not currently assign it a global category rank there.
Is Gemma 4 E4B good for coding and programming?
Gemma 4 E4B has visible benchmark coverage in coding and programming, but BenchLM does not currently assign it a global category rank there.
Is Gemma 4 E4B good for reasoning and logic?
Gemma 4 E4B has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.
Is Gemma 4 E4B good for agentic tool use and computer tasks?
Gemma 4 E4B has visible benchmark coverage in agentic tool use and computer tasks, but BenchLM does not currently assign it a global category rank there.
Is Gemma 4 E4B good for multimodal and grounded tasks?
Gemma 4 E4B has visible benchmark coverage in multimodal and grounded tasks, but BenchLM does not currently assign it a global category rank there.
Is Gemma 4 E4B good for instruction following?
Gemma 4 E4B has visible benchmark coverage in instruction following, but BenchLM does not currently assign it a global category rank there.
Is Gemma 4 E4B open source?
Yes, Gemma 4 E4B is an open weight model created by Google, meaning it can be downloaded and run locally or fine-tuned for specific use cases.
Which sibling models are related to Gemma 4 E4B?
Gemma 4 E4B belongs to the Gemma 4 family. Related variants on BenchLM include Gemma 4 31B, Gemma 4 26B A4B, Gemma 4 12B, Gemma 4 E2B.
Does Gemma 4 E4B have full benchmark coverage on BenchLM?
Not yet. Gemma 4 E4B currently has 14 published benchmark scores out of the 323 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.
What is the context window size of Gemma 4 E4B?
Gemma 4 E4B has a published context window of 128K, which determines how much text it can process in a single interaction.
Related Resources
Choose with this week’s evidence
Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.
Free. One email per week.