Benchmark profile
BullshitBench v2
A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.
How BenchLM shows BullshitBench v2
BenchLM mirrors the published BullshitBench v2 leaderboard using the official snapshot generated on July 16, 2026 at 9:22 PM UTC. The public view reports per-model clear-pushback rates across 100 nonsense prompts, scored by a 3-judge panel.
BullshitBench is a useful reasoning sanity check, but BenchLM currently keeps it display only rather than weighted. The public leaderboard is highly variant-specific and exposes reasoning-effort settings directly, so BenchLM treats it as a mirrored external benchmark instead of a canonical ranking input.
Clear pushback rate on BullshitBench v2 — July 16, 2026 at 9:22 PM UTC
BenchLM mirrors the published clear pushback rate view for BullshitBench v2. Claude Opus 4.8 (none) leads the public snapshot at 95% , followed by Claude Opus 4.8 (xhigh) (94%) and Claude Sonnet 4.6 (high) (91%). BenchLM does not use these results to rank models overall.
Claude Opus 4.8 (none)
Anthropic
anthropic/claude-opus-4.8@reasoning=none
Claude Opus 4.8 (xhigh)
Anthropic
anthropic/claude-opus-4.8@reasoning=xhigh
Claude Sonnet 4.6 (high)
Anthropic
anthropic/claude-sonnet-4.6@reasoning=high
Clear pushback rate table (182 models)
ScoreThe published BullshitBench v2 snapshot places Claude Opus 4.8 (none) first at 95%. The third row is 4.0 points behind. The broader top-10 range is 16.0 points, so the table still separates the published systems.
182 models have been evaluated on BullshitBench v2. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. BullshitBench v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About BullshitBench v2
Year
2025
Tasks
Nonsensical and flawed prompts across multiple domains
Format
Prompt challenge and refusal evaluation
Difficulty
Robustness and critical reasoning
BullshitBench evaluates a crucial real-world capability: knowing when NOT to answer. Models that score highly recognize flawed premises, impossible physics scenarios, and logical contradictions rather than hallucinating plausible-sounding responses. V2 includes harder and more diverse challenge categories.
BenchLM freshness & provenance
Version
BullshitBench v2 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does BullshitBench v2 measure?
A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.
Which model leads the published BullshitBench v2 snapshot?
Claude Opus 4.8 (none) currently leads the published BullshitBench v2 snapshot with 95% clear pushback rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on BullshitBench v2?
182 AI models are included in BenchLM's mirrored BullshitBench v2 snapshot, based on the public leaderboard captured on July 16, 2026 at 9:22 PM UTC.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.