GPT-5.4 wins the broad comparison. Claude Opus 4.6 has the higher coding estimate. GPT-5.4 has the slightly stronger agentic result and the lower price. Both have newer successors, so this is now a version-specific decision rather than a verdict on OpenAI versus Anthropic.
Current snapshot
| Metric | GPT-5.4 | Claude Opus 4.6 |
|---|---|---|
| Overall rank | #9 | #15 |
| Overall score | 74.18 Supported | 68.52 Supported |
| Coding rank | #15 | #11 |
| Coding score | 64.17 Supported | 65.66 Estimated |
| Agentic rank | #18 | #20 |
| Agentic score | 59.82 Supported | 58.87 Supported |
| Price in/out | $2.50/$15 | $5/$25 |
| Context window | 1.05M | 1M |
The evidence label prevents a false tie-break. Opus 4.6 is 1.49 points ahead in coding, but that row is Estimated. GPT-5.4's coding row is Supported. The right conclusion is that Claude is a strong coding candidate, not that a small modeled gap guarantees it will be better on every repository.
Raw benchmark comparison
| Benchmark | GPT-5.4 | Claude Opus 4.6 | Lead |
|---|---|---|---|
| HLE | 48 | 53 | Claude |
| GPQA | 92.8 | 91.3 | GPT |
| MMLU-Pro | 93 | 82 | GPT |
| SWE-bench Pro | 57.7 | 74 | Claude |
| SWE-bench Verified | 84 | 80.8 | GPT |
| LiveCodeBench | 84 | 76 | GPT |
| Terminal-Bench 2.0 | 75.1 | 65.4 | GPT |
| OSWorld-Verified | 75 | 72.7 | GPT |
| BrowseComp | 82.7 | 83.7 | Claude |
| SimpleQA | 97 | 72 | GPT |
| LongBench v2 | 95 | 92 | GPT |
| MRCRv2 | 97 | 92 | GPT |
| IFEval | 96 | 95 | GPT |
| MMMU-Pro | 81.2 | 77.3 | GPT |
| OfficeQA-Pro | 96 | 94 | GPT |
The rows disagree because they test different things. Claude's SWE-bench Pro lead is large. GPT-5.4 leads SWE-bench Verified and LiveCodeBench. A model can win one harness and lose another without either result being wrong. BenchAlign normalizes each benchmark against its own field before combining them, rather than averaging raw percentages from tests with different difficulty.
Coding
Opus 4.6 ranks eleventh on the current coding surface at 65.66. GPT-5.4 ranks fifteenth at 64.17. If your work is repository engineering, put both on the same issue set and measure accepted patches, tests passed, regressions, and repair loops.
The comparison should also include current versions when procurement allows. Claude Opus 4.8 ranks fourth in coding at 73.49 with Supported evidence. GPT-5.6 Sol ranks third at 74.08 with Estimated evidence. Choosing Opus 4.6 or GPT-5.4 without pricing the successors can save money, but it is no longer a frontier-only bake-off.
Agentic work
GPT-5.4 scores 59.82 and Opus 4.6 scores 58.87. Both rows are Supported and the gap is under one point. Treat them as close until a workflow evaluation says otherwise.
Tool reliability is not the same as one benchmark pass rate. Measure permission handling, recovery after a failed call, intervention rate, wall-clock completion, and cost per successful run. GPT-5.4's lower token price makes a tie operationally meaningful.
Price
GPT-5.4 costs $2.50/$15 per million input/output tokens. Claude Opus 4.6 costs $5/$25. For one million input and 200,000 output tokens, that is $5.50 for GPT-5.4 and $10 for Opus 4.6.
Claude can still be cheaper if it avoids enough retries or human review. That is a total-cost question, not a token-price claim.
Which one should you use?
Use GPT-5.4 when you want the stronger supported overall and agentic profile at the lower price. It is the safer default for broad workloads between these exact two versions.
Use Claude Opus 4.6 when coding is the center of the workload and it wins a direct repository trial. Its coding estimate is higher and SWE-bench Pro result is materially stronger.
Use neither by default when the current frontier is in scope. Compare Claude Opus 4.8, GPT-5.6 Sol, and Claude Fable 5 too. A version-specific SEO page should answer the version-specific question without freezing the entire market in March.
→ Compare models · Coding leaderboard · Agentic leaderboard · Overall ranking
Reader questions
Frequently asked questions
01Is Claude Opus 4.6 better than GPT-5.4?
GPT-5.4 ranks higher overall at 74.18 versus 68.52 for Claude Opus 4.6, both with Supported evidence. Claude Opus 4.6 ranks higher in coding at 65.66 versus 64.17, but its coding row is Estimated. GPT-5.4 has the slightly stronger Supported agentic score.
02Where does Claude Opus 4.6 beat GPT-5.4?
Claude Opus 4.6 ranks above GPT-5.4 on BenchLM's coding surface and leads raw rows including HLE and SWE-bench Pro. Its coding evidence is Estimated, so the exact gap should be tested on your repositories.
03How much does Claude Opus 4.6 cost compared with GPT-5.4?
Claude Opus 4.6 costs $5 input and $25 output per million tokens. GPT-5.4 costs $2.50/$15. Claude is 2x the input price and about 1.7x the output price.
04Are Claude Opus 4.6 and GPT-5.4 still frontier models?
They remain capable models, but both have newer family alternatives. Claude Opus 4.8 ranks fourth overall, while GPT-5.6 Sol ranks third. Use this comparison when these exact versions are in your stack, not as a proxy for the current Anthropic and OpenAI frontier.
05Should I use Claude Opus or Claude Sonnet?
Claude Opus 4.6 scores 68.52 overall, while Claude Sonnet 4.6 scores 65. Opus has the stronger broad row; Sonnet costs less at $3/$15. Test whether the capability difference reduces enough errors or review time to cover the price gap.
Continue with live BenchLM data
Share or save