GPT-5.4 is the stronger broad and lower-cost default. Claude Opus 4.6 has the higher coding estimate. The models are close on agentic work, where both rows are Supported.
Headline comparison
| Metric | Claude Opus 4.6 | GPT-5.4 |
|---|---|---|
| Overall | 68.52 Supported | 74.18 Supported |
| Overall rank | #15 | #9 |
| Coding | 65.66 Estimated | 64.17 Supported |
| Coding rank | #11 | #15 |
| Agentic | 58.87 Supported | 59.82 Supported |
| Agentic rank | #20 | #18 |
| API price | $5/$25 | $2.50/$15 |
Where Claude wins
Claude leads raw HLE, 53 to 48, and SWE-bench Pro, 74 to 57.7. Its combined coding estimate is also higher. Those are real reasons to test it on complex repository work.
The caveat is evidence strength. Claude's coding row is Estimated, while GPT-5.4's is Supported. A 1.49-point lead is not strong enough to skip a direct trial.
Where GPT-5.4 wins
GPT-5.4 leads overall, agentic work, and price. It also leads raw SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.0, OSWorld-Verified, SimpleQA, MMLU-Pro, LongBench v2, and MRCRv2 in this matchup.
The broad pattern favors GPT-5.4. The price reinforces it: for one million input and 200,000 output tokens, GPT-5.4 costs $5.50 and Opus 4.6 costs $10.
Which should you use?
Use GPT-5.4 when you want the safer broad default between these two exact versions. Use Claude Opus 4.6 when coding dominates and it wins on your repository test.
If current frontier models are allowed, add Claude Opus 4.8, GPT-5.6 Sol, and Claude Fable 5. Opus 4.6 versus GPT-5.4 is a version-specific comparison now, not a proxy for the best model each lab sells.
→ Full version comparison · Side-by-side data · Coding ranking
Reader questions
Frequently asked questions
01Is Claude Opus 4.6 better than GPT-5.4?
GPT-5.4 is higher overall at 74.18 versus 68.52 for Claude Opus 4.6, both Supported. Claude ranks higher in coding at 65.66 versus 64.17, but Claude's coding row is Estimated. GPT-5.4 has the slightly stronger Supported agentic score.
02What is the price difference?
GPT-5.4 costs $2.50 input and $15 output per million tokens. Claude Opus 4.6 costs $5/$25, which is 2x the input price and about 1.7x the output price.
03Which is better for coding?
Claude Opus 4.6 ranks eleventh in coding at 65.66 Estimated; GPT-5.4 ranks fifteenth at 64.17 Supported. The modeled gap is small enough that a repository evaluation should decide.
Continue with live BenchLM data
Share or save