Skip to main content

Claude Opus 4.6 vs GPT-5.4: Where Each Model Wins

Claude Opus 4.6 vs GPT-5.4 on current overall, coding, agentic, raw benchmark, and pricing data, with the evidence caveats that decide the close calls.

Published
Last updated
Reading time
6 min
External sources
0
Tags: comparison, claude, gpt, benchmarksData and scoring methodology
In this article4 sections

GPT-5.4 is the stronger broad and lower-cost default. Claude Opus 4.6 has the higher coding estimate. The models are close on agentic work, where both rows are Supported.

Headline comparison

Table 1
Metric Claude Opus 4.6 GPT-5.4
Overall 68.52 Supported 74.18 Supported
Overall rank #15 #9
Coding 65.66 Estimated 64.17 Supported
Coding rank #11 #15
Agentic 58.87 Supported 59.82 Supported
Agentic rank #20 #18
API price $5/$25 $2.50/$15

Where Claude wins

Claude leads raw HLE, 53 to 48, and SWE-bench Pro, 74 to 57.7. Its combined coding estimate is also higher. Those are real reasons to test it on complex repository work.

The caveat is evidence strength. Claude's coding row is Estimated, while GPT-5.4's is Supported. A 1.49-point lead is not strong enough to skip a direct trial.

Where GPT-5.4 wins

GPT-5.4 leads overall, agentic work, and price. It also leads raw SWE-bench Verified, LiveCodeBench, Terminal-Bench 2.0, OSWorld-Verified, SimpleQA, MMLU-Pro, LongBench v2, and MRCRv2 in this matchup.

The broad pattern favors GPT-5.4. The price reinforces it: for one million input and 200,000 output tokens, GPT-5.4 costs $5.50 and Opus 4.6 costs $10.

Which should you use?

Use GPT-5.4 when you want the safer broad default between these two exact versions. Use Claude Opus 4.6 when coding dominates and it wins on your repository test.

If current frontier models are allowed, add Claude Opus 4.8, GPT-5.6 Sol, and Claude Fable 5. Opus 4.6 versus GPT-5.4 is a version-specific comparison now, not a proxy for the best model each lab sells.

Full version comparison · Side-by-side data · Coding ranking

Reader questions

Frequently asked questions

01Is Claude Opus 4.6 better than GPT-5.4?

GPT-5.4 is higher overall at 74.18 versus 68.52 for Claude Opus 4.6, both Supported. Claude ranks higher in coding at 65.66 versus 64.17, but Claude's coding row is Estimated. GPT-5.4 has the slightly stronger Supported agentic score.

02What is the price difference?

GPT-5.4 costs $2.50 input and $15 output per million tokens. Claude Opus 4.6 costs $5/$25, which is 2x the input price and about 1.7x the output price.

03Which is better for coding?

Claude Opus 4.6 ranks eleventh in coding at 65.66 Estimated; GPT-5.4 ranks fifteenth at 64.17 Supported. The modeled gap is small enough that a repository evaluation should decide.

Share or save

Share on XShare on LinkedIn

Keep reading

All research

These rankings update with every new model. Join 2,000+ readers for one email a week on what moved, why, and what still needs proof.