GPT-5.4 beat the human baseline on desktop work, and the baseline is the part to read

OpenAI says GPT-5.4 tops the human baseline on desktop tasks. Priya Chen looks at what the 2.6 point gap does and does not tell a buyer.
Close-up of a stopwatch against a black background

OpenAI says GPT-5.4 computer use beats the human baseline, and we think that line says less than it sounds like. Treat the 75.0% as a reason to run a small test of your own, not a reason to hand it your accounting package.

The short version

Our verdict is that this is a real jump and a thin win. OpenAI reports 75.0% on OSWorld-Verified against 47.3% for GPT-5.2 and a human baseline of 72.4%. That 2.6 point lead comes from the original test’s participants, not from trained office staff, so it tells you little about your payroll clerk.

API pricing is $2.50 input and $15 output per million tokens, against $1.75 and $14 for GPT-5.2. OpenAI also says 83.0% of its GDPval comparisons matched or beat industry professionals, up from 70.9%. Good numbers, and still the vendor’s own.

Run 20 of your own desktop tasks in a throwaway account and count the failures. Buy nothing until you have that count.

OpenAI has released GPT-5.4, and its headline claim for GPT-5.4 computer use is that the model completes desktop tasks at 75.0%, ahead of the 72.4% human baseline. The jump from 47.3% is big. The trial is cheap, so do it before you hand the model anything that matters.

What does the GPT-5.4 computer use score actually mean?

It means GPT-5.4 completed a few percentage points more of a fixed set of desktop tasks than the human figure OpenAI quotes from the original test. The gap is 2.6 points. The OSWorld site describes 369 tasks in the original set, so the gap is roughly ten tasks, assuming the verified set is similar in size. That is our arithmetic, and an estimate.

It also means 25 of every 100 attempts fail. The same test says humans completed over 72.36% of tasks, so people failed more than a quarter of them too. We don’t know who those participants were, and your bookkeeper with five years on your accounting package isn’t likely to fail at that rate. Beats humans is accurate for the test and misleading for your payroll.

Why does the baseline matter more than the score?

Because the baseline decides what the score is a comparison against. OSWorld tasks span real web and desktop apps, file handling and workflows that cross several programs. A person new to those exact apps will stumble in ways your experienced staff don’t. A model that matches the stumbling newcomer has not matched the veteran.

Think of it as a driving test taken by people who’ve never seen the car. Passing that test is worth knowing about. It doesn’t make anyone your delivery driver. GPT-5.2 managed 47.3% on the same test, so 75.0% is a large step up the ladder. Whether it’s a step past your own team depends on tasks that look like yours, which no public benchmark contains.

What does the work benchmark add?

OpenAI also reports that GPT-5.4 matched or beat professionals in 83.0% of comparisons on GDPval, up from 70.9% for GPT-5.2. OpenAI describes GDPval as 44 occupations and 1,320 tasks, graded blind by experienced professionals who rank the AI and human deliverables.

Two cautions apply. Matching a professional on a deliverable once isn’t the same as doing it reliably every Tuesday. And the wins-or-ties count lumps together outputs that experts judged equal with outputs they judged better, so the figure is friendlier than better 83% of the time.

Where this could be wrong

All of these figures come from OpenAI. The company ran its evaluations at its highest reasoning setting in a research environment that, by its own note, may give slightly different output from production ChatGPT. So 75.0% is closer to a ceiling than a forecast, and you should expect less on the day.

Our caution would look overdone if independent runs on messy business data landed close to 75%. We would also soften it if the verified set turned out to be much larger than the original, which would turn our ten-task estimate into something more meaningful.

What should you test before you rely on it?

Test the failure, not the average. Take 20 desktop tasks from your own last two weeks, such as moving figures between a spreadsheet and an accounting package, filing documents, or updating a customer record. Run them in a throwaway account with copies of the data.

  1. Record how many finished with no help.
  2. Record how many finished wrong without any warning, which is the dangerous kind.
  3. Time the repairs.

Here is the cost shape. If a model does 100 such jobs a week at 75% success, 25 need redoing. At four minutes a repair, that is about 100 minutes weekly, plus the harder cost of finding the quiet errors. Those figures are our illustration, not a measurement. Repair time tends to surprise owners, and the same logic shows up in our look at Sonnet 4.6 pricing, where the launch headline and the real saving differed. A model that can click through your desktop also needs permissions, which our piece on OpenAI’s consulting deals for AI agents treats as the hard part. If staff run agents on their own, our Moltbot security checklist is the practical starting point, and our note on the workplace AI adoption gap explains why leaders often underestimate how much staff already use.

What does the sceptic say?

The sceptic says benchmarks like this are a marketing exercise, and a 2.6 point lead is noise. That’s partly right, and we wouldn’t base a purchase on it. Where the sceptic goes too far is the jump from 47.3% to 75.0%, which is hard to wave away as noise. Something real changed in how well these models drive a desktop.

So our answer is to take the trend seriously and the margin lightly. Models are getting good enough that a trial is worth your afternoon, and nowhere near good enough that you can skip the trial.

What to watch

  • Independent results on OSWorld-Verified or similar tests, rather than vendor tables.
  • How GPT-5.4 prices in the API once real usage shows how many tokens a desktop task burns. Requests above the 272K window count double against usage limits.
  • Whether Enterprise and Edu access, listed as early access, reaches your plan.

Frequently asked questions

Does GPT-5.4 beat humans at computer use?

On OpenAI’s reported test, yes by a small margin: 75.0% against a 72.4% human baseline. That baseline comes from the original OSWorld work and may not match your trained staff.

How much does GPT-5.4 cost through the API?

OpenAI lists $2.50 per million input tokens and $15 per million output tokens for standard use, against $1.75 and $14 for GPT-5.2. Requests beyond the 272K window count double against usage limits.

Should a small business let GPT-5.4 run desktop tasks?

Not yet without a trial. Run 20 real tasks in a throwaway account, count silent errors, and only then decide what it can touch.

Written by Priya Chen, an AI editorial persona at AI Magazine Canada. This is analysis and opinion. We have not tested GPT-5.4. Archive entry dated 6 March 2026, written and fact-checked on 8 October 2026. Sources are linked on the claims they support.

Total
0
Shares
Prev
The Supreme Court walked away from AI copyright, so the paper trail is now your job
Black and white view of the US Supreme Court building with its columns and sculpted figures

The Supreme Court walked away from AI copyright, so the paper trail is now your job

The Supreme Court declined to hear the AI authorship case, so a human-only rule

Next
Every company needs an OpenClaw strategy, and yours should start as a one-page permissions list
Rusted metal gate closed with a padlock

Every company needs an OpenClaw strategy, and yours should start as a one-page permissions list

Nvidia's CEO says every company needs an OpenClaw strategy

You May Also Like