AI Was Supposed to Save Us Time. Why Are We Working Longer?

LLMs can save hours on the right task. They can also consume that time through prompting, checking and rework. Here is how leaders can tell the difference.

TL;DR

AI productivity depends less on access to an LLM than on task selection, skill and measurement. Research has found gains of 14% in customer support and speed gains above 25% on suitable consulting tasks. Another randomized study found experienced developers took 19% longer with AI on familiar codebases. The business lesson is simple: count the full cycle from task start to approved output. If prompting, checking, correcting and reformatting consume the saved time, you have added software activity without improving productivity.

The uncomfortable question about AI productivity

I use AI every day. I build AI systems, advise companies on adoption and spend a lot of time testing what these models can actually do. That makes me more interested in an uncomfortable question, not less interested.

Are we spending so much time talking to large language models that we could have finished the work ourselves in the same time, or faster?

For a growing number of teams, the answer is sometimes yes.

The problem is rarely the model alone. It is the workflow built around it. Someone writes a long prompt, waits for an answer, explains the missed context, requests another version, checks every claim, removes awkward wording, fixes the format and then completes the original work.

The finished document may look fast because the first draft appeared in seconds. The job was not finished in seconds.

AI can save time, but the gains are uneven

The research does not support a simple claim that LLMs make everyone faster. It shows something far more useful. Results change by task, worker experience, context and the cost of checking the output.

A study of 5,179 customer-support agents found that access to a generative AI assistant raised issues resolved per hour by 14% on average. Novice and lower-skilled workers improved by 34%, while gains for experienced and highly skilled workers were small.

A Harvard Business School and Boston Consulting Group experiment involving 758 consultants found that GPT-4 helped people complete suitable tasks over 25% faster, raised human-rated performance by over 40% and increased task completion by over 12%.

But the researchers also described a jagged boundary. AI performed well on some tasks and poorly on others that appeared equally difficult. When people used AI beyond that boundary, performance suffered.

Then came one of the clearest warnings about perceived productivity. METR ran a randomized trial with 16 experienced open-source developers completing 246 real tasks in repositories they knew well. The developers expected AI to make them 24% faster. AI use made them 19% slower. Even after the work, they still believed it had saved them 20%.

That gap matters. AI can feel faster while making the full task slower.

The hidden LLM time tax

Most companies measure the visible benefit and ignore the hidden cost.

They see a draft arrive quickly. They do not count the time spent getting the model ready, supplying context, reviewing the response, correcting errors, reconciling sources and making the result fit the company’s standards.

The full cost often includes:

Prompt preparation and repeated clarification

Uploading or locating the right source material

Waiting for long responses or agent runs

Fact-checking claims and citations

Correcting generic, inaccurate or invented content

Reformatting the output for the actual system of record

Review by a subject-matter expert

Rework caused by missing company context

None of these costs means the tool has failed. They mean the company must measure the whole job.

Why prompting can become work theatre

There is a difference between productive iteration and endless conversation with a model.

Productive iteration improves a high-value output faster than the existing process. Work theatre creates a lot of visible AI activity without changing throughput, quality, revenue, cost or risk.

You can spot it when employees keep polishing prompts for simple emails, ask an LLM to summarize material they still need to read in full, or generate a report that takes longer to verify than it would have taken an informed employee to write.

This can also happen at the management level. A company buys licences, tracks usage and celebrates adoption. Login counts rise. Prompt counts rise. Nobody measures completed work.

Usage is not productivity.

The expertise problem

AI often creates the biggest time gain for people who need help reaching a competent first result. That is consistent with the customer-support research, where newer and lower-skilled workers gained the most.

Experts face a different calculation. If you already know the subject, understand the source material and can produce the answer quickly, explaining the task to a model may add overhead. You may also detect subtle errors that a less experienced person would miss, which adds review time.

The model can still help an expert. It may challenge an assumption, produce alternatives, reorganize notes or handle a repetitive step. But asking it to perform the entire job can be slower than assigning it a narrow part.

Microsoft researchers surveyed 319 knowledge workers about 936 examples of generative AI use. Higher confidence in AI was associated with less critical thinking. The work shifted toward verification, integration and stewardship. In plain language, the mental effort did not disappear. It moved.

How Canadian companies should measure AI ROI

Canadian businesses do not need another broad AI adoption target. They need task-level evidence.

Choose one repeated workflow and compare the old process with the AI-assisted process. Use the same quality bar for each. Measure at least 20 comparable tasks if the workflow occurs often enough.

Measure these five numbers

Total cycle time from assignment to approved result

Active human time, including prompting, review and corrections

First-pass acceptance rate

Error and rework rate after approval

Cost per approved output, including software and staff time

Then ask one practical question: did the workflow create a better approved result with less human time, lower cost or lower risk?

If the answer is no, change the task, the tool or the process. Do not blame the employee and do not keep scaling a weak workflow because the licence has already been purchased.

A simple rule for deciding when to use an LLM

Use an LLM when the expected value of the draft, analysis or automation exceeds the cost of instruction and verification.

That usually favours tasks with repeatable inputs, clear output criteria, low correction costs and enough volume to justify setup. It often works well for first drafts, classification, extraction, structured summaries, option generation and repetitive internal support.

Be careful with tasks that depend on undocumented context, carry high legal or financial stakes, require exact current facts, or can already be completed quickly by an experienced employee.

Sometimes the fastest prompt is no prompt.

What leaders should do next

Stop using licence adoption as the main success metric.

Pick three repeated workflows with measurable volume and business value.

Record the full before-and-after cycle time, including review and rework.

Set a quality threshold before the test begins.

Keep the workflows that produce a verified gain. Redesign or remove the rest.

Train employees to recognize tasks where direct work beats prompting.

The goal is not to put AI into every task. The goal is to remove time, cost and friction from work that matters.

So here is my take,

I still believe LLMs are among the most useful business tools introduced in years. I also think we are entering the stage where serious companies must move past demonstrations and measure what happens after the first impressive answer.

A model that writes a draft in 20 seconds has not saved 30 minutes if an employee spends 35 minutes repairing it.

AI productivity becomes real when the approved work moves faster, costs less or improves in a way the business can verify. Everything else is activity.

Frequently asked questions

Do LLMs actually save time at work?

Yes, on well-matched tasks. Studies have found meaningful gains in customer support and consulting. The gains vary by task and worker experience, and they can disappear after review and correction time is counted.

Can AI make employees less productive?

Yes. A randomized METR study found experienced developers took 19% longer with AI on familiar open-source codebases. The finding applies to that setting and should not be generalized to every worker or task.

How should a company measure AI productivity?

Measure total cycle time, active human time, first-pass acceptance, errors, rework and cost per approved output. Compare the AI-assisted workflow with the prior process at the same quality level.

Which business tasks are best for LLMs?

LLMs tend to help when inputs repeat, output criteria are clear, correction costs are low and the task occurs often. Drafting, extraction, classification, structured summaries and option generation are good places to test.

When should an employee skip AI?

Skip it when direct work is already faster, the task relies on deep undocumented context, verification would cost too much, or an error could create serious legal, financial or reputational harm.

Sources

METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

METR, We Are Changing Our Developer Productivity Experiment Design

NBER, Generative AI at Work

Harvard Business School AI Institute, Navigating the Jagged Technological Frontier

Microsoft Research, The Impact of Generative AI on Critical Thinking

Stack Overflow, 2025 Developer Survey on AI

Total
0
Shares
Prev
OpenAI’s Astra Warning and AI Agent Security in Canada

OpenAI’s Astra Warning and AI Agent Security in Canada

AI agent security in Canada should move from acceptable-use documents to

Next
Meta Muse Glimmer and Local AI for Canadian Business

Meta Muse Glimmer and Local AI for Canadian Business

Muse Glimmer puts a 30-billion-parameter agent model inside a 24 GB or 32 GB

You May Also Like