Insights

The Future Is Now. Where Is the Cost?

AI is transforming MSP operations — but the human cost of adoption rarely makes it into the business case. Before raising targets or redesigning roles, the impact on your engineers deserves a line in the same plan as the gains.

Andrew Yager · Director and CTO AI Business Work Culture

AI is transforming IT operations for service providers. That much is real. We sell it, we run it in production rather than in pilots, and at Crayon Connect earlier this year I told a room that organisations which don't change will die, and that an MSP expecting to still run L1 helpdesk in 2030 would find life hard. I stand by all of it.

Several talks at ARN Edge this week have pitched AI Ops and AI-assisted support as the next big transformation for MSPs. Not as an experiment, but as the thing that changes how our support teams work and wins out in the end. I do think they're right.

What I haven't heard anyone account for is what the change does to the engineer.

An engineer who adopts AI properly stops being the person who produces the work and becomes the person who directs it and then checks it. The undemanding stretches leave the day first, because they're the easiest thing to hand over. What's left is denser, more ambiguous, and needs a judgement call more or less continuously. That's a different job to the one before it, not the same job at a faster pace, and it asks more per hour while giving back fewer places to rest. We have watched that land on people repeatedly, and it follows a shape we now expect: about a fortnight of remarkable acceleration, then output falling through the floor from around day 30, then a recovery at roughly day 91 to something like a fifth better than where they started. Those numbers matter, but only as the visible trace. What they're tracing is a person carrying more than they were.

That has consequences for the organisation, and they aren't the ones on the slide. Capacity you measured in week two isn't capacity you can staff against. Because the curve attaches to the person rather than the company, rolling out gradually doesn't avoid the trough, it distributes it, so a team that keeps onboarding people has somebody in it more or less permanently. On a team dashboard that averages into a mild unremarkable dip that never quite resolves, and nobody goes looking for a cause they can't see.

Then there's the part that should be easiest for our industry to grasp, because we do it for a living. None of us would run a platform migration without budgeting for change management. We all know the licence is the cheap part and the change is where the money goes. Two months of degraded output per adopting engineer, plus the management attention required to carry someone through it, is a change management cost. It's forecastable, it's repeatable, and I have never seen it in a business case for AI adoption. Including mine.

And then the part nobody can answer. We can see day 91. We can't see day 910. Whether that 120% is a floor or a point on a slope is the single most consequential thing about this and there is no data on it, because none of us have been doing this long enough.

This isn't an argument to slow down. It's an argument to stop treating adoption as free acceleration. Before anyone raises targets, reduces headcount, or redesigns engineering roles around this, the cost to the people doing the work needs to sit in the same business case as the gain.

What Has Actually Worked

Our own evidence for the transformation is substantial.

AI has changed how our engineers solve problems, and the effect is strongest at the senior end, which I didn't expect. The most valuable shift is watching engineers move into adjacent territory they knew but were never deep enough to drive alone. That gap used to be a hard boundary. It isn't now. We also solve problems faster, analyse data quicker, and iterate faster than we used to.

More recently we've put Claude in front of more of our engineers and integrated Claude Tag into many of our internal Slack channels. That second part matters more than it sounds. Claude is available where the work already happens, in the threads where the team is already talking, rather than in a separate tool someone has to decide to open.

Two things came out of that. The first is retrieval. Staff started finding documentation that always existed but they couldn't locate, or that was framed in a way that didn't match how they'd think to ask. The knowledge was already in the building. What changed was being able to ask in whatever words occurred to you.

The second is that it's shared. Because it happens in a channel rather than a private window, the question and the answer are visible to the thread. Junior engineers see how senior ones frame a problem, which is something we used to get from sitting near each other and have been quietly missing since.

I'll flag one thing about it and come back to it. No friction to asking also means no friction to starting. There's no longer a moment where you decide to begin.

What the Roadmap Sells, and What It Leaves Out

ConnectWise keynoted with a five-phase roadmap to 2030. Operationalise AI, Extend the Workforce, Supervise AI, Autonomous Ops, Predictive Intelligence, with the first three marked as delivering real results today. Phase 5 in 2030 is roughly the future I described at Crayon Connect, and we're already running things the roadmap puts a year or two out. The destination isn't where I'd pick a fight.

But a roadmap isn't only a sequence of technical capabilities. It's a sequence of changes to human work, and that half isn't costed.

Phase 3 is called Supervise AI, dated 2026 to 2027. Which is now. It may well be an accurate description of the next operating model. It's also the point where an engineer's role shifts from producing work to continuously evaluating machine output, and the evidence we have says that shift is cognitively expensive and unevenly distributed. It appears on the roadmap as a milestone rather than as a cost.

Credit where it's due, because plenty of vendors talk as though the human evaporates and this roadmap doesn't. Naming a supervision phase is more honest than pretending there isn't one. But two years of engineers supervising machine output is a cost centre as well as a rung on a ladder, and I haven't heard that half meaningfully costed by anyone this week. I didn't cost it either, when I had the microphone.

Phase 4 may not resolve it. Autonomous Ops reads like the point where supervision load falls away, and the human factors literature suggests the opposite. The more autonomous the system, the rarer and stranger the events that reach a human, and the less practised that human is when one does.

What We Observed

The caveats first. No control group, no clean instrumentation, and figures I'd call educated estimates rather than measurements. They're built from tickets solved, projects delivered and similar operational measures.

What makes me take it seriously anyway is that it repeats. We're not extrapolating from a single rollout. We've watched the same shape appear on different people at different times, often enough that we now expect it.

The first fortnight is terrific. By around day 15 an adopting engineer is clearing work at roughly three times their old rate. I'd defend the rough scale of that sooner than the exact figure. It holds up through about day 30.

Then from roughly day 31 to day 90 it declines, and not by a little. Output goes below where that person started. The word people used at the time was burnout, and I don't think it was hyperbole.

Around day 91 it comes back, to about 120% of where they began.

I'm publishing that rather than tidying it away, because the response is the part that matters. Once we recognised the shape we stopped treating it as an individual problem. We now baseline with a validated measure rather than reading throughput, we stage adoption so we aren't carrying several people through the trough at once, and we've moved verification-heavy work back to conventional automation where it belonged. The rest of this piece is largely what we learned doing that.

Schematic chart of output per adopting engineer after AI adoption, indexed to their day zero baseline, from an internal estimate rather than measurement. Output climbs to roughly 300 percent by day 15, holds to around day 30, falls below baseline through days 31 to 90, then recovers to roughly 120 percent at day 91. Dashed projections show the open question of whether 120 percent holds or keeps sliding.

Read the whole curve rather than any point on it. The 300% shows what the technology can unlock. The trough shows what the transition costs. The 120% is the sustainable state, and a fifth more output per person is a good result. The unknown after that is why measurement has to continue.

What the curve mainly demonstrates is that the gain and the cost don't arrive at the same time. At day 15 the organisation sees acceleration. The adaptation, the oversight load and the lost recovery land on the individual later. By day 91 the result is still positive, but it's a different result from the one implied at launch.

Which makes when you measure more consequential than what you measure. Fourteen days in, you'd call it a triumph and start reallocating people. Forty days in, you'd kill the project. Neither is right, and nearly every pilot review and vendor case study I've seen lands in one of those windows. The costlier mistake is the headcount you reduce against a day-15 figure, which leaves you staffed for capacity nobody sustains. On this curve you'd find that out around day 50, with two months of trough still to run.

And this applies to my own numbers as much as anyone's. Every figure I've given you is throughput: tickets closed and projects shipped. Those are the numbers an MSP has to hand and the ones I reached for without thinking. They're also the wrong instrument. Throughput tells you what someone produced, not what producing it cost, which means it can't tell you whether that 120% is a floor or a point on a slope.

What the Research Explains

Nobody has published our curve, and I'm not going to claim it as a general law. What the literature does provide is several mechanisms consistent with what we keep seeing, which is enough to justify measuring rather than assuming.

Oversight is the tiring part, not usage. The BCG Henderson Institute and UC Riverside surveyed 1,488 US workers, published in Harvard Business Review in March. 14% reported mental fatigue from using, interacting with or overseeing AI beyond their cognitive capacity. Marketing was worst affected at around 26%, with software development, HR, finance and IT also elevated. Those affected reported 33% more decision fatigue, 39% more major errors, and intent to quit rising from 25% to 34%. Productivity peaked at around three AI tools running at once and fell away at four. The people most at risk weren't the reluctant adopters. As one of the BCG authors put it, they'd noticed this happening to people regarded as high performers. This is the Phase 3 finding, and it's why I'd want that phase examined before it's designed into roles.

One finding complicates my argument. In the same study, AI users reported less burnout than non-users, and replacing repetitive tasks with AI reduced burnout scores by around 15%. It was oversight load that did the damage, not the removal of routine work. So the simple version of my thesis, that automating routine work is what hurts, doesn't survive. The narrower version does: automating routine work is fine and often good, until the work that replaces it is continuous evaluation of machine output. The researchers' explanation matters more than the result. Burnout instruments measure emotional and physical exhaustion, and acute cognitive fatigue is a different thing the standard scales don't catch. A team can be draining down while every wellbeing dashboard reads normal.

Work intensifies even when nobody demands it. Aruna Ranganathan and Xingqi Maggie Ye of UC Berkeley's Haas School spent eight months embedded in a 200-person tech company, published in HBR in February. AI didn't reduce work, it intensified it, through task expansion, blurred boundaries and multitasking. One worker put it plainly: you don't work less, you just work the same amount or more. Because AI made starting frictionless, people began slipping work into moments that had been breaks. Lunch, meetings, the minute spent waiting for a file to open. Several worked out in hindsight that downtime had stopped being restorative. These were volunteers.

That last part is the Slack integration I mentioned, described back to me by researchers who've never seen our environment. Putting Claude in the channel where the team already talks is the best thing we've done this year, and it removes the last friction between a spare moment and starting work. Both true at once, which is roughly the whole problem.

Where the recovery goes. Occupational health psychology has a model called effort–recovery, set out by Meijman and Mulder in 1998. Mental effort produces measurable load reactions in the body, from raised cortisol through to fatigue and flattened mood. Those are harmless and fully reversible on one condition, which is that the systems involved return to baseline during periods of low demand. Much of that happens in the troughs of an ordinary day. Waiting for a build. Boilerplate you could write half-asleep. Reading code you already understood to find one line.

That's the work we automated first, because it was easiest to automate and most obviously wasteful. It was also where the troughs were. Lisanne Bainbridge described the consequence in 1983 in "Ironies of Automation": automation takes the easy parts and hands the operator the residue, the difficult and ambiguous fragments. "By taking away the easy parts of the task," she wrote, "automation can make the difficult parts of the human operator's task more difficult." The paper is about power stations and aircraft, and it describes a Phase 3 service desk reasonably well.

Diagram showing how automating low-demand work draws down the recovery reserve. Automating boilerplate, waiting and routine reading produces two effects: the leftover principle, where remaining work is uniformly high-demand, and the disappearance of low-demand stretches within the day. Both lead to continuous evaluative attention with no unwinding. Reduced headcount and the law of stretched systems amplify it. The result is that the recovery reserve draws down.

A 2022 meta-analysis in PLOS ONE across 22 samples found micro-breaks improve vigour and reduce fatigue, but mainly for less demanding tasks. Recovering from genuinely hard work takes longer breaks. If the day is now uniformly hard, the short natural pauses that used to do the job won't. I'll be honest that no published study isolates "AI removed the low-demand work that was doing my recovery" and measures what follows. The mechanism is coherent. The direct test hasn't been run.

And the aggregate numbers are ambiguous. Google's DORA research found AI adoption associated with better throughput and worse delivery stability. METR's randomised trial had experienced developers estimate AI made them 20% faster on real tasks in their own repositories; measured, they were 19% slower. Faros AI's telemetry across 22,000 developers found work restarts and long-stalled tasks both up sharply. ActivTrak's analysis for the Wall Street Journal, across 164,000 workers, found focused uninterrupted work fell around 9% after adoption. And Stack Overflow, asking the same questions annually, found favourability falling from 72% to 60% and trust in accuracy from 40% to 29% while adoption rose to 84%.

None of that proves our curve generalises beyond us. It's enough to say the cost we keep seeing is plausible, consequential, and consistent with a reasonable body of research.

What Better Design Looks Like

The useful conclusion is that much of this is a design problem rather than something inherent to AI. Some of the cost is a consequence of how you divide the work rather than of AI itself.

What we're aiming at is closer to co-working with AI than supervising it.

Send deterministic work to deterministic automation. If a task runs the same way every time, it belongs in a script or a runbook, or in the automation engine you're already paying for. Output you can trust without reading costs nothing to supervise; output you have to verify costs you every time. Some of what gets pitched as AI Ops is work conventional automation already did reliably, moved to a tool whose output now needs checking. That transfers effort into verification rather than removing it.

Keep the engineer as the author. The moment AI produces the artefact and a human inspects it, you've built the Phase 3 job and its oversight load. Where the engineer is still making the thing, AI contributes to a decision rather than handing over something to be audited.

Point the intelligence at the questions rather than the output. Asking what we haven't considered. Surfacing context that isn't obvious from the ticket. Noticing that a customer's last three incidents share a cause nobody connected. The reason that division of labour matters is the asymmetry in error cost: a bad question costs nothing because you read it and discard it, while bad output you didn't catch costs you an incident. Insight-shaped use carries a fraction of the verification burden of production-shaped use.

The drawback is that this model demos badly. It won't produce a 300% number in week two, because you're improving the decisions inside the work rather than replacing the work. Which may be why it isn't being sold from a stage.

The Regulator Is Heading Here Anyway

Psychosocial hazards are already a regulated duty under Australian WHS law. Safe Work Australia's model Code names job demands and cognitive overload, and it names fatigue, which SafeWork guidance defines as not enough recovery driven by hours, shifts, workload, or how the work is structured. That last clause is worth reading twice. A workflow that removes within-day recovery is a structural decision about how work is organised.

NSW has gone further. The Work Health and Safety Amendment (Digital Work Systems) Act 2026 passed on 12 February 2026 and received assent on 18 February, adding a new section 19(3)(c1) to the primary duty, covering any use of a digital work system, plus a new section 21A covering the allocation of work by one. "Digital work system" is defined broadly as an algorithm, artificial intelligence, automation or online platform. Under 21A a PCBU must consider whether allocating work that way creates unreasonable workloads, unreasonable performance metrics, or excessive monitoring.

It hasn't commenced yet, and the timing is worth understanding rather than assuming. Only the provisions requiring SafeWork NSW to develop guidelines on the new entry-permit power took effect on assent. Everything else commences on proclamation, on a date not yet announced.

Here's the part I'd initially got wrong. Most provisions can't commence earlier than a month after those guidelines are published, but the commencement clause carves out the schedule items containing the duties themselves. Herbert Smith Freehills reads that as the new duties likely taking effect reasonably soon, with the delay applying to the right-of-entry powers rather than to sections 19(3)(c1) and 21A. So don't plan on the guidelines consultation buying you time on the duty.

If you've lifted delivery targets assuming AI covers the difference, or you're tracking AI usage as a performance metric, 21A is worth reading now rather than after it commences. We're not lawyers and this isn't advice on your circumstances.

What To Actually Do

We sell AI acceleration work, so weigh this accordingly.

Measure the right construct, before there's a problem you can't see. Given what the “brain fry” study found about burnout scales, throughput and burnout surveys won't tell you what you need. Maureen Dollard's PSC-4 out of the University of South Australia is free, four questions, and Australian. Add the Recovery Experience Questionnaire to catch detachment. Baseline now, repeat at three months and six. A rising need-for-recovery score moves well before anyone presents with burnout.

Treat recovery as a design problem rather than a personal responsibility. Telling people to take breaks doesn't work when the work has no natural stopping points. Put review into blocks with real gaps after them. Watch the tool count, since three concurrent tools appears to be where returns stop. And audit where your verification burden comes from, because any repeatable task where a model produces output someone reads is a candidate to move back to conventional automation.

Budget the change, not just the licence. If the trough is roughly two months per adopting engineer, that's forecastable and it belongs in the plan the same way you'd cost the change management on a platform migration. Which means staging adoption so you're not carrying too many people through it at once, giving whoever is in the trough somewhere to go for help, and not committing the day-15 capacity to a customer while it's still notional.

Then decide deliberately where the gain goes. Every improvement to a system tends to get consumed by higher expectations rather than kept as slack. That's the law of stretched systems, an observation from safety research credited to Lawrence Hirschhorn and popularised by David Woods and Richard Cook, and it has held up well. If your engineers are faster and also busier, the gain went to output. That's a legitimate choice. Just make it on purpose, and put it in the ledger alongside everything else.

Because that's the ask. Not that anyone slows down. If our people are integral to the results, we count what the results cost them in the same document where we count the gain.

Move, But Know What You're Moving Into

The vendors are right that the future is arriving. The question is whether we're treating it as a product to buy or an operating model to understand.

If AI changes what people do, how continuously they concentrate, when they recover and how much uncertainty they carry, those effects belong in the business case before the gain is banked. So think twice before you bank a pilot number and three times before you take headcount out against one. Ask anyone quoting you a productivity figure how long after go-live it was measured, and what happened to the people in the eight weeks after that.

We can see day 91. Nobody I've spoken to this week can tell me what day 910 looks like, and that's the number that decides whether any of this was worth what it cost. If you've run it longer than we have, I'd like to hear from you, including if your numbers say I've got this wrong.

Want to work through where your team sits? Call us on 1300 798 718.


Sources

Enjoyed this? Subscribe.

New posts on cybersecurity, cloud and the real-world problems we solve — straight to your inbox.

Email me about

We’ll email you new posts and you can unsubscribe anytime. See our privacy policy.

Want to talk it through?

If this raised questions about your own setup, call us — no pressure, just a conversation.

1300 798 718