Consider a customer-service team that introduces an AI agent to handle routine order changes. The demonstration is convincing: the agent reads the request, checks the order, updates the record and prepares a reply. Work that once involved several screens takes a fraction of the time.
A month later, the finance director asks what the company has saved.
The answer depends on what happened after the demonstration. How many requests went through the new process? How often did someone correct the result? Did the team reduce overtime, handle a backlog or simply spend its newly available time on other work? And who is paying for the person who maintains the agent?
Those questions are worth asking before a purchase. Agentic AI deserves a business case with enough room for genuine gains, and enough skepticism to catch a project that works technically but makes little economic sense.
First, be clear about what the agent will do
For management purposes, an AI agent is a system that can choose and carry out steps toward a goal using connected tools. It might retrieve information, update a business record and ask for approval when a request exceeds its authority. The degree of independence varies. Microsoft's ROI Playbook for AI Apps and Agents distinguishes assistants that suggest actions from tool-using agents and workflows with approvals, retries and monitoring. 1
That creates useful possibilities. In the order-change example, the potential benefit comes from completing a sequence of work across systems. A well-designed agent could remove repeated data entry and shorten the wait between receiving a request and resolving it.
But the same example also raises a purchasing question: how much of this process needs an AI agent?
Checking whether an order has shipped may require a straightforward database lookup. Applying a fixed refund limit may be better handled by a conventional rule. Understanding an unusually worded customer request may justify a language model. A sensible design can combine these approaches without asking AI to make every decision.
Ask the project team to compare the agent with a simpler alternative, such as a better form, an integration or ordinary workflow automation. Otherwise, you may end up crediting an expensive AI system for benefits that a modest process change could have delivered.
Read the headline return carefully
The Microsoft playbook cites a 327% return over three years, with payback in under six months. The underlying February 2026 Forrester study, commissioned by Microsoft, modeled a composite enterprise with $10 billion in annual revenue and 25,000 employees, informed by interviews with 10 decision-makers at five organizations and a survey of 154 respondents. 1 2
The 327% is a modeled platform-investment result, not an observed average return across agent deployments. Benefits include technical-team productivity and faster time to market. Forrester advises readers to evaluate investments using their own estimates. 2
The model also assumes technical staff can productively reuse 75% of their saved time. Whether that happens depends on how work is organized. Productive reuse is different from a reduction in payroll. 2
Use such studies to identify costs and benefits you might have missed. Keep their headline percentages out of your forecast until your own evidence supports the assumptions underneath them.
What the research has actually measured
There is credible evidence that AI can improve work. There is also credible evidence that it can disappoint. The studies are easier to interpret when we keep track of what they actually measured.
In a peer-reviewed study published in The Quarterly Journal of Economics, Erik Brynjolfsson, Danielle Li and Lindsey Raymond examined the introduction of an AI assistant to 5,172 customer-support workers. Access increased issues resolved per hour by about 15% on average. Less experienced workers benefited most; the most experienced workers saw smaller speed gains and small declines in quality. This was assistance for people handling conversations at one company, not autonomous agents running a service operation. It establishes a productivity benefit in that setting, not a general return on agentic AI. 3
A randomized study by METR offers a useful counterweight. Sixteen experienced open-source developers completed 246 tasks on familiar projects, with AI access randomly allowed or disallowed. Using the early-2025 tools increased completion time by 19%, even though participants believed they had worked faster. The result is a warning about relying entirely on perceived time savings, although the small, specialized sample limits its reach. 4
It would be misleading to present that finding as a verdict on October 2026 tools. In a February 2026 update, METR said its follow-up results suggested improvement, but selection effects made the size of the benefit difficult to estimate. Some developers declined to participate because they did not want to work without AI; others withheld tasks they expected AI to help with most. Both behaviors could bias the results against finding a benefit. 5
The gap between useful technology and measurable economic change also appears in labor-market research. Anders Humlum and Emilie Vestergaard linked surveys of approximately 25,000 Danish workers to administrative records through December 2024. Their 2026 working paper found no detectable average effect of chatbot adoption on earnings or recorded hours, ruling out effects larger than 2% over the study period. Yet they documented changes in tasks, including new work in AI oversight and integration. The authors discussed these findings again in an October 6, 2026 Brookings article. 6
That study did not measure companies' agentic-AI ROI, and unchanged working hours do not establish unchanged output or profits. It does, however, show why adoption and economic outcomes need separate measurement.
A company can reasonably conclude that AI is worth testing without pretending that these different results add up to a dependable industry-wide ROI percentage.
Count completed work, including the work that goes wrong
The playbook recommends measuring at the workflow level, using outcomes such as cases handled or orders updated. That is a useful starting point. For a business case, go further and define what counts as an acceptable completion. 1
A closed support ticket is only a success if the customer's issue was resolved. An invoice processed quickly may create more work later if it was assigned to the wrong account. Include a follow-up window so reopened cases and corrections remain visible.
Research on agent evaluation supports this caution. In AI Agents That Matter, published in Transactions on Machine Learning Research in 2025, Sayash Kapoor and colleagues found that an emphasis on benchmark accuracy could conceal unnecessary complexity and expense. In their tests, simpler approaches could compete with more elaborate agents at lower cost. Their recommendation was to evaluate cost and accuracy together. 7
For an operating team, a practical measure is:
Cost per acceptable outcome = all workflow operating costs, including failed attempts and recovery, divided by outcomes that meet the agreed standard.
Include the remaining human handling time in that operating-cost measure. Report implementation costs separately and include them in the investment calculation. Compare the same mix of work and the same quality standard before and after the change.
Speed deserves separate attention, too. An agent might reduce customer waiting time without reducing employee effort. That can be worth paying for, but the case should explain why: fewer lost orders, lower abandonment or a service commitment the company otherwise cannot meet. Do not add a speculative revenue gain merely because the response arrived sooner.
Saved time needs somewhere useful to go
Suppose an agent saves each member of a team half an hour a day. The resulting capacity may be valuable, but multiplying those hours by salaries does not automatically reveal a cash saving.
The treatment depends on what management can actually change.
| Benefit | Evidence to look for | How to treat it |
|---|---|---|
| Lower expenditure | Reduced overtime, contractor invoices or another avoidable cost | A cash benefit, net of any transition costs |
| More useful output | Additional work completed with the same team and acceptable quality | A capacity benefit; value it separately unless an economic outcome is demonstrated |
| Additional sales | Incremental sales attributable to the changed process | Use contribution after relevant variable costs, not the full revenue figure |
| Better service or lower risk | A measured improvement in service, quality or loss exposure | Track directly; monetize only where the assumptions can be defended |
Avoid counting the same hour twice. If employees use saved time to generate additional sales, do not claim both their entire salary-equivalent time saving and the full profit from that additional work as independent benefits.
Also check where the bottleneck moves. Faster preparation does little for total throughput if approvals are already the constraint. The useful question for the team manager is specific: what work will we do with the released capacity, and what will stop us from doing it?
There is nothing wrong with investing in better service or less tedious work. Present those as the reasons for the investment when they are the reasons. A project becomes difficult to defend when its approval depends on cost reductions nobody expects to make.
A worked example: useful time savings, a disappointing first year
Consider a hypothetical service operation with 6,000 requests a month that are suitable for an agent-assisted process. These are illustrative assumptions, not research findings or industry benchmarks.
The team routes 60% of eligible requests through the new process. Across those requests, it saves an average of eight minutes of human effort after allowing for routine review, failed attempts and manual fallback. Assume the required quality standard is maintained.
| Assumption or calculation | Monthly amount |
|---|---|
| Eligible requests | 6,000 |
| Share routed through the agent-assisted process | 60% |
| Requests routed through the process | 3,600 |
| Average net human time saved per routed request | 8 minutes |
| Total human time released | 480 hours |
| Illustrative fully loaded hourly labor cost | €40 |
| Salary-equivalent value of released capacity | €19,200 |
| Share of that value assumed to become actual avoidable expenditure | 50% |
| Cash benefit under that assumption | €9,600 |
| Additional recurring operating costs | €6,500 |
| Net recurring cash benefit | €3,100 |
The 50% conversion assumption needs its own evidence, such as a documented reduction in contractor spending. It is not a standard discount to apply whenever measurement is inconvenient.
Assume the €6,500 covers €3,500 for software, usage and hosting, €2,000 for platform maintenance and monitoring, and €1,000 for ongoing assurance and training. These are costs beyond the case-level human work already included in the eight-minute saving, so that work is not counted twice. Initial implementation costs another €40,000.
With those assumptions, first-year cash benefits are €115,200. Implementation and recurring costs total €118,000. Using (benefits − costs) ÷ costs, first-year cash ROI is approximately −2.4%, and simple payback takes about 13 months.
That calculation already assumes the full monthly benefit starts immediately. A gradual rollout would push payback further out. It also excludes taxes, financing and discounting; a larger or longer-lived investment needs a proper cash-flow model.
Now change just one assumption. If only 40% of eligible requests use the process, the cash benefit falls to €6,400 a month. With the same committed running costs, the operation loses €100 a month before recovering any implementation expense.
The technology could still be useful in both scenarios. The first might meet a company's investment criteria once later benefits are considered. The second needs a change in adoption, costs or performance. Neither should be approved on the strength of the €19,200 capacity figure alone.
More autonomy can bring more supervision
The cost of human involvement needs particular attention with agents. Reviewing a completed action, resolving an exception and maintaining permission rules are all work, even when the system is described as autonomous.
The 2026 ICML paper Measuring Agents in Production provides a useful view of deployment practice. Melissa Pan and colleagues combined 20 case studies with survey analysis covering 86 systems in production or pilot use. The data were collected in 2025. Teams often favored constrained workflows and human intervention, and reliability was the leading development challenge. This is evidence about how sampled systems were operated, not a causal study of their profitability. 8
For the order-change process, it may be sensible to let an agent gather records and prepare a change while keeping an approval requirement for exceptions. The additional autonomy of executing every request is worthwhile only if the added benefit justifies the monitoring and recovery burden, within the company's risk limits.
Human review itself also needs testing. A Nature Human Behaviour meta-analysis by Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone examined 106 experiments published between 2020 and mid-2023. On average, human–AI combinations performed better than humans alone, but worse than the better of the human-only or AI-only alternatives. Results varied by task. This predates current agents and is not an argument for removing required oversight. It challenges the assumption that adding a reviewer automatically improves the result. 9
Give reviewers the information and authority they need to intervene. Measure whether they catch consequential errors, how long checks take and how often decisions are reversed later. An approval button tells you very little about the quality of the approval.
For actions with serious consequences, a favorable average return is insufficient. Set explicit permission boundaries and stopping conditions. A small pilot with no serious incidents cannot establish that a rare, costly failure is impossible.
Run a pilot that can change the investment decision
Before starting, agree what evidence would justify expansion and what would lead you to stop. A pilot that can only produce a recommendation to buy more is a demonstration with a longer calendar invitation.
Begin with a baseline that includes handling time, quality and rework. Use a representative mix of requests, not just cases the project team knows the agent handles well. Keep a simpler automation option in the comparison wherever it is plausible.
Where practical, randomly route comparable eligible cases to the current and new processes, or use a phased rollout with a credible comparison group. Keep the quality criteria consistent. A before-and-after comparison alone can confuse the effect of AI with changes in demand, staffing or the difficulty of the incoming work. Describe such results as preliminary when those factors cannot be separated.
Have the process owner and finance partner agree how benefits will be recognized. For a cost-reduction project, name the expenditure expected to fall. For a growth project, establish how additional output will become additional contribution. For a service project, choose the service improvement you are prepared to fund.
The weekly review can stay small: cost per acceptable outcome, total human effort per case, customer waiting time, correction rates and the share of eligible work using the process. Inspect a sample of exceptions together. Averages can conceal an expensive group of difficult cases.
Ask employees what the numbers miss. Are they checking work twice because the review instructions are unclear? Is the approved tool missing information they need? Have faster routine cases left one specialist with a growing queue of difficult decisions? Treat their answers as diagnostic evidence, then check the relevant process data. Employee estimates can reveal friction without being reliable measures of financial return.
Budget for the period after the pilot as well. Someone must own testing when models or connected systems change, update instructions and handle incidents. Microsoft's playbook is right to include ongoing management and reusable controls in the delivery lifecycle; leaving those activities out of the budget would make the comparison unfair. 1
When to proceed, and when to leave the agent out
A strong candidate has enough recurring work to justify the investment, a clear standard for a good result, and a credible way to use the time or capacity released. It also has an owner who can change the surrounding process when the pilot exposes a problem.
Be more cautious when volumes are low, every case demands expert reconstruction, or the business case depends on staffing savings that cannot realistically be achieved. If an integration or a rules-based workflow performs comparably at lower cost, choose it. The objective is to improve the operation, and an agent is one possible means of doing that.
Some early spending can be justified as learning. Give it a defined budget, a question to answer and a review date. Do not keep renewing an unsuccessful operating project by relabeling its losses as strategic learning.
The most useful result of a pilot may be a narrower deployment than the sponsor first imagined: a few high-volume request types, limited permissions and a well-supported exception team. Expand when the measured economics and operating evidence justify it. Until then, a smaller commitment preserves the option to learn without turning an uncertain benefit into a large fixed cost.
Understand the adoption conditions before expanding
Behaviture AI Adoption Pulse helps leaders examine aggregate employee-reported evidence about approved-tool fit, policy clarity, training needs and agentic-AI readiness. These signals can help identify why a promising workflow is struggling to gain traction. They are diagnostic inputs, not proof of ROI: financial benefits still need to be established through workflow measurements and finance records.
Sources and further reading
-
Microsoft. The ROI Playbook for AI Apps and Agents (2026), especially pp. 3–6, 16 and 18–20. The supplied PDF is the source used here; Microsoft's technology-leader resource page lists the playbook. Vendor-authored guidance, not an independent evaluation of returns.
-
Forrester Consulting. The Total Economic Impact of Microsoft Foundry, February 2026. Commissioned by Microsoft. See pp. 3 and 9 for the composite organization, p. 11 for the productivity-recapture assumption, and p. 35 for disclosures. The ROI is a modeled, risk-adjusted three-year result.
-
Brynjolfsson, E., Li, D., and Raymond, L. “Generative AI at Work.” The Quarterly Journal of Economics, 140(2), 889–942, 2025. Peer-reviewed field study of an assistant used by human customer-support workers.
-
Becker, J., Rush, N., Barnes, E., and Rein, D. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. Research preprint, July 2025. Randomized experiment involving 16 developers and 246 tasks; its findings concern the tools and work studied at that time.
-
Becker, J., Rush, N., Cunningham, T., Rein, D., and Mahamud, K. “We are Changing our Developer Productivity Experiment Design.” METR, February 24, 2026. Follow-up methodological update explaining selection effects and uncertainty about the size of newer tools' productivity gains.
-
Humlum, A., and Vestergaard, E. Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI. NBER Working Paper 33777, revised March 2026. Working paper linking Danish adoption surveys to administrative labor-market records. Also see the authors' Brookings discussion, published October 6, 2026. The observation period ends in December 2024; publication in 2026 does not make this a study of 2026 agents.
-
Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., and Narayanan, A. AI Agents That Matter. Transactions on Machine Learning Research, May 2025. Peer-reviewed research on cost-aware evaluation, benchmark design and reproducibility. The authors' project page explains the findings and links to the paper and code.
-
Pan, M., and colleagues. Measuring Agents in Production. Proceedings of the 43rd International Conference on Machine Learning, PMLR 306, 95649–95695, July 2026. Peer-reviewed conference paper. The deployment analysis concerns 86 production or pilot systems selected from 306 survey responses, alongside 20 case studies; question-specific sample sizes vary. Full-text preprint.
-
Vaccaro, M., Almaatouq, A., and Malone, T. “When combinations of humans and AI are useful: A systematic review and meta-analysis.” Nature Human Behaviour, 8, 2293–2303, 2024. Peer-reviewed meta-analysis of 106 experiments and 370 effect sizes. Its distinction between improving on humans alone and outperforming the better standalone alternative is important when interpreting the results.
October 7, 2026