In February 2024, Klarna announced that its AI customer-service assistant, built in partnership with OpenAI, had handled 2.3 million customer conversations in its first month, equivalent to the workload of 700 full-time agents, and that the deployment had reduced average resolution time from 11 minutes to under 2 minutes [1]. In May 2025, Klarna's chief executive publicly reversed course and announced that the company would resume hiring human customer-service agents, citing a quality gap on complex and emotional interactions that the AI deployment could not close [2]. The institutional finding is not about the AI tool. It is about the workforce decision architecture that produced both the reduction and the reversal.
The cost-of-hire literature establishes the asymmetry that makes the workforce decision high-stakes. The U.S. Department of Labor estimates that a poor hiring decision costs an enterprise at least 30 percent of the employee's first-year earnings; Society for Human Resource Management research places the total cost of replacing a mid-level employee at between one-half and two times annual salary, with higher multiples in technical and senior roles [3]. Gallup's State of the Global Workplace 2024 report documented engagement at 23 percent globally and estimated USD 8.9 trillion in annual lost productivity from disengagement; Gallup's 2025 update reported a further decline to 21 percent engagement and USD 438 billion in incremental productivity loss in 2024 alone [4].
The MIT NANDA initiative's 2025 report The GenAI Divide: State of AI in Business 2025, based on 150 leader interviews, 350 employee surveys, and 300 deployment analyses, found that approximately 95 percent of corporate generative AI pilots produced no measurable P&L impact despite USD 30 to 40 billion in enterprise investment. The primary cause was organizational, not technological: brittle workflows, lack of contextual learning, and misalignment with day-to-day operations [5].
The architectural conclusion across the three data sets is consistent. The workforce decision, hiring, retention, deployment, and the integration of new tools into existing workflows, is where the operational test of an enterprise's productivity strategy sits. Where the decision architecture is sound, the cost-of-hire and engagement cost lines move in favor of the business. Where the decision architecture is partial, automated by tooling that does not fit the actual workflow, or driven by a cost target detached from the operating consequence, the costs accumulate in the form of reversed deployments, replacement hiring, disengaged retention, and the productivity gap the engagement data documents.
The diagnostic question is not whether a company has a workforce strategy. It is whether the workforce decisions, taken under cost and growth pressure, are made by people accountable for the operational consequence of those decisions, and whether the architecture distinguishes the kinds of work that compound enterprise value from the kinds of work that look identical on the org chart but produce very different operational outcomes.
The Klarna case has been read as a story about the limits of artificial intelligence in customer service. The more durable reading is about the workforce decision that produced both the original announcement and the reversal.
In February 2024, Klarna and OpenAI published a joint case study reporting that Klarna's AI assistant had handled 2.3 million customer chats in its first month, the equivalent of the workload of 700 full-time human agents; that the assistant had reduced average customer resolution time from 11 minutes to under 2 minutes; that it had produced a 25 percent drop in repeat inquiries; and that Klarna projected approximately USD 40 million in profit improvement from the deployment in 2024 [1]. The announcement was widely interpreted as evidence that AI could substitute for substantial portions of customer-facing workforce.
In May 2025, Klarna chief executive Sebastian Siemiatkowski reversed the public position. The company announced it would resume hiring human customer-service agents, with a recruitment model targeting students, rural residents, and existing Klarna users; agents would be paid 400 Swedish krona per hour starting and could choose their schedule and location within Sweden [2]. The stated reason was that the AI deployment had not met quality standards on complex and emotional customer interactions and had produced a customer-experience gap that human agents would address.
The institutional finding is architectural, not technological. The 2024 deployment did substitute for the agent capacity Klarna would otherwise have needed to add during a growth phase; the productivity gain was real on the routine inquiries that constituted the majority of volume. The 2025 reversal acknowledged that the substitution had not been complete on the segment of work that determined the customer-experience quality the company competed on. The workforce decision in 2024, made under a productivity and cost target, did not distinguish the routine-inquiry capacity from the complex-and-emotional capacity at the level of architecture; the 2025 reversal corrected the architecture under operational pressure.
The same pattern is observable across the 2024-2026 record of corporate AI-driven workforce decisions that were reversed or qualified within twelve to eighteen months of announcement. The pattern is not that AI fails as a productivity tool. It is that workforce decisions made on aggregate cost targets, without architectural distinction of the kinds of work the workforce performs, produce reversals when the kinds of work that were grouped together turn out to behave differently under operational pressure.
The workforce decision is high-stakes because the cost of the wrong decision is asymmetric to the cost of the right one. The U.S. Department of Labor's long-standing estimate places the cost of a poor hiring decision at a minimum of 30 percent of the employee's first-year earnings; the figure is referenced as a conservative benchmark in workforce-management literature [3]. Society for Human Resource Management research places the total cost of replacing a departing employee at between one-half and two times annual salary depending on the role; for technical and senior positions, replacement cost multiples in excess of 100 percent of salary are commonly cited [3].
The components of the cost are not exotic. Direct cost includes recruiting, interviewing, screening, onboarding, training, and the productivity lost during the ramp period of the replacement. Indirect cost includes the disruption to the team carrying the workload during transition, the customer or stakeholder impact of the gap, the institutional knowledge that exits with the departing employee, and the time of managers and senior personnel diverted to the search and onboarding process. For positions where the work is highly leveraged, including engineering, senior commercial roles, regulated compliance roles, and roles with primary customer or counterparty relationships, the indirect cost component routinely exceeds the direct cost.
The asymmetry is sharper for the kinds of work that the Klarna reversal isolated. The work of routine inquiry handling can be substituted at high throughput and acceptable quality by a well-deployed tool; the cost of an imperfect substitution is modest. The work of complex and emotional customer interaction is highly leveraged on the institutional reputation of the firm; the cost of an imperfect substitution is high and may be invisible in real time. The same dynamic applies to engineering work, where the cost of an imperfect substitution on routine code generation is modest and the cost of an imperfect substitution on architectural judgment is high; to legal work, where the cost of an imperfect substitution on document review is modest and the cost of an imperfect substitution on counsel on a contested matter is high; and to operational work, where the cost of an imperfect substitution on data entry is modest and the cost of an imperfect substitution on escalation judgment is high.
The architectural distinction the cost-of-hire literature requires is between the work that can be substituted at lower marginal cost without substantial downside risk and the work that cannot. The workforce decisions that hold under operational pressure are the ones that draw the distinction at the level of role architecture, not at the level of the aggregate cost target.
Gallup's State of the Global Workplace 2024 report documented global employee engagement at 23 percent, with the report estimating USD 8.9 trillion in annual productivity loss to the global economy from disengagement, equivalent to approximately 9 percent of global GDP [4]. The 2025 update reported engagement at 21 percent, a fall from 23 percent the previous year, with USD 438 billion in incremental productivity loss attributable to the single-year decline; the data set further documented that manager engagement, which had moved from 30 percent to 27 percent in the relevant period, was the principal driver of the decline rather than frontline contributor engagement [4].
The productivity-loss figures are not the institutional finding. The finding is that engagement, the operating-system condition under which the workforce decision becomes operational, is in measurable decline at the same time that companies are making the workforce decisions documented in section I, often under cost targets that do not internalize the engagement cost of the decision. A workforce reduction that achieves a near-term cost target while degrading the engagement of the workforce that remains can produce a net productivity loss that exceeds the announced cost saving; the engagement cost is delayed, distributed across the remaining workforce, and not attributed to the decision that caused it.
Gallup's data further establish that manager engagement is the leading indicator. Managers are the operating-system layer between the workforce decision the executive team makes and the operational performance the workforce produces. Where managers are disengaged, the workforce decisions made above them and the operational performance produced below them are decoupled. Workforce reductions or tool deployments announced by senior leadership do not translate into the operating consequences the leadership intended, and the gap between the announcement and the operational outcome appears as productivity loss rather than as a workforce-architecture failure.
The MIT NANDA initiative's 2025 report The GenAI Divide: State of AI in Business 2025, based on 150 leader interviews, 350 employee surveys, and analysis of 300 corporate AI deployments, found that approximately 95 percent of corporate generative AI pilot programs had produced no measurable P&L impact despite USD 30 to 40 billion in enterprise investment. The report's central finding was that the cause of failure was organizational rather than technological: brittle workflows, lack of contextual learning, and misalignment with day-to-day operations [5]. The 5 percent of pilots that succeeded had specific characteristics: deep integration with existing workflows, partnerships with specialized vendors rather than internal builds, and empowerment of line managers rather than central AI labs to drive adoption.
The MIT finding is not a verdict on AI. It is a verdict on the workforce decisions that surround AI deployment. The 95 percent failure rate documents that enterprises that approached AI as a workforce-substitution exercise without architectural distinction of the work AI could substitute, the workflows AI needed to integrate with, and the line-manager capability AI deployments needed to draw on, produced no measurable productivity gain. The 5 percent that succeeded made the workforce decision architecturally: which workflows, which roles, which managers, which integration path. The cost-of-failure asymmetry is the same as the cost-of-hire asymmetry; the architecture is what determines which side of the asymmetry the decision falls on.
The same architectural lens explains the Klarna sequence. Klarna's 2024 deployment was a partnership with a specialized vendor (OpenAI), the integration was deep, and the line-manager engagement was substantial; the deployment achieved real productivity on the routine inquiry workload. The 2025 reversal corrected the segment of the deployment where the workforce decision had not distinguished the work the AI could substitute from the work it could not. Klarna's outcome is not a 5-percent success and not a 95-percent failure; it is a mixed result that was corrected under operational pressure when the workforce architecture proved partial.
The workforce decision is distributed across functions whose incentives diverge, in a pattern structurally similar to the supply chain due diligence allocation problem.
The chief executive and chief financial officer hold the aggregate cost target and the public commitment to the productivity strategy. The chief human resources officer holds the workforce policy, the talent architecture, and the engagement program. The chief operating officer or the relevant business-unit leader holds the operational consequence of the workforce decision and the workflow-level integration. Line managers hold the day-to-day translation of the workforce decision into operational performance. Where these functions hold the workforce decision as a single integrated obligation, the decision becomes a process the enterprise can defend. Where the allocation is partial, the workforce decision becomes the sum of what each function chooses to contribute under its own incentive structure.
The Klarna reversal is consistent with an architecture in which the chief executive and chief financial officer held the cost target, the chief human resources officer held the workforce reduction, the customer-service organization held the operational consequence, and the line-manager translation produced the quality gap that the reversal corrected. The architecture was not absent; it was partial. The 2025 correction added the customer-experience consequence into the workforce decision at the level the original 2024 decision had not.
For boards, the implication is the same allocation question that applies across compliance, quality, and supply chain decisions. A workforce strategy that documents headcount targets, cost-per-employee ratios, and aggregate productivity metrics is the documented strategy. A strategy that documents the operational metrics, the segmentation of the workforce by the kind of work performed, the engagement state of managers responsible for translating the strategy into operation, the integration-quality of any tools deployed against any workflow, and the cycle time from a workforce decision to the operational consequence it produces, is the operational strategy. The U.S. Department of Justice's September 2024 Evaluation of Corporate Compliance Programs, in directing prosecutors to evaluate whether a compliance program is "adequately resourced and empowered to function effectively," operationalizes at the level of compliance the same architectural standard that the engagement and productivity data establish at the level of workforce [6].
The 2024-2026 record of AI-driven workforce decisions and their reversals will be read as a story about the limits of artificial intelligence in the workplace. The more durable reading is about workforce decision architecture.
The Klarna case documents that a workforce decision made on aggregate cost grounds, without architectural distinction of the kinds of work the workforce performs, produces reversal when the kinds of work that were grouped together behave differently under operational pressure. The cost-of-hire literature documents that the cost of the wrong workforce decision is asymmetric to the cost of the right one, and that the asymmetry compounds in highly leveraged roles. The Gallup engagement data documents that the operating system around the workforce decision, the engagement of the managers who translate executive-level decisions into operational performance, is in measurable decline at the same time the decisions are being made. The MIT 95-percent finding documents that the AI deployments that succeeded were the ones that respected the architectural conditions of integration and managerial engagement that the failed deployments did not.
The leadership inquiry that follows is architectural. The workforce decision is the operational test of every productivity strategy. Where the decision is made by leaders accountable for the operational consequence, where the workforce is segmented at the level of role architecture rather than at the level of headcount, where line-manager engagement is treated as a precondition rather than as a residual, and where the integration of new tools is held to the same operational test as the workforce they are deployed alongside, the cost-of-hire and engagement-cost lines move in favor of the business. Where these architectural conditions are partial, the workforce decision will produce the same pattern Klarna's 2024 announcement and 2025 reversal made visible: a near-term cost saving that the operational consequence subsequently corrects, at a cost the architecture failed to internalize when the original decision was made.
For chief executives, chief financial officers, chief human resources officers, and the boards that supervise them, the 2024-2026 record describes operational implications that are time-bound.
Segment the workforce by the kind of work performed, not by department or grade. For each segment, identify the leverage profile (the consequence of an imperfect substitution at that segment) and the engagement state of the managers who translate executive decisions into operational performance at that segment. Identify the segments where the leverage is high and the manager engagement is at or below the Gallup global manager benchmark; these are the workforce decisions where the cost-of-failure asymmetry is highest.
For any tool deployment that has substituted for a workforce segment in the prior 18 months, including AI-enabled customer service, AI-enabled coding assistance, AI-enabled document review, or AI-enabled operational support, conduct an architectural review. The questions are: which segment of the workforce did the deployment substitute for, what is the quality gap the deployment has produced on the highly leveraged portion of that segment, what is the line-manager engagement state on the deployment, and what is the integration depth between the deployment and the surrounding workflow. The review should be structurally separated from the function that authorized the deployment.
Restructure the board-level workforce report to distinguish documented workforce metrics from operational workforce metrics. Documented metrics include headcount, cost per employee, attrition rate, hiring throughput, and aggregate engagement scores. Operational metrics include the segment-level leverage profile, manager engagement at the segments where leverage is highest, the integration quality of any tool deployment against any workflow, the cost-of-replacement experience versus the cost-of-replacement benchmark for the segment, and the cycle time from a workforce decision to the operational consequence it produces. Establish a board-level review of workforce architecture at least semi-annually, with the compensation or audit committee as the standing forum.
Treat engagement as a leading indicator, not a lagging one. Where manager engagement at any leveraged segment of the workforce declines materially between reporting cycles, escalate to the relevant board committee on the same cadence as a material financial-performance signal. Build feedback loops between workforce decisions, the operational consequences they produce, the engagement cost they impose, and the replacement cost they ultimately generate.
A material revision to the Gallup engagement methodology or the headline engagement and productivity-cost figures, a sector-specific cost-of-hire data update that supersedes the U.S. Department of Labor benchmark, a Society for Human Resource Management methodology change to the cost-of-replacement multiplier, or a follow-up to the MIT GenAI Divide report that materially modifies the 95-percent finding or the architectural success conditions would require recalibration of the segmentation, deployment-review, and reporting design.
Klarna scope: The findings cited from the Klarna 2024 deployment and the 2025 reversal reflect the company's own statements and the OpenAI joint case study at the publish dates referenced. Subsequent operational performance of the human-rehiring program, the architecture of any further AI deployments, and quantitative validation of the customer-experience quality gap that prompted the reversal will continue to develop during the post-reversal period and may refine the architectural reading cited here.
Gallup methodology: The 2024 and 2025 State of the Global Workplace figures cited reflect Gallup's published methodology and respondent base. Sector-specific applicability requires confirmation against industry-specific engagement and productivity data. The USD 8.9 trillion and USD 438 billion productivity-cost figures are macro-economic estimates derived from Gallup's engagement methodology and should be referenced as directional, not as precise enterprise-level cost figures.
MIT GenAI Divide scope: The 95-percent failure-rate figure reflects MIT NANDA's analysis of 300 corporate generative AI deployments as of the report's publication. The architectural conditions of the 5-percent success cohort identified in the report (vendor partnership, workflow integration, line-manager empowerment) are described here as the conditions the data set associates with success; alternative architectures may produce successful outcomes outside the data set.
Cost-of-hire data: The U.S. Department of Labor 30-percent benchmark and the SHRM cost-of-replacement multipliers cited reflect commonly cited workforce-management estimates and should be confirmed against the company's own historical replacement experience for the specific roles to which the estimates are applied. Sector- and role-specific cost-of-replacement figures vary materially.
Cross-sector application: The architectural lens described here applies across regulated industries in which the workforce performs heterogeneous work at varying levels of leverage. The specific applicable workforce regulations differ by sector and jurisdiction; the operational test the architecture proposes does not depend on a single regulatory framework.
[1] OpenAI and Klarna, "Klarna's AI assistant does the work of 700 full-time agents" (joint case study), February 2024. URL: https://openai.com/index/klarna/
[2] Fortune (Sasha Rogelberg), "Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver," May 9, 2025. URL: https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/
[3] U.S. Department of Labor, long-cited estimate that a poor hiring decision costs at least 30 percent of the employee's first-year earnings, as referenced in workforce-management literature; Society for Human Resource Management, cost-of-replacement research placing total replacement cost at between one-half and two times annual salary for mid-level roles, with higher multiples in technical and senior positions. URLs: https://www.dol.gov/ and https://www.shrm.org/
[4] Gallup, "State of the Global Workplace: 2024 Report" and "State of the Global Workplace: 2025 Report." URL: https://www.gallup.com/workplace/349484/state-of-the-global-workplace.aspx
[5] MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025," 2025. URLs: https://nanda.media.mit.edu/ and https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
[6] U.S. Department of Justice, Criminal Division, "Evaluation of Corporate Compliance Programs (Updated September 2024)," September 23, 2024. URL: https://www.justice.gov/criminal/criminal-fraud/page/file/937501/dl
[7] Klarna, "Klarna AI assistant handles two-thirds of customer service chats in its first month," International Press Release, February 27, 2024. URL: https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/
[8] Fortune (coverage of MIT NANDA report), "MIT report: 95% of generative AI pilots at companies are failing," August 18, 2025. URL: https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
[9] TechCrunch, "Klarna CEO says company will use humans to offer VIP customer service," June 4, 2025. URL: https://techcrunch.com/2025/06/04/klarna-ceo-says-company-will-use-humans-to-offer-vip-customer-service/
[10] Pragmatic Engineer (Gergely Orosz), "Klarna's AI chatbot: how revolutionary is it, really?" April 2024. URL: https://blog.pragmaticengineer.com/klarnas-ai-chatbot/