Klarna Said Its AI Did the Work of 700 People. Then It Started Hiring Humans Again


The most-cited AI deployment in enterprise history has a second chapter almost nobody reads carefully — and it isn’t the story you think it is.


On 27 February 2024, Klarna published a press release that would end up in more consulting decks than possibly any other document of the AI era.

The Swedish buy-now-pay-later company announced that an OpenAI-powered customer service assistant had handled 2.3 million conversations in its first month — two-thirds of all its support chats — and was doing the equivalent work of 700 full-time agents. Resolution time had dropped from 11 minutes to under two. Repeat inquiries fell 25%. The projected profit impact for the year was $40 million. Customer satisfaction, the company said, was on par with human agents.

Fifteen months later, the same CEO told Bloomberg the company had gone too far, that quality had suffered, and that Klarna was recruiting humans again.

That’s the story most people know: AI hype, AI reversal, lesson learned. It’s also wrong in a way that matters, because it misses what happened in between — and what Klarna’s numbers were doing the entire time the correction was being reported.

Two years on, with an IPO, several earnings calls and a broader industry reckoning now on the record, it’s possible to tell the whole thing properly. It’s a better story than the headline version, and a considerably more uncomfortable one.

The number that wasn’t what you think

Start with the famous figure, because almost everything downstream depends on misreading it.

“The work of 700 full-time agents” was a workload-equivalence calculation, not a headcount action. It described volume handled. It did not describe 700 people who lost their jobs on launch day. The reduction that followed came largely through attrition and a hiring freeze — and much of the displaced labour was outsourced agents supplied by third-party vendors, not Klarna employees at all. Siemiatkowski would later say directly that people weren’t laid off; the company stopped using outsourced agents.

The confusion had a source, and it was almost comically unlucky. Klarna had laid off roughly 700 people in 2022, during a period when its valuation collapsed from about $45.6 billion to $6.7 billion. Same number, different event, two years apart, no causal link. The company told reporters the two were unconnected. The headline — “Klarna replaces 700 workers with AI” — wrote itself anyway, and has been repeated ever since by outlets that never checked.

Two other qualifiers went missing on the way to the boardroom:

The $40 million was a projection. It was framed as a forward-looking profit improvement for 2024, published on day one of the deployment. Not an audited saving. Not a realised number. An estimate made at launch by the party with the most to gain from it being large.

The satisfaction claim had no score attached. No published CSAT figure, no sample size, no breakdown by conversation type. Just the phrase “on par with human agents.”

Hold onto that last one. It turns out to be the whole case.

Why Klarna, and why then

Deployments don’t happen in a vacuum, and two pressures explain why Klarna moved faster and louder than anyone else.

The first was survival economics. The 2022 valuation collapse was brutal and public. Cost discipline stopped being a preference and became the organising principle of the company.

The second was the IPO. Klarna was building a story for public markets, and an AI-efficiency narrative is unusually legible to investors: it converts a speculative technology bet into a margin line you can put on a slide. In 2024, that narrative was worth real money.

It’s also worth understanding how total the commitment was. This wasn’t a support-desk experiment. Klarna was the first European company and first fintech globally to ship a ChatGPT plugin. Siemiatkowski publicly described the company as OpenAI’s favourite guinea pig. Klarna dropped Salesforce and Workday and retired something like 1,200 smaller software services — not primarily to save licence fees, but to consolidate its data so it could be fed into AI systems more easily. Ninety percent of employees were reported to be using generative AI tools daily.

So when the February 2024 announcement landed, it wasn’t a product launch. It was the proof point for a company-wide thesis. And critically, it was co-published with the vendor — OpenAI ran it on its own site, with its COO quoted praising Klarna’s implementation.

Two parties, both with a strong interest in the number being big, and neither with any incentive to publish the confidence interval.

What Klarna actually built

The technical work deserves more credit than the discourse gives it. This wasn’t a chatbot bolted onto a help page, and if you’re building something similar the architecture is the most useful part of the case.

The assistant ran on a GPT-4-class model licensed from OpenAI, heavily prompted for BNPL and banking context. Responses were grounded in retrieval over Klarna’s help centre, policy documentation and historical ticket resolutions — a real RAG implementation rather than a model answering from parametric memory.

It had live authenticated access to account data, transaction history and payment schedules, so it knew who you were and what you’d bought before your first message. And critically, it could act: reschedule a payment, issue a refund, close an account. That’s the difference between a deflection tool and an agent, and it’s why the automation numbers were as high as they were. It covered refunds, returns, disputes, cancellations, invoice errors and identity protection across 35+ languages, 24/7.

Two pre-launch decisions are worth stealing outright:

Klarna cleaned up its help-centre content before deploying. Inconsistent or contradictory policy documents are a direct hallucination source — the model faithfully retrieves your mess and presents it with confidence. Budget real time for this.

The January 2024 beta ran in human-review mode, with every response checked before it went out. That’s an expensive way to run a pilot and the correct one.

The execution was good. The execution was never the problem.

How the claim travelled

There’s a mechanism here worth naming, because it will happen to your numbers too.

Klarna’s press release was, read closely, quite carefully worded. “Equivalent work of.” “Estimated profit improvement.” “On par with.” Every load-bearing claim carried a hedge.

Then it entered secondary coverage. Forbes ran it. The trades ran it. LinkedIn ran it approximately nine million times. At each hop, a qualifier fell off — because “equivalent workload of 700 FTEs during a growth phase, per company estimate” is not a headline and “AI replaced 700 workers” is.

By the time it came back around, it had been laundered from company-reported forward projection into established benchmark for what AI does to customer service headcount. Executives started arriving at meetings asking why their support function couldn’t do what Klarna’s had.

The claim was never independently audited. It didn’t need to be. It had been repeated enough to feel verified.

The correction

On 8 May 2025, Siemiatkowski told Bloomberg that Klarna was hiring human agents again.

His stated reasoning was specific: cost had become too dominant a factor in how support was organised, and the result was lower quality. Investing in the quality of human support, he said, was the way forward for the company. He framed it primarily as a brand and trust issue — that it was critical for customers to know a human is always available if they want one.

Note what that is and isn’t. It isn’t “the model hallucinated” or “the technology failed.” It’s an admission that the evaluation criteria were wrong — that optimising a support organisation on cost produces something that costs less and is worse.

But look at what Klarna actually did next, because it is not “rehire the call centre.”

The Uber-style workforce

The replacement model was genuinely novel, and it’s the least-discussed part of the case.

Klarna began recruiting a gig-structured support workforce: remote agents setting their own hours, logging in and out as availability allowed, piloted initially in Sweden with pay starting around 400 Swedish krona. It targeted students, people in rural areas, and — most unusually — its own most engaged customers.

The logic there is actually sharp. A passionate Klarna user already understands the product, the edge cases and the emotional texture of a payment going wrong. Onboarding cost is low. Motivation is high. Siemiatkowski described hiring the company’s own customers as the most exciting project underway.

And the explicit target was replacing the thousands of outsourced agents Klarna had been contracting from third-party vendors.

That reframes the entire episode. This was never humans replacing AI. It was a shift from outsourced human labour to AI plus a differently-structured, more precarious, better-informed human layer. By September 2025, Klarna was also reassigning engineers, marketers and analysts into support through an internal talent pool.

Whether gig-model support work is an improvement on outsourced contact-centre work is a genuinely open question, and one the coverage almost entirely skipped in favour of the tidier “AI reversal” frame.

The part everyone skips

Here’s what breaks the morality tale.

Klarna never scaled the AI back. On its Q3 2025 earnings call, the company reported the assistant doing the work of 853 agents — up from 700 — with $60 million in savings, up from the original $40 million projection. Siemiatkowski said Klarna continued to see demonstrable value and would keep investing. Company NPS was cited at 73. Klarna also stated in its Q1 2025 earnings that there had been no drop in consumer satisfaction following the rollout.

So the AI got bigger while the humans came back. Both things happened simultaneously, and any account that presents this as a retreat from AI is simply not describing the data.

Forrester analyst Kate Leggett’s read was that Klarna had overpivoted on its AI strategy — that it optimised for cost containment without weighing the longer-term customer experience consequences. That’s a more precise description than “reversal,” and Klarna itself has pushed back on the reversal framing, pointing out that the assistant still handles the equivalent of 800-plus roles.

By 2026 the strategy had settled into something explicitly two-tier. At London SXSW, addressing the rehiring headlines directly, Siemiatkowski offered a line about two things being true at the same time — and described human customer service as something that will always be a VIP offering, drawing an analogy to handmade furniture retaining its value after mass production arrived. AI takes the routine tier. Humans become the premium tier.

And the long-run automation ambition never reversed at all. By early 2026 he was describing headcount having fallen from roughly 7,000 to 3,000, with an expectation of fewer than 2,000 by 2030. Revenue per employee, he said, had gone from around $300,000 to $1.3 million since 2022, with a chunk of the salary savings redirected into cash and equity compensation for remaining staff.

By the financial metrics Klarna set for itself, the strategy worked. That’s what makes the case interesting rather than merely cautionary.

So what actually went wrong?

If the deployment hit every metric it declared and still required a public correction, the metrics were the problem. Four failure modes, in rough order of importance.

1. Everything measured was throughput

Volume, speed, cost, workload equivalence. Not one launch metric captured resolution quality on difficult cases, escalation friction, or trust. The organisation then optimised precisely what it measured, which is what organisations always do.

2. A blended satisfaction average is designed to hide the failure that matters

This is the technical heart of the case and it generalises far beyond Klarna.

Suppose AI handles the routine two-thirds of contacts at genuine parity with humans, and degrades meaningfully on the remaining third — fraud claims, financial hardship, disputes. Your blended CSAT barely moves. The average is dominated by the easy majority. Meanwhile the exact interactions that drive churn, complaints, regulatory exposure and word-of-mouth damage are getting worse, and your dashboard is green.

Klarna’s “on par with human agents” claim, published without any breakdown by interaction type, was structurally incapable of detecting this. It’s not that they measured and got it wrong. It’s that the measurement they chose could not have surfaced the problem.

3. Support tickets are not uniform in stakes

“Where’s my order” is high-volume, low-consequence and highly repetitive — a perfect automation candidate. “I’ve lost my job and I can’t make this payment” is low-volume, high-consequence, emotionally loaded, and requires judgment about ambiguity and discretion about policy.

A system tuned for the first category will fail the second in ways that are expensive and statistically invisible. Volume-weighted metrics will always tell you to automate more. The stakes distribution is what should tell you where to stop.

4. Automation shipped without a decision architecture

The missing layer wasn’t a better model. It was the governance around it: a defined confidence threshold at which the AI hands off, real-time detection of boundary breaches, a monitored escalation path, and a designed handoff experience for someone who is already frustrated.

There’s a related trap here. Reporting suggests the assistant often functioned as a gateway to human support rather than a terminal resolution. That inflates your automation-rate metric while adding a step to the customer’s journey — the number goes up and the experience goes down.

An honest caveat, which most write-ups omit

I want to be straight about the evidence, because a lot of commentary on this case is not.

The “quality collapsed” narrative rests almost entirely on the CEO’s own admission. Klarna’s own reporting denies a satisfaction drop. Dozens of articles assert that CSAT fell and complaints rose. Very few cite a published figure, because there isn’t one.

So what we actually have is an executive describing a quality problem against his own company’s interest — which is meaningful evidence, and probably the most credible kind available here — plus no independent measurement of its size.

That absence is itself the lesson. The same lack of granular public data that let the launch claim go unchallenged now makes the correction impossible to verify. Nearly every famous number in this story — 700, 853, $40M, $60M, “on par,” NPS 73 — is company-reported and unaudited.

The headcount figures don’t even reconcile with each other. Depending on the source and date, the reduction is 5,527→3,422 (per the IPO prospectus), 5,500→3,000, 7,400→3,000, or 7,000→3,000. Different measurement dates and different treatment of contractors explain some of it. The rest tells you how loosely the topline numbers were being handled by everyone, including the company.

If you cite this case, grade your figures. Regulatory filings are one thing; a number from an earnings call is another; a number from a vendor’s blog post about the case is a third thing entirely.

Klarna wasn’t an outlier — it was early

The most important development since 2025 is that Klarna stopped being a single anecdote.

Workforce planning firm Orgvue found that 39% of business leaders had made staff redundant because of AI — and that 55% of them later concluded the decision was a mistake. Staffing firm Robert Half reported that around 32% of US hiring managers who eliminated a role citing AI later restaffed the same or a similar position. Forrester’s 2026 predictions suggest more than half of AI-attributed layoffs will be quietly reversed.

The concrete cases are instructive because they rhyme with Klarna’s:

Ford deployed automated quality-control systems and roughly 900 AI cameras, then found the technology was missing defects that experienced technicians caught. It brought back somewhere in the region of 300–350 veteran engineers, and subsequently topped the JD Power 2026 Initial Quality Study. A Ford executive’s summary was blunt: the technology is only as good as the data used to train it.

Commonwealth Bank of Australia reversed more than 40 customer service roles after an AI voice bot struggled under real call volume. IBM automated large portions of its HR operations, then announced plans to triple US entry-level hiring — executives noting that humans remained essential for judgment, ethics and complex decisions.

Meanwhile the cutting continues: Challenger, Gray & Christmas recorded 87,714 AI-attributed job cuts through May 2026, already past the 54,836 logged across all of 2025.

The pattern across all of them is identical, and it is not “AI doesn’t work.” It’s that organisations cut based on generic benchmarks rather than task-level analysis of what a role actually absorbed — including the invisible work, the judgment calls, and the institutional knowledge that never appeared in anyone’s process documentation.

Klarna’s distinction is only that it went first, went loudest, and then had the unusual honesty to say so publicly.

Seven things to take from this

  1. Publish the denominator. A workload-equivalence figure without a stated automation rate, a segment breakdown and an audit basis is marketing, not measurement. Apply that scepticism to other companies’ numbers and to your own.
  2. Segment your quality metric before you launch, not after. Report satisfaction by interaction type — routine, complex, high-stakes. The blended average is the number that will lie to you most convincingly.
  3. Map the stakes distribution, not just the volume distribution. Volume tells you where automation saves the most. Stakes tell you where it costs the most when it’s wrong. Only one of those appears on a standard support dashboard.
  4. Design the escalation path first. What does the AI decide alone? At what confidence does a human take over? How do you detect a boundary breach in real time? What does the handoff feel like to someone who’s already upset? Klarna had weak answers to all four in 2024.
  5. Distinguish cost avoided from cost eliminated. Avoided hiring during growth and displaced vendor contracts are not layoffs. Conflating them is how a defensible efficiency story becomes an indefensible headline.
  6. Audit what the role actually absorbed. The recurring failure across Ford, CBA, IBM and Klarna is the same: work that nobody documented, done by people everyone treated as interchangeable headcount.
  7. A guaranteed human option is a brand asset, not an efficiency leak. Klarna’s landing position — scale the AI aggressively and guarantee human access — treats them as complements. That’s the actual lesson, and it took an expensive detour to reach it.

The uncomfortable ending

There’s a version of this story where everyone learns AI can’t replace people, we all feel reassured, and nothing changes. That version isn’t true either.

Klarna’s real endpoint is that talking to a human becomes a premium tier — something you get if you’re valuable enough to the company to warrant it. Siemiatkowski’s own analogy is handmade furniture: mass production didn’t destroy craft, it made craft a luxury good.

Apply that to financial services and the implication is genuinely uncomfortable. The customers least likely to qualify for VIP treatment are disproportionately the ones in hardship, dispute or fraud situations — exactly the people who most need a human being with discretion, patience and the authority to make an exception. A two-tier support system doesn’t just ration convenience. It rations judgment.

That’s the question the 700-agent headline never asked. It’s a harder one than whether the chatbot worked, and it’s the one that will still matter in five years.


All Klarna performance figures cited here are company-reported and have not been independently audited. Principal sources: Klarna’s February 2024 press release and OpenAI’s joint announcement; Bloomberg (May 2025); TechCrunch and IT Pro (June 2025, London SXSW); Klarna Q1 and Q3 2025 earnings materials; CX Dive, Fortune and Forbes (2024–2026); CNBC (July 2026) on the broader reversal pattern; Orgvue, Robert Half and Forrester survey data as reported.

References

Primary company and vendor material

  1. Klarna. “Klarna AI assistant handles two-thirds of customer service chats in its first month.” Press release, 27 February 2024. https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/the original 700-agent, 2.3M-conversation, $40M claim
  2. OpenAI. “Klarna’s AI assistant does the work of 700 full-time agents.” Customer story. https://openai.com/index/klarna/the joint vendor announcement; source for the 90% internal AI adoption figure and the OpenAI COO quote

The launch and its reception (2024)

  1. Torres, Gene. “Klarna’s New AI Tool Does The Work Of 700 Customer Service Reps.” Forbes, 13 March 2024. https://www.forbes.com/sites/quickerbettertech/2024/03/13/klarnas-new-ai-tool-does-the-work-of-700-customer-service-reps/representative of how the claim entered mainstream business coverage
  2. Orosz, Gergely. “Klarna’s AI chatbot: how revolutionary is it, really?” The Pragmatic Engineer, 2024. https://blog.pragmaticengineer.com/klarnas-ai-chatbot/the best contemporaneous technical scrutiny of the deployment

The correction (May 2025)

  1. Bloomberg News. Interview with Sebastian Siemiatkowski, 8 May 2025 — the origin of the “gone too far” and “lower quality” admissions. Cited here via items 6–8.
  2. Kahn, Jeremy. “Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver.” Fortune, 9 May 2025. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/the cost-as-dominant-criterion quote; the Uber-style pilot and pay details
  3. Shibu, Sherin. “Klarna Is Hiring Customer Service Agents After AI Couldn’t Cut It on Calls.” Entrepreneur, 9 May 2025. https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396the “always a human if you want” framing; recruitment targeting
  4. CX Dive. “Klarna changes its AI tune and again recruits humans for customer service.” 9 May 2025. https://www.customerexperiencedive.com/news/klarna-reinvests-human-talent-customer-service-AI-chatbot/747586/the goal of replacing outsourced agents
  5. FinTech Weekly. “Klarna Reverses Course on AI Customer Support, Resumes Human Hiring.” 12 May 2025. https://www.fintechweekly.com/magazine/articles/klarna-hires-customer-service-after-ai-pivotthe gateway-to-human-support observation
  6. eMarketer. “Klarna backtracks AI customer service plans.” https://www.emarketer.com/content/klarna-backtracks-ai-customer-service-plansIPO timing context

The two-tier synthesis (2025–2026)

  1. Wiggers, Kyle / Sawers, Paul. “Klarna CEO says company will use humans to offer VIP customer service.” TechCrunch, 4 June 2025. https://techcrunch.com/2025/06/04/klarna-ceo-says-company-will-use-humans-to-offer-vip-customer-service/London SXSW; “two things can be true”; the Salesforce/Workday and 1,200-services consolidation
  2. IT Pro. “Klarna CEO might be calming his tone on AI replacement – but the future of customer service could be a two-tier nightmare.” June 2025. https://www.itpro.com/technology/artificial-intelligence/klarna-ceo-sebastian-siemiatkowski-calming-ai-replacement-customer-service-nightmarethe “VIP thing” quote and the luxury-goods analogy
  3. Curry, Rachel. “AI enabled Klarna to halve its workforce—now, the CEO is warning workers…” Fortune, 10 October 2025. https://fortune.com/2025/10/10/klarna-ceo-sebastian-siemiatkowski-halved-workforce-says-tech-ceos-sugarcoating-ai-impact-on-jobs-mass-unemployment-warning/headcount and post-IPO market cap
  4. CX Dive. “Klarna says its AI agent is doing the work of 853 employees.” 20 November 2025. https://www.customerexperiencedive.com/news/klarna-says-ai-agent-work-853-employees/805987/Q3 2025 earnings: 853 agents, $60M, NPS 73; Forrester’s Kate Leggett on overpivoting; the Q1 2025 no-satisfaction-drop claim
  5. CX Dive. “Klarna pursues ‘Uber’ style customer service model.” 20 February 2026. https://www.customerexperiencedive.com/news/klarna-pursues-uber-style-customer-service-model/812763/headcount trajectory and the sub-2,000 target for 2030
  6. Time / Charter. “What Klarna Learned from its Ambitious AI Rollout.” 28 April 2026. https://time.com/partner-content/charter/7378651/what-klarna-learned-from-its-ambitious-ai-rollout/hiring customers as agents; revenue per employee
  7. Bigeye. “Klarna’s AI customer service deployment — AI Autopsy 002.” June 2026. https://www.bigeye.com/blog/klarnas-ai-customer-service-deploymentclose reading of the press release’s hedged wording

The wider reversal pattern (2026)

  1. Lee, Justina. “Employers who laid off workers for AI are reversing their decisions.” CNBC, 1 July 2026. https://www.cnbc.com/2026/07/01/employers-who-laid-off-workers-for-ai-are-reversing-their-decisions.htmlOrgvue 39%/55%; Robert Half 32%; Ford, Commonwealth Bank and IBM cases
  2. Ellison, Dana. “The AI Rehire Correction Paradox.” Forkast, 4 August 2026. https://forkast.news/the-ai-rehire-correction-paradox-87k-cuts-ytd-but-a-third-of-companies-are-already-reversing-course/Challenger, Gray & Christmas cut totals; Ford’s 900 cameras and JD Power result; Forrester Predictions 2026; Klarna’s dispute of the “reversal” framing
  3. American Bazaar. “Companies reverse AI-driven layoffs as human workers return.” 1 July 2026. https://americanbazaaronline.com/2026/07/01/companies-reverse-ai-driven-layoffs-as-human-workers-return-483875/corroborates the Orgvue and IBM figures

Secondary analysis consulted but not relied on for facts

  1. Okhrem, Paul. “Klarna’s AI: 700 Agents, Then a Rehire.” https://paul-okhrem.com/companies-using-ai/klarna-ai/useful for its explicit source-reliability grading; the IPO prospectus headcount figures are cited there
  2. Twig. “Klarna AI Saved $40M on Support — Then Walked It Back.” https://www.twig.so/blog/klarna-ai-customer-support-efficiencyvendor blog; the automation-rate arithmetic is worth reading, the conclusions are commercially motivated

Leave a Reply

Discover more from Geeky Codes

Subscribe now to keep reading and get access to the full archive.

Continue reading