AI transformation

How to Tell Whether Your AI Agent Actually Works

You often hear executives say: “We’ve deployed AI.” But ask for the details — what exactly changed in the way work gets done, which tasks the agents actually took over, how many steps left the process — and the picture blurs. Is the business easier to run now? Usually there is no answer, because nobody agreed up front on what was supposed to change.

This is not a local problem. MIT’s “The GenAI Divide: State of AI in Business 2025” examined 300 enterprise deployments: against $30–40 billion invested, 95% of pilots produced no measurable impact on the P&L. McKinsey’s “State of AI 2025” shows the same gap from the other side: 88% of companies use AI regularly, 39% see any effect on operating profit, and only 6% report a material one — above 5% of EBIT. Almost everyone deployed. Almost nobody earned.

The difference isn’t the model. Those 6% change a process and measure a process; the rest buy a tool and measure impressions.

What deploying an AI agent actually means

Deploying an agent is not “plugging in a neural network.” It is handing over a complete piece of work. Not “answer questions about contracts,” but “receive the incoming document, identify its type, register it, assign an owner and set a deadline.” An agent is a colleague with duties, system access and boundaries: what it does alone, and what it must escalate to a human.

This is exactly where most projects break. The agent is launched on top of email, spreadsheets and paper — on top of a process that doesn’t exist inside a system. It has no history of decisions to learn from, no route to validate against, nobody to hand the result to. So it answers in generalities, everyone shrugs, and the project joins the 95%.

Why we turned this into a methodology

We learned it on ourselves and formalised it as DATM — the Documentolog AI Transformation Methodology. It rests on one hard rule: an AI agent only enters a process that already lives inside the system. The order never changes:

  1. Digitise. The process moves into d8n in full — routes, roles, deadlines, documents. Not “some requests in the system, some in WhatsApp” — the entire flow in one channel.
  2. Describe. Regulations, nomenclature, routing rules, the knowledge base of the process. An agent needs not just data, but the rules by which that data is read.
  3. Only then — the agent. It takes the routine: classification, registration, choosing the assignee, tracking deadlines.

Between step two and step three we put a formal gate: at least 80% of the process document flow must run through the system for four consecutive weeks. If the threshold isn’t met, no agent is launched — however tempting the demo. This isn’t bureaucracy, it’s protection from disappointment: an agent on an unprepared process works at roughly a third of its potential.

The second DATM rule is equally inconvenient and equally important: every agent has a named C-level owner who is accountable for its results and runs a weekly review of its work. Not “IT deployed it” — a specific executive with that agent in their KPIs.

We also fix quality thresholds below which an agent never reaches users: hallucinations under 1% (0.3% for financial processes), answer relevance above 85%, citation accuracy above 90%, adoption above 70%, user NPS above +40.

The full methodology, with stages and timelines, is on our DATM page. Incidentally, the same MIT report explains why a methodology with an external partner matters at all: 67% of projects reach production when outside expertise works alongside the client’s internal team, against 22% for those going it alone with their own IT.

What it looks like on a live process: the Records agent

Take the registry office — handling incoming correspondence. A process every organisation has, and everyone considers simple until they measure it.

What the person did before the agent:

  • opened each incoming item (email, portal, paper) and read it in full to grasp the substance;
  • identified the document type and the right nomenclature entry;
  • created the record card by hand: details, sender, date, deadline;
  • decided from memory and experience which department and which person it should go to — and if the sender was new, went to ask colleagues;
  • set the due date according to regulation;
  • manually tracked who had missed their deadline and chased them.

Every step costs minutes, and a non-standard document costs hours of waiting for whoever “knows where this goes.” Registering a single incoming item took up to half an hour.

What the person does after the agent:

  • reviews a queue that is already processed: type identified, card filled in, assignee proposed, deadline set;
  • confirms or overrides the agent’s decision where the case is non-standard;
  • handles exceptions — what the agent honestly flagged as ambiguous and escalated;
  • watches the overall deadline picture instead of chasing individual reminders.

The role shifts from executor to supervisor. By our internal data, routine operations in this process fell by roughly 80% — and only then did we apply the same method to HR, support, meetings and executive analytics.

Measuring the process — before

Before deploying anything, take one process and answer four questions honestly:

  1. How long does it take from start to result?
  2. How many times does a human touch the task by hand?
  3. Where does it get stuck most often?
  4. What does each error or missed deadline cost — in money, in clients, in reputation?

That’s your baseline. Without it, any conversation about effectiveness turns into an exchange of feelings. And take it before you start: measured in hindsight, a baseline always turns out conveniently shaped.

Which metrics to set

A metric must describe the process, not the technology. Working reference points for typical processes — the figures on the right are targets we set when scoping work, not promises of outcome:

ProcessMetricHow to countTarget
Records / incomingDocument registration timefrom arrival to completed card−50% or better
Records / assignmentsShare of overdue assignmentsoverdue / total per month−30%
Any document flowManual touches per documentaverage human actionsmultiple-fold reduction
SupportFirst response timefrom ticket to first replyminutes instead of hours
SupportShare of tickets reaching humansescalations / all ticketsdown at equal CSAT
SalesLost enquiriesunprocessed per periodto zero
Any processAgent qualityhallucinations, relevance, citation≤1% / ≥85% / ≥90%

Note the last row: agent quality is a metric too, and no less important than speed. A fast agent that makes mistakes costs more than a slow human.

How the ones who succeeded do it

The best public example of counting correctly is Klarna. Look at the metrics they reported for their support assistant: resolution time fell from 11 minutes to 2, repeat enquiries dropped 25%, the assistant handled two-thirds of chats — the equivalent of 700 agents and roughly $40 million of profit impact in a year. Not a word about a clever model, only before-and-after process numbers. A year later Klarna publicly admitted it had cut too deep and brought people back into premium support. That’s the second lesson of the same case: an agent needs boundaries, and moving them based on data is normal.

Our own examples are smaller in scale, identical in logic.

Sales funnel. We audited our own funnel and found we had lost 14 enquiries in a single week: they arrived at night and on weekends, while qualification depended on a human who sleeps. We handed the agent one step — overnight qualification of inbound enquiries. The metric is simple: lost enquiries. It was 14 a week. It is now zero.

Customer support. An agent runs first-line support: routine questions it closes itself, complex cases go to engineers with the context already gathered. The load on the support team dropped by 70% — and this isn’t about layoffs: people moved to work nobody previously had time for.

Assessing effectiveness — after

In a month or two, return to the same four questions and compare:

What you measureBeforeAfter
Time from request to resulthours / days?
Manual touches per taskN?
Errors and missed deadlines per periodN and their cost?
Where the freed-up time wentsomething useful, or thin air?

The effect is the difference between “before” and “after” minus what you paid to deploy and run it. I deliberately won’t quote an “average AI payback period” — it doesn’t exist; every process has its own. Klarna counts minutes per chat; we count lost enquiries and the share of routine removed.

The calculation stays open

This is not a teaser or a result hidden behind registration. It is the calculation itself, using the example data. For a records process with 300 operations per month, the model uses 30 → 6 minutes, 6 → 2 manual touches, errors at 12% → 4%, overdue work at 18% → 6%, and a 50% time-reuse factor.

ROI worksheet inputs: the process before and after an AI agent

LeverBefore, KZT/yearAfter, KZT/yearImpact, KZT/year
Process time2,250,000450,0001,800,000
Manual touches900,000300,000600,000
Errors and rework1,728,000576,0001,152,000
Overdue work5,184,0001,728,0003,456,000
Total annual impact10,062,0003,054,0007,008,000
First-year ownership cost2,319,300
Net first-year impact4,688,700
Payback4.0 months

ROI worksheet result: KZT 7,008,000 impact, KZT 4,688,700 net impact and 4.0-month payback

The fifth lever — manageability — is deliberately excluded from the sum. There is no honest universal formula for turning decision quality into cash.

Five impact levers: time, errors, SLA, operation cycle and manageability

Get the working ROI file

Enter your email. We will send a one-time link valid for seven days; no attachments.

The principle is the same: don’t count the price of the AI, count the price of the process before and after. That is a language your CFO, your board and your accountant all understand. And it is exactly what separates the 6% who earned from the 95% who merely deployed.

Where to start: pick the one process that hurts most and bring it to a demo — we’ll calculate its before-and-after economics together, on your data. If you’d rather see the whole route first, it’s laid out on our DATM page.

CEO DIGEST

Read new materials first

Longreads when published and one weekly digest. Email confirmation is required.