When you're weighing whether an agent could hold part of your company, the picture in your head is probably a chat window. It drafts something, you judge the prose, and the decision gets made on small talk.
You'd never hire a person that way. You hand a hire access to the systems, with the permissions set properly and someone governing what they can touch, and then you shadow them, teach the job, correct the small things. An agent that genuinely holds a role got that same treatment. Access governed, corrections written down so they stick. Chat was the interview. Nobody should run part of a company on the interview.
Here are the four things I'd want answered before I signed anything, taken from my last company's numbers.
The operation happens in writing you can open. 32 procedures, about sixty thousand words in total, on top of a 5,341-word statement of how the company decides things and where it stops. You can read what the business will do before it does it. 21 of those 32 name something that needed my approval before it was allowed to happen, so a person holds the stops. That's also the part you can do: approve or block, and let the rest run.
The old baseline for that much written judgment was a consultant's binder, or more often nothing at all, the rules living in the head of whoever had been around longest.
Two people ran the company across 18 countries and 11 languages. The normal company's answer to that footprint is a hire per region or per language cluster, plus the layer that manages those hires.
Ramp is where it gets unfair. For an average role the model starts stronger than a new person does on day one, and you train it in writing, so a correction lands once and stays landed. A quarter of hand-holding becomes days of it. And this is the worst these models will ever be, which means next year's agents start from a higher floor than this year's.
Payroll and the managers of that payroll. It also removes the daily doubt about whether anyone is following the rules, because the rules sit inside the work instead of in a document.
30 of the 56 use no AI model at all. Their part of the job was simple enough that fixed rules cover it, and you can't hire a person for only the easy slice of a job. A single human gives you exactly one setting. With agents the thinking scales up or down per part of the job.
Two thirds of everything written down names a point where the work stops and waits for a person. That ratio is the governance. It's also the part nobody shows you in a demo.
Switch the agents off and work stops across 56 roles. Refilling them means hiring across 18 countries and 11 languages, which is a program with a budget, not an afternoon. The payroll you avoided comes due. What you keep either way is the sixty thousand words, because that's the asset. Procedures on a shelf just don't run.
An agent has no instinct to escalate, and it never becomes suspicious. We had a check whose whole job was to stop anything reaching a customer who had asked not to be contacted. It spent weeks approving every single message and reporting it had blocked nothing, because it was pointing at a file that didn't exist. Nothing failed and nothing alarmed. I've stopped presenting agents as better than people, because of stories like that one.
A person doing that job walks over and tells you the empty report looks wrong. That instinct is doing more work inside most companies than anyone credits, and no agent brings it. So the honest construction is written checks, a person reviewing outcomes on a schedule, and the judgment-heavy points routed to an approver, which is why 21 of those procedures routed back to me. You get capacity and a company you can read. You give up the suspicious employee.
If you want to know which parts of your operation an agent could genuinely hold and which parts still need a person, email me.