The first agent in any organization is a demo. It has a sponsor, a dedicated team, a Slack channel, and an audience. Everyone knows what it does and who to call when it misbehaves.
Eighteen months later the same organization has fifty agents, and nobody can answer three basic questions: how many are actually running, who owns each one, and which of them do the same thing twice.
That is not a failure story. It is what success looks like now. Agents became cheap to build, so every team built one. The problem quietly moved from construction to operations, and most organizations are still staffed, tooled, and budgeted for construction.
Sprawl is a symptom of success
There is a familiar shape to this. It happened with macros, with departmental databases, with RPA bots, with microservices. Lower the cost of creating something and you get more of it than you can track. Agents are the fastest round of this cycle yet, because the build cost dropped further and faster than for anything before them.
The difference is that an unmaintained macro corrupts a spreadsheet. An unmaintained agent holds credentials, calls production APIs, and makes decisions on behalf of the company. Sprawl with that kind of reach is not clutter. It is unmanaged operational risk.
The fleet problems nobody plans for
The problems of the fiftieth agent are not smarter versions of the problems of the first. They are different in kind:
- Ownership. The team that built an agent dissolves or moves on. The agent keeps running. When it starts failing, the incident lands on whoever notices, which usually means nobody.
- Duplication. Three departments build three "invoice status" agents against the same ERP, each slightly different, each maintained separately, each producing slightly different answers to the same question.
- Credential hygiene. Every agent holds secrets: API keys, service accounts, tokens. Multiply by fifty and you have a rotation problem, an expiry problem, and an access-review problem that no one is watching.
- Silent breakage. Agents sit on top of APIs, schemas, and prompts that other people change. An agent does not crash when its world shifts; it degrades. It starts giving confidently wrong answers, and there is no failed deployment to alert anyone.
- Cost attribution. Token spend, per team, per agent, per use case. If you cannot answer who spends what, you cannot decide what to keep.
None of these shows up in a pilot. All of them show up in a fleet.
A control plane is not a dashboard
The market has noticed. IBM introduced an Agentic Control Plane in watsonx Orchestrate in mid-2026, pulling operations, governance, and an agent catalog into one place, and followed up within months with cross-platform agent discovery and evaluation tooling. Every major vendor is converging on the same shape, which tells you where the pain actually is.
But it is worth being precise about what a control plane has to do, because "we have a dashboard" is where many programs stop. A dashboard shows you the fleet. A control plane lets you run it:
What running the fleet actually requires
- A catalog where agents are published with versions, owners, and dependencies, so a new team searches before it builds.
- Policy enforcement at the platform level: guardrails, data masking, and access controls that apply to every agent because of where it runs, not because its builder remembered.
- Credential health as a monitored property, not a wiki page.
- Evaluation as a gate. An agent update that has not passed its eval set does not ship, the same way untested code does not ship. This is where the audit-trail work we described in Auditable agents plugs in: per-decision traceability is the micro level, fleet evaluation is the macro level.
Tooling makes this possible. It does not make it happen. Somebody still has to own the operating model.
Treat the fleet like a product portfolio, not a pile of scripts
The organizations that handle this well make one mental shift: an agent is a product with a lifecycle, not a script with an author.
Products get owners, and ownership survives reorgs. Products get versioning and release notes. Products get deprecated: sunset dates, migration paths for their consumers, and an actual shutdown. In a healthy fleet, the retirement rate matters as much as the build rate. If agents are only ever added, the fiftieth-agent problem becomes the two-hundredth-agent problem on a schedule you can predict.
This is also where BPM and integration teams turn out to be sitting on exactly the right instincts. They have spent years running estates of long-lived automation: versioned, monitored, owned, decommissioned on purpose. The stack is new. The discipline is not.
What we see fail
Three patterns, over and over:
- The catalog that arrives after the sprawl. Retrofitting an inventory onto sixty existing agents is archaeology; publishing every agent through a catalog from day one is free.
- The agent with no owner. If the answer to "who owns this?" is a person who left, the real answer is nobody, and nobody is on call for it.
- The update without an eval. Someone improves a prompt on Friday, the agent's answers shift over the weekend, and the business finds out from a customer.
All three are cheap to prevent and expensive to repair. The common thread is that each one is an operating-model decision postponed until it became an incident.