AI rarely looks dangerous in a demo.
A model answers a few hundred questions, an agent completes a handful of tasks, and the failure rate appears small enough to ignore. Then the same system reaches production, handles millions of interactions and begins touching real customers, infrastructure and business processes. The tiny error that disappeared inside a pilot suddenly becomes a problem that repeats at scale.
That was the fault line Matt Rosoff, Editor in Chief of The Register, kept returning to during his HumanX Amsterdam conversation with Robert Blumofe, EVP and CTO of Akamai. The session, Trusting AI in the Wild, was officially framed around what happens when AI leaves controlled environments and enters networks and systems where threats are active and outcomes become harder to predict.
Rosoff eventually asked the question every enterprise experimenting with agents has to confront: what changes when AI moves from the test environment into a real production workflow?
Blumofe reduced the answer to two words: scale and reliability.
“It’s all fun and games in the lab,” he said. A failure once in several thousand interactions can be practically invisible when a team runs only hundreds of tests. At millions or billions of interactions, that same error rate stops looking small. Scale does not only multiply AI’s useful outputs; it multiplies its mistakes.
That observation led the conversation somewhere more interesting than the usual debate over whether the next model will be more accurate.
Blumofe argued that companies may need to stop giving agents every capability available and then trying to restrain them afterward. Borrowing from cybersecurity’s principle of least privilege, he proposed another rule: “least capability.”
Rosoff immediately challenged the idea. Large language models are general-purpose and inherently unpredictable, so how do you restrict capability in practice?
Blumofe pointed to something much simpler than retraining the model: the tools surrounding it.
Modern agent frameworks may arrive with web search, browsing, memory, command-line access and code execution. But most enterprise agents do not need all of them. A retail assistant whose job is to recommend a shirt does not need arbitrary code execution. If that capability is never exposed, the agent cannot unexpectedly use it.
The principle is almost disarmingly simple: start with the job, then give the agent only the abilities required to finish it.
That distinction becomes especially important inside businesses. Blumofe separated broad personal assistants from application-specific agents. A personal agent may legitimately need a wide range of tools; an agent helping a bank customer, recommending a car or assisting a shopper has a far narrower purpose. For those systems, he argued, least capability can become part of the architecture from the beginning.
The conversation then moved from what an agent should be allowed to do to where a company should watch it doing it.
Rosoff asked how teams could build resilience when an agent begins drifting toward an unwanted result. Blumofe shifted attention away from the language model and toward the agentic harness, the software that turns model output into actions.
An LLM, he noted, generates text. The harness is what invokes a tool, passes the parameters and receives the result. That makes it one of the most useful control points in the entire system. Every tool call can be observed, logged and checked as it happens.
That changes the trust problem significantly.
Instead of asking an almost philosophical question such as “Can we trust what the model is thinking?”, an enterprise can ask something much more enforceable:
What is the agent trying to do next, and should the system permit it?
Rosoff then pushed on another popular answer to AI uncertainty: put a person in the loop.
Blumofe’s response produced the session’s sharpest exchange.
“Please don’t put a human in the loop.”
He was not suggesting that humans disappear. Quite the opposite. His objection was to reducing people to the final safety mechanism for an automation system, staring at machine output and repeatedly clicking approve or deny.
“I don’t want to be a guardrail,” he said.
Blumofe instead argued for meaningful human-agent interaction: software should handle the repetitive controls it can enforce consistently, while humans remain involved where judgment, creativity and responsibility actually add value.
That is a more demanding design philosophy than simply inserting an approval step into every workflow. It asks companies to decide not only when a human must intervene, but why that intervention deserves human attention in the first place.
The discussion was grounded in a broader infrastructure shift already underway inside Akamai. Blumofe explained that the company, after expanding from content delivery into cybersecurity and cloud computing, has added GPU capabilities because it expects AI to become part of much of the enterprise software stack. He also expects inference to become increasingly distributed as applications require lower latency and more bandwidth closer to users, devices and other agents.
But Rosoff saved the cleanest question for the end:
Is there anything we should not trust AI agents with?
Blumofe answered without trying to rescue the word “trust”:
“I don’t think you want to fully trust it with anything.”
The statement sounds pessimistic until the second half of his argument. He was not saying AI cannot operate in consequential domains. He was saying trust should belong to the engineered system, not to one probabilistic component inside it.
A well-designed application can combine AI with deterministic software, constrained tools, monitoring, security controls and meaningful human involvement. In that architecture, the model is allowed to be useful without being treated as infallible.
That may have been the most important idea to leave the HumanX stage.
The AI industry spends enormous energy asking whether models are becoming smart enough to trust. Blumofe reframed the challenge: the model does not have to become perfectly trustworthy if the system around it is engineered to expect imperfection.