Enterprise AI deployment gap: Most 'agents' are just chatbots, surveys show
A series of VentureBeat Pulse Research surveys finds that 71% of enterprises report a quarter or fewer of their deployed "agents" are true multi-step workflows rather than simple chatbot wrappers, while infrastructure spending outpaces the ability to measure, secure, and trust what is being built.
Enterprise AI is being built faster than it can be controlled, according to a new wave of VentureBeat Pulse Research surveys spanning more than 500 organizations across five separate studies. The central finding: a widening gap between the ambition of autonomous agent deployment and the reality of what is actually in production.
The most telling signal comes from a survey on agent orchestration, which asked 101 enterprises to assess their own portfolios honestly. Seventy-one percent said a quarter or fewer of their deployed "agents" are true multi-step orchestrated workflows rather than single-prompt chatbot wrappers. Only 10% of organizations have crossed the halfway mark. The orchestration layer is being built well ahead of the orchestrated portfolio it is meant to run.
Yet enterprises are consolidating fast onto the major model platforms. Anthropic's Claude Platform & Agent Skills leads as the primary orchestration platform for 40% of respondents, more than double any rival, followed by Microsoft AI Foundry/Copilot Studio at 18% and OpenAI's Agents SDK/Responses API at 13%. The choice is driven by "model gravity" — native alignment with a state-of-the-art base model (21%) — and success is judged by reliable multi-step execution (task completion reliability 32%, multi-step workflow management 28%).
**The compute visibility gap**
Infrastructure spending is accelerating well ahead of the ability to track its economics. A separate survey of 107 enterprises found that only one in five (21%) run AI in production at scale, yet spending intentions are outrunning that maturity. The single largest planned area enterprises intend to evaluate over the next year is AI-specialized clouds (45%), a layer almost none use today.
The compute already in place runs cold. Eighty-three percent of organizations report GPU utilization of 50% or less, and fewer than half (44%) can rigorously track what their AI compute costs. A clear majority (64%) plan to switch or add an infrastructure provider within twelve months, and 38% within the next quarter. When choosing, enterprises prioritize integration with the existing stack (41%) and total cost of ownership (35%), not headline token price (8%).
**The security gap**
AI agents are being given real access to systems and data while controls lag behind, according to a second survey of 107 enterprises. More than half (54%) have already experienced a confirmed agent security incident (18%) or a near-miss caught before harm (36%). The structural weakness is identity: only about a third (32%) give every agent its own scoped, managed identity, while the rest report that some agents share credentials or run on shared API keys. Only three in ten enterprises (30%) isolate their highest-risk agents in sandboxes.
The security stack is overwhelmingly borrowed from model providers rather than purpose-built for agents. OpenAI's guardrails (51%) lead, followed by Google's and Microsoft's cloud controls. Satisfaction with this borrowed stack averages 4.2 out of 5, yet a clear majority plan to change tooling within the year. Enterprises are satisfied with controls they are simultaneously preparing to replace.
**The context gap**
A survey of 101 enterprises on RAG and context infrastructure reveals that retrieval-augmented generation is already the default context source (38%), and provider-native retrieval — OpenAI's file search (40%) and Google's Vertex AI Search (38%) — has overtaken dedicated vector databases. Yet a majority of organizations (57%) report that in the past six months their AI agents produced confident, wrong answers traced to missing or inconsistent business context. More than half of those said it happened more than once.
The infrastructure to fix the problem is being built: 58% already run or are building a governed semantic layer, but for most it is not yet in production. A plurality (36%) say they intend to keep best-of-standalone tools rather than consolidate onto a provider's native context stack, even as actual usage leans provider-native.
**The evaluation gap**
A survey of 157 enterprises on agentic reliability found that half of organizations (50%) have deployed an agent or LLM feature that passed their internal evaluations and then caused a customer-facing failure in the past year, with a quarter seeing it happen more than once. Trust in the tests themselves is thin: only 5% say they fully trust automated evaluation today, and the single most-cited limitation is that evaluations align poorly with real-world outcomes (29%).
Yet two-thirds of organizations (66%) already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to allow it within twelve months (33%). The most common primary evaluation tools are the model providers' native evals, tied with having no dedicated tooling at all (17% each). Only about a quarter of enterprises run real-time quality checks on live production traffic. The autonomy is arriving faster than the assurance.
**The cost of autonomy**
Across all five surveys, a consistent pattern emerges: enterprises are spending aggressively on AI infrastructure while lacking the visibility, security, context reliability, and evaluation rigor to manage what they are building. Fiscal control is a particular laggard. More than a quarter (27%) of organizations in the orchestration survey have no real-time way to stop a runaway agent before the bill arrives.
The surveys collectively paint a picture of an industry that is deploying faster than it can measure, securing agents with borrowed controls, trusting context that is already failing, and shipping evaluations that do not align with real-world outcomes. As one survey put it, the gap is between the autonomy enterprises are granting their agents and the controls in place to contain them.
Artículos relacionados
También te puede interesar




