by datastudy.nl

Friday, July 24, 2026

AI

OpenAI Presence bundles engineers with enterprise AI agents

OpenAI Presence is a managed service for enterprise AI agents sold as a project with engineers attached. It resolves 75% of support line issues without humans.

Pie chart showing 75% of inbound issues on OpenAI's support line resolved by the agent without human assistance, and 25% requiring human handoff. Enterprise AI agents managed service data from AI News reporting on OpenAI figures.
OpenAI's support line agent resolves 75% of inbound issues without human assistance, with 25% requiring human handoff. Source: AI News reporting on OpenAI figures. Data Today benchmark.

OpenAI has spent three years selling API keys and telling enterprises to build their own enterprise AI agents. On July 22, it changed the deal. OpenAI Presence is a managed service where the company's Forward Deployed Engineers show up at your business, scope a single workflow, and run the agent into production themselves. You do not buy it online. You cannot self-serve it. The price is not published.

The headline number comes from OpenAI's own phone support line, where the agent now resolves 75% of inbound issues without human assistance. Those are OpenAI's figures on OpenAI's channel, measured against OpenAI's grading criteria. The transparency is welcome, the numbers are not independently verified, and they tell you where the enterprise AI agents market actually breaks.

Enterprise AI agents have a deployment problem that model quality does not fix. Gartner has warned that more than 40% of agentic AI projects will be cancelled by the end of 2027, attributing failures to governance gaps, undefined business value, and weak operational discipline rather than model capability. The companies shipping agents into production have learned the same lesson: the hard part is integration, permissions, change management, and evaluation. A vendor that sends engineers to do that work is responding to what buyers have actually been failing at.

What exactly did OpenAI ship on July 22?

Presence is a managed product delivered through a limited general availability programme, as AI News reported. It is not an API endpoint or a seat licence. Each engagement starts with a single job: resolving a billing dispute, handling an insurance claim, clearing an IT service request. The agent receives only the knowledge and system access that the job requires. The customer writes the rules governing what it can do, when it needs sign-off, and when a person takes over.

OpenAI's documentation is candid about the labour involved. The help centre lays out a six-stage process: scoping business outcomes, security and privacy and legal review, simulation and acceptance testing, staged rollout, and post-launch iteration. A Presence agent does not become production-ready by ingesting documents. After launch, Codex reads production sessions and escalations, then proposes changes the customer's team tests and approves before rollout.

The delivery model borrows directly from Palantir. Forward Deployed Engineer is a Palantir title for staff embedded in customer operations for months at a time, and the economics look nothing like metered inference. Software scales; engineers cleared into a bank's core systems do not.

Does the managed model solve the real problem?

The Gartner diagnosis identifies governance, business value, and operational discipline as the failure modes for agentic AI. Almost everything Presence bundles targets that diagnosis directly. Simulations and graders test whether an agent reached the right outcome, followed policy, used its tools correctly, and escalated when it should, all before anyone outside the company speaks to it. Guardrails intervene when an interaction moves past defined boundaries. Session records and action histories give reviewers something to audit. Escalation paths hand a person structured context rather than a cold transcript.

This is the gap we flagged when we looked at the enterprise AI agent evaluation gap: half of enterprises ship agents without rigorous evaluation, and the agents break in predictable ways. Presence is OpenAI's answer to that gap, and it is an answer priced in engineering time rather than tokens.

The model makes business sense for OpenAI in a way that pure API sales do not. Enterprise buyers who have been burned by agent deployments are wary of another dashboard that leaves the integration work to them. A vendor that takes responsibility for the deployment, even at the cost of scaling slowly, can charge enterprise prices and create relationships that API access alone cannot lock in.

How good are the proof points?

The strongest single proof point is OpenAI's own English-language phone support line, 1-888-GPT-0090. The company says the agent met or exceeded its internal benchmarks for frontline human support within weeks, now resolves 75% of inbound issues without human assistance, and cut human handoffs by 15 percentage points in ten days through the Codex improvement loop. If the current handoff rate is 25%, the 15 percentage point reduction implies a starting handoff rate of 40%, meaning the agent's resolution rate climbed from 60% to 75% in ten days.

The chart below shows the before and after of those two metrics.

Dumbbell chart showing agent resolution rate rising from 60% to 75% and human handoff rate falling from 40% to 25% on OpenAI's support line, a 15 percentage point shift over ten days through the Codex improvement loop.
OpenAI's support line: agent resolution rate rose from approximately 60% to 75% over ten days, while human handoffs fell from approximately 40% to 25%. Before figures derived from OpenAI's stated 15 percentage point reduction. Source: AI News reporting on OpenAI figures. Data Today benchmark.

Those are OpenAI's numbers, measured against OpenAI's grading criteria, on OpenAI's own channel. Take them seriously and treat them with appropriate caution. The same report names three customers: BBVA is exploring voice support for everyday banking in Mexico. SoftBank is testing Japanese-language conversations. IAG is exploring support during high-demand events such as severe weather. Daniel Ordaz, head of AI transformation at BBVA Mexico, calls the bank a design partner helping shape voice experiences for financial customer service. All three are in exploration or testing. None is presented as running Presence at scale, which is worth holding alongside the word "battle-tested" in OpenAI's framing.

What does Presence mean for your build versus buy call?

If you are scoping an enterprise agent deployment, Presence changes the calculus in three concrete ways:

  • Delivery capacity, not model capability, is the binding constraint. OpenAI says access depends on workflow fit, implementation readiness, and available delivery capacity. The third item is a consulting constraint. If your workflow fits and you can get on the schedule, you trade control for speed. If you cannot, you are back to building with the API.
  • Accountability lines need to be written into the contract. When the model vendor is also the implementation partner, the contract must specify who is responsible for a policy misapplied in production, how quickly issues get fixed, and what happens when the model configuration changes. Teams that have spent the past year building evaluation suites against specific versions will want the contract to say what they are held to when the configuration moves.
  • Pricing is opaque by design. Implementation scope and cost are set per customer and per deployment, which is ordinary for enterprise services and still leaves buyers without a public reference point for cost per resolved contact against an incumbent contact-centre vendor. You will need to benchmark against your current cost per contact, including the fully loaded cost of human agents, to know whether Presence is competitive.

The model behind Presence is not named. The documentation says it uses OpenAI models with configuration selected for the workflow, subject to change as the workflow evolves. That flexibility is defensible engineering, because pinning a production agent to a frozen model version ages badly. It also means you are buying a service, not a version, and the service level is what the contract says it is.

The agent cost problem we tracked earlier does not disappear under Presence. It shifts from per-token pricing to per-project pricing, which may be more predictable but is not necessarily cheaper. The cost question moves from "how many tokens did the agent burn" to "how many engineering hours did the deployment require, and what is the ongoing cost of the Codex improvement loop."

What is the real scaling risk?

Forward Deployed Engineers do not multiply the way API endpoints do. OpenAI has put its own FDEs and named global systems integrators at the front of every deployment, which works while volumes are small. It becomes more complicated when the waiting list grows and the integrators OpenAI relies on to scale are also its competitors in the consulting layer.

The model also puts a governance question on the table for buyers. When the vendor that built the model is also running it in your production environment, the lines of accountability for a policy misapplied in production need to be written down rather than assumed. Data handling follows the same pattern: the signed architecture and contract are the governing record, not any published policy. Channel support during limited GA covers voice or chat, with contact-centre integration, routing, authentication, and handoff design confirmed deployment by deployment.

Presence sits apart from ChatGPT Workspace Agents, which remain the self-serve path for teams building inside ChatGPT and Slack. Voice customers keep API access to OpenAI's frontier models. OpenAI now offers broadly the same capability three ways, separated less by what the technology can do than by who does the work. That leaves buyers choosing on delivery capacity as much as on model capability, and on OpenAI's own account, delivery capacity is the part being rationed.

The contract is the product

Presence is a bet that the bottleneck for enterprise AI agents has moved from the model to the plumbing: the guardrails, the evaluation, the integration, the change management, the audit trail. OpenAI is selling the plumbing the way Palantir sells it, with engineers embedded in your operations and a contract that governs what happens when things go wrong. If you are buying, read the contract before you read the model card. If you are building, note that the most valuable AI company in the world just decided the hard part is everything around the model.

Sources