The OpenAI Agents API and Agents SDK place operational responsibility in different places. The API provides a managed agent runtime. The SDK runs the agent loop in your application. Choosing between them starts with who should operate the workflow, where tools execute, and which product state your application must own.
OpenAI's September 29 DevDay recap highlighted the Agents API, making this a timely architecture decision for teams adding AI workflows to SaaS products. This guide uses the official documentation checked October 7, 2026 and a worked product scenario. The recommendations are MUBBITS's design guidance, not claims that either option guarantees a reliable business process.
Key takeaways
- Choose runtime ownership before comparing convenience features.
- A self-hosted execution environment attached to the Agents API still uses a managed agent runtime.
- Keep customer authorization, business records, and final write decisions in your application design.
- Compare both options on one representative workflow, including interruptions and recovery.
What is new, and what decision does it create?
OpenAI's DevDay 2026 recap describes the Agents API as bringing Codex capabilities such as multi-agent orchestration, tools, context compaction, and computer use into a managed offering for developers. The relevant product question is whether delegating more of the runtime removes work your team actually wants to stop operating.
Start from a real workflow. A short classification inside an existing request handler may not need a long-running agent at all. A research task that works with files, pauses, and continues later creates different demands. List those demands before selecting an architecture because it appears in a launch announcement.
Agents API: managed sessions and orchestration
The API overview describes a managed Codex harness with durable sessions, orchestration, context compaction, and recovery. Applications provide instructions and tools, submit work, and receive progress through streams or webhooks. An environment is optional when the task does not need its own compute or files.
For a product team, the useful question is which operational components this replaces in your proposed system. Map the managed capabilities against your existing queue, application records, support tools, and failure handling. Avoid keeping duplicate orchestration merely because the prototype already has it.
Managed execution still needs a product contract. Define what the customer requested, which inputs were authorized, how completion is recognized, and who can view the artifact. Those requirements exist regardless of where the agent loop runs.
Agents SDK: application-owned deployment and integration
OpenAI's SDK documentation describes a code-first route in TypeScript or Python. Your server owns deployment, tool implementations, storage, and approval decisions while the SDK runs the agent loop. It also provides routes for orchestration, guardrails, resumable state, and observability.
Consider this approach when the workflow needs to live closely inside existing product code. Examples include a business rule engine that already controls each step, a team with established background-job operations, or an integration whose lifecycle depends on application-specific state. These are reasons to prototype the SDK, not automatic disqualifications for the API.
Budget for the operations you retain. Assign owners for worker availability, retries, upgrades, monitoring, and incident response. Direct control is valuable when somebody is equipped to exercise it; otherwise it can become an unplanned maintenance commitment.
Separate the agent runtime from its execution environment
The API architecture has three parts: OpenAI's harness, an optional execution environment, and your application server. An OpenAI-hosted environment is provisioned by OpenAI. With a self-hosted environment, your code supplies compute and manages its lifecycle while the harness remains hosted. No environment is also an option for suitable tool-based tasks.
Draw these boundaries before a security or infrastructure review. The location of command execution does not, by itself, establish where every request, model input, event, or stored record goes. Document the actual data flow and review the relevant terms for your requirements.
The following ownership map is deliberately narrow. It identifies who operates the loop and compute; it is not a claim about certification, complete data residency, or whether a particular private system can be connected without additional controls.
| Approach | Agent loop | Execution compute |
|---|---|---|
| API, hosted environment | OpenAI-managed | OpenAI-managed sandbox |
| API, self-hosted environment | OpenAI-managed | Your provisioned environment |
| SDK in your application | SDK in your application | Your chosen runtime and tools |
OpenAI documentation: harness, environment, and application architecture
Keep runtime, compute, and business ownership distinct
Worked example: a customer account review report
Imagine a SaaS feature that summarizes an account's activity, checks permitted support records, and drafts a renewal briefing. The deliverable is an editable report for an authorized team member. Sending that briefing to a customer is a separate action. This is a hypothetical design example, not a deployed MUBBITS case study.
Create an application job with a customer identifier, initiating user, permitted source scope, and requested output. Have each data-access tool validate that scope on the server. A generated account identifier is an input to validate, not authority to read another tenant's records.
Let the agent produce a draft and source references. Save a versioned artifact, show the user which sources were available, and allow correction. If the product later supports sending the report, bind that action to the exact reviewed version and recipient. A revised draft should not silently inherit permission granted to an earlier one.
For this scenario, either runtime can be viable. The differentiator is how well it fits the team's job lifecycle, tool interfaces, review experience, and operating capabilities. Test those boundaries before spending time on a polished chat interface.
Which agent architecture fits your product?
Bring one business workflow and the systems it must access. We can map ownership, prototype constrained tools, and test recovery before expanding the feature.
Keep business state independent of transient events
Use an application-owned job record to explain progress to users and support staff. For example, your product might use requested, running, awaiting review, completed, and failed. These are suggested application states, not OpenAI API status names. Define which validated event or action permits each transition.
Give side-effecting operations a stable identity and a server-side duplicate check. If an export succeeds but the response is lost, a retry should locate that result or safely continue instead of creating an unintended second export. The exact mechanism depends on the destination service and its transaction model.
Do not equate receiving an event with committing a business outcome. Store the result and the relevant audit context, then acknowledge completion according to the integration's contract. Design a reconciliation path for jobs that remain in progress after a handler crash.
Provide a useful cancellation experience too. Explain whether cancellation prevents new work, interrupts current execution, or only hides the job. Already-completed external actions need their own recovery behavior. These product semantics should survive a future runtime change.
Compare total cost and failure behavior in a prototype
Build the same small workflow using the architecture you currently favor, then use the alternative only where a meaningful uncertainty remains. Recreating the whole product twice is rarely a sensible first step. Measure accepted report quality, elapsed time, tool failures, review effort, and support visibility.
For budgeting, the API overview identifies model usage, tools, and hosted sandbox charges as relevant billing components. Your business model also needs engineering operations, storage, monitoring, and human review. Use current rates and measured usage rather than assuming managed infrastructure is always cheaper or always more expensive.
Include interrupted jobs in the comparison. A design that looks efficient on successful runs can become costly when staff repeatedly investigate missing artifacts or duplicate actions. Record which failures require manual intervention and how long that intervention takes.
- Submit the same request twice and inspect whether it creates duplicate business outcomes.
- Interrupt the application handler and confirm that progress can be reconciled.
- Deny a tool permission and verify that the report explains the missing evidence.
- Change access rights before a later tool call and verify the server checks current authority.
- Exhaust the task budget and confirm the product offers a clear, recoverable next step.
Make the architecture decision reviewable
Write a short decision record with the workflow, expected volume, tool boundaries, state owner, cost assumptions, and recovery responsibilities. Include a diagram and the evidence from the prototype. State the conditions that would make the team revisit the decision.
Favor the managed API when its operational model fits the work you need to delegate. Favor the SDK when integrating the loop into your existing application is the clearer ownership model. Keep a simpler deterministic workflow in consideration when the task does not benefit from agent planning.
MUBBITS can help turn that decision into a scoped implementation: a working job lifecycle, constrained tools, representative evaluations, and a release plan. The useful first deliverable is evidence that the selected architecture completes your business task reliably enough for its intended use.
Frequently asked questions
Is the Agents API the same as the Agents SDK?
No. The API uses a managed runtime, while the SDK runs the agent loop within your application. Compare operational ownership and integration needs before choosing.
Does a self-hosted sandbox mean the whole agent is self-hosted?
No. In the Agents API architecture, self-hosting the execution environment does not move the hosted harness into your infrastructure. Review every relevant data flow separately.
Do we still need an application database for a managed agent?
For most business products, you still need authoritative customer permissions, job records, and committed outcomes. Decide which records belong to your application instead of treating an agent conversation as the entire business database.
Which option is cheaper for a SaaS product?
There is no universal answer. Compare measured model, tool, and compute usage with engineering operations and human review for your workload, including failed and interrupted jobs.
Let’s work through your next step.
Tell us what you’re building, what you’ve tried, and where you need a hand. We’ll work with you to define a practical way forward.

