OpenAI’s September 29 DevDay announcements included a public beta of the Agents API, a managed service for running agents with tools and long-running sessions. The interesting part for developers who already use MCP servers and agent skills is not another agent demo. It is the decision to put those two kinds of capability inside a hosted agent environment.
That changes where you make deployment decisions. A tool connection, a skill’s instructions, and the environment in which commands run are related, but they are not interchangeable.
What OpenAI announced
OpenAI describes the Agents API as powered by the Codex harness, with orchestration, long-running sessions, and tool use. Its DevDay recap also names computer use, multi-agent capabilities, tool search, tool calling, and context compaction among the capabilities brought into applications. Those are product claims from OpenAI, not evidence that a particular third-party integration works unchanged in the beta.
The developer overview gives builders a useful vocabulary. An agent contains the model, instructions, tools, and MCP servers available to it. An optional environment is the sandbox or computer where it can access files, load skills, and run commands. A session is a durable instance that works on tasks and responds to input. This division is worth keeping in mind even if you never use OpenAI’s service.
It tells you what to test separately. Can the agent discover the right tool? Can the environment actually load the skill? Does a later turn in the session have the context it needs? A single successful prompt proves very little about the other two.
MCP servers provide actions; skills provide procedure
An MCP server exposes an interface to a system. An agent skill supplies task-specific instructions, examples, and sometimes scripts or other assets. The distinction is easy to blur when both appear in an agent’s configuration. The failure modes are different.
Consider a research assistant preparing a product brief. An MCP connection might let it search a knowledge base or retrieve current issues. A research skill could specify how to compare sources, distinguish quotes from inference, and organize a brief. If the MCP connection fails, the assistant lacks evidence. If the skill is missing, it may have the evidence but no reliable method for using it. Our skills-versus-MCP guide covers that boundary in more detail.
The Agents API’s environment model makes another boundary visible: installing a skill somewhere does not mean a hosted agent can read it. Check where the skill files live, which environment the session runs in, and what loading mechanism that environment supports. Do not assume that a local Claude Code skill or a directory install command becomes portable just because the hosted service says it can load skills.
What to verify before moving a workflow
Start with one narrow task and record its inputs and expected output. Then inspect the surfaces separately:
- Tool access. List the MCP servers the agent can actually reach, their authentication method, and whether a tool call can write or spend. A server’s presence in a directory does not grant the hosted environment permission to call it. Use our MCP evaluation checklist for the trust questions.
- Skill loading. Confirm the exact skill files are present in the environment and that the agent uses their instructions on a test task. A skill’s compatibility label is not a substitute for a run in your target environment.
- Session behavior. Run a task across turns, including an interruption and a resumption. Inspect what was retained, what was compacted, and whether the agent can still cite the source of a decision. Context compaction is useful, but a summary is not an audit log.
- Side effects. Put approval in front of messages, deployments, purchases, or record changes. Tool search may make an action easier to find; it should not silently widen the agent’s authority.
If a workflow includes browser or computer use, test it with an account and environment that can tolerate mistakes. A successful read-only run should come before a run that can change a live system. Screenshots, prompts, and tool results can all contain sensitive information; decide what you retain before turning on detailed traces.
What this does not settle
The announcement does not make MCP server quality uniform, certify directory install commands, or guarantee that skills written for other agents will behave the same way in a hosted environment. It also does not eliminate the choice between a managed execution service and a self-hosted agent. Those are questions about permissions, portability, observability, and cost, not just model capability.
For teams already experimenting with multiple agents, the most useful response is a small compatibility test. Pick one read-only MCP server and one skill with a clear output contract. Compare the hosted run with your current setup, including a failed tool call and a resumed session. If those cases are legible, expand the workflow. If they are not, adding more tools will only make the diagnosis harder.
FAQ
Does the Agents API replace MCP servers?
No. OpenAI’s developer overview explicitly includes MCP servers among an agent’s available capabilities. The API provides a managed place to run an agent; a server still provides access to a specific outside system.
Will an existing agent skill work without changes?
Not necessarily. Verify that the environment can load the skill’s files and run any scripts or commands it requires. Compatibility depends on the actual environment and skill, not the word “skill” in both products.
Is this generally available?
OpenAI introduced the Agents API in public beta. Check the current documentation for availability and supported features before designing a production dependency around it.