Pick Pydantic AI if you want typed agent code that stays portable across model providers. Pick the OpenAI Agents SDK if OpenAI models and its native sandbox or harness are already part of your product boundary. The Pydantic AI vs OpenAI Agents SDK decision is portability versus vendor-native execution, and the points worth comparing are provider lock-in, output typing, sandbox ownership, migration cost, and tool APIs.
To see how the two behave on the same work, I ran one inventory task through both SDKs on my M1 Mac with gpt-5.4-mini, three times with default settings and three times with matched settings. Both returned the correct structured result in all 12 runs. Pydantic AI had the lower median wall time in both cells (2682.36 ms versus 3034.64 ms by default) and used fewer tokens per run (406 versus 501). One task is a data point, not a speed ranking.
What is Pydantic AI
Pydantic AI is a Python agent framework built by the team behind Pydantic, the data-validation library that underpins FastAPI. It defines agents as typed objects: an Agent takes a model string, an output_type for structured results, and a set of tools, and returns validated Python objects instead of raw JSON. The framework works against OpenAI, Anthropic, Google, Groq, Mistral, Bedrock, and several other providers through one consistent API.
The V2 release on June 23, 2026 rebuilt the framework around "capabilities" as a core primitive, bundling tools, hooks, and model settings into one composable unit. In practice it replaced scattered Agent constructor arguments with a single capabilities list. Instrumentation, tool preparation, and MCP toolsets each became a capability object instead of a keyword argument, and the default end_strategy flipped from 'early' to 'graceful', so function tools now run alongside a successful output tool rather than getting skipped, according to the Pydantic AI changelog. Pydantic recommends upgrading to v1.100.0 first to clear deprecation warnings before jumping to V2, since the API surface changed enough that a direct migration risks silent breakage.
Repository star counts change daily and do not decide which framework fits a production system. Use the release history, compatibility notes, and the same-task evaluation below instead of treating popularity as a capability measure. That history has kept moving since V2 (the benchmark below used v2.33.0), so pin the version under test and read its migration notes before comparing capabilities or upgrade cost.
What is the OpenAI Agents SDK
The OpenAI Agents SDK is OpenAI's own agent-building framework, available in Python and TypeScript. An Agent in this SDK takes a name, a model, a list of tools, handoffs to other agents, and input_guardrails/output_guardrails for validation. It is the successor to OpenAI's earlier Assistants API patterns and is built to work best with OpenAI's own models, though it can call other providers through compatible endpoints.
On April 15, 2026, OpenAI added native sandbox execution and an open-source, model-native harness to the Python SDK, letting agents run inside isolated environments from Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, or Vercel with file and tool access scoped to the sandbox. "This launch, at its core, is about taking our existing Agents SDK and making it so it's compatible with all of these sandbox providers," said Karan Sharma of OpenAI's product team, as quoted by TechCrunch. TypeScript did not get the same harness and sandbox support until the OpenAI changelog entry dated May 6, 2026, roughly three weeks later, not "June" as the update has sometimes been described.
The Python package has 27,730 GitHub stars and the TypeScript package has 3,348, which I read as both a head start and OpenAI's larger install base among API-first teams.
Pydantic AI vs OpenAI Agents SDK at a glance
| Criteria | Pydantic AI | OpenAI Agents SDK |
|---|---|---|
| Model providers | OpenAI, Anthropic, Google, Groq, Mistral, Bedrock, and more | Best with OpenAI models; other providers via compatible endpoints |
| Core abstraction | Typed Agent + capabilities (v2.0.0, June 23, 2026) | Agent + tools/handoffs/guardrails |
| Sandbox/execution | No first-party sandbox; relies on caller's runtime | Native sandbox (Blaxel, Cloudflare, Daytona, E2B, Modal, Runloop, Vercel), shipped April 15, 2026 |
| TypeScript support | No official TypeScript package | Yes, harness/sandbox parity since May 6, 2026 |
| Version used in the benchmark | v2.33.0 | v0.22.0 |
| Release context | Pin the Pydantic AI version under test | Pin the OpenAI Agents SDK version under test |
| Structured output | Native, via output_type and Pydantic models | Via tool schemas and output_type |
Both frameworks are free and open source, so licensing cost is not a differentiator; the real tradeoff is portability versus vendor-native tooling.
Typing and provider portability: the core decision
The single biggest reason teams pick Pydantic AI over the OpenAI Agents SDK is that a Pydantic AI agent does not know or care which model provider it is talking to. Swap model="openai:gpt-5.2" for model="anthropic:claude-sonnet-5" and the rest of the code, including the output_type validation, keeps working. The OpenAI Agents SDK's Agent class assumes the opposite: it is built first to be the best possible interface to OpenAI's own models.
Choosing Pydantic AI does not mean giving up OpenAI models: it treats OpenAI as one of several supported providers and, since V2, uses OpenAI's Responses API as the default transport. If your team ships to multiple model vendors, or wants to swap models without touching agent logic, that portability is worth the smaller ecosystem. If your stack is already OpenAI end to end, the native sandbox and harness reduce the amount of infrastructure you have to write yourself.
Same-task benchmark: default and controlled parity
I ran the same inventory task through both SDKs with gpt-5.4-mini, the same prompt, one local tool, the same Pydantic output schema, and the same fixture. The default cell leaves each SDK's ordinary settings in place. The parity cell aligns tool choice, output-token cap, request limits, and timeouts as closely as each SDK allows. These results describe this task, not a universal speed ranking.
| Cell (three runs each) | Pydantic AI median wall time | OpenAI Agents SDK median wall time | Tool calls and requests per run |
|---|---|---|---|
| Default settings | 2682.36 ms | 3034.64 ms | 1 tool call, 2 requests for both |
| Controlled parity | 8985.54 ms | 10232.88 ms | 1 tool call, 2 requests for both |
Both SDKs returned the correct structured inventory result in every run of both cells, with one successful tool call and two requests per run, so the differences are in speed and tokens rather than outcome. Pydantic AI had the lower median wall time in both cells, and both were slower under parity than with their defaults. Parity narrows the configuration differences, but the SDKs remain different internally.
The parity cell used a median of 406 total tokens for Pydantic AI, versus 501 for the OpenAI Agents SDK; the input/output split was 356/50 and 456/45. At the gpt-5.4-mini rates on OpenAI's pricing page, that comes to an estimated $0.000492 and $0.0005445 per run. These are rate-based estimates, not provider invoices.
A per-run token cost like that is only half the comparison, because a framework that fails more often spends the same tokens twice. Divide the same measurement by accepted outcomes, as in the AI agent cost per successful task method, before treating the cheaper SDK as the cheaper choice.
Sandboxing and execution safety
The OpenAI Agents SDK's sandbox support is the newest and most consequential difference between the two frameworks in mid-2026. Before April 15, an OpenAI Agents SDK user who wanted an agent to run shell commands or edit files had to wire up their own sandbox. Now, as TechCrunch reported, the SDK supports seven sandbox providers out of the box and standardizes filesystem tools, apply_patch for file edits, and AGENTS.md-style custom instructions.
Pydantic AI has no equivalent first-party sandbox. It assumes the caller already runs inside whatever isolation layer their infrastructure provides, whether that is a container, a Modal function, or a CI runner. That is not a gap so much as a design choice: Pydantic AI stays a thin, portable layer, and leaves execution isolation to the host application. Teams building anything that runs untrusted or agent-generated code should weigh this difference carefully; it is the one area where the OpenAI Agents SDK is meaningfully ahead as of this writing.
Migration and breaking changes
Pydantic AI's V2 migration is the bigger lift right now. The changelog lists the moves: Agent(instrument=...) becomes capabilities=[Instrumentation(...)], Agent(mcp_servers=[...]) becomes Agent(toolsets=[...]), and the ModelProfile type changed from a dataclass to a TypedDict. None of this is exotic, but it touches most existing agent definitions, so budget real time for the upgrade rather than treating it as a patch bump.
The OpenAI Agents SDK's recent releases have been comparatively low-drama. According to OpenAI's changelog, point releases through June and July 2026 fixed sandbox artifact boundary handling and realtime model defaults without breaking the public Agent API. If your team values release stability over the newest primitives, that track record matters.
Which should you pick
The framework choice and the MCP-versus-RAG choice are separate decisions. If the agent needs current structured records, compare MCP with document retrieval in the 12-query test before choosing the SDK. If document retrieval is the right path, compare open source vector databases by workload before you commit to a storage layer.
If this framework choice follows a retiring no-code workflow, the OpenAI Agent Builder migration guide covers that path. For untrusted code execution, pair the framework decision with the E2B self-hosting guide.
Choose Pydantic AI when you need more than one model provider or want other teams to plug their models into the same typed agent layer. Choose the OpenAI Agents SDK when your product is built around OpenAI models and native sandbox execution reduces infrastructure you would otherwise own. The decision is portability versus vendor-native execution, not a popularity contest.
Neither choice is permanent, but migration cost is real. Pin both packages, run the same tool and structured-output tests, and compare the migration notes before switching. Portable code matters when provider changes are likely; native sandboxing matters when OpenAI-specific execution is already part of the product boundary.
How to use the same-task benchmark in a framework decision
My recorded run makes the comparison more concrete than a framework feature list, but it is one task and one environment. First choose the architecture requirement: model portability and typed dependency injection tend to favor Pydantic AI; close alignment with OpenAI's agent primitives tends to favor its SDK. Then rerun the same prompt, tools, error case, and tracing setup on your target versions. Keep output quality, developer effort, latency, and provider lock-in as separate observations rather than collapsing them into one winner.
How I tested this
I ran the benchmark on August 24, 2026, on an Apple M1 Mac with Python 3.13.5. A clean uv environment installed pydantic-ai==2.33.0 and openai-agents==0.22.0 in 1932 ms. Each SDK got the same prompt, one local inventory tool, the same Pydantic output schema, and the same fixture against gpt-5.4-mini: three runs with default settings and three under parity settings (tool choice auto, a 400-token output cap, a 90-second timeout, and run limits matched as far as each API allows). All 12 runs were real OpenAI API calls with no mocks. Token costs use the OpenAI rates as listed on August 22, 2026.
I did not test the sandbox providers, the TypeScript SDK, non-OpenAI models, multi-tool workflows, or error handling, so these numbers say nothing about them. The raw results list every run, the pinned versions, and the fixture hash. The published record omits runtime logs and redacts local filesystem paths; the measurements are unchanged.
Related coverage
- MCP vs RAG: A 12-Query Test
- OpenAI Agent Builder Migration
- E2B Pricing and Limits
- Pydantic AI vs Microsoft Agent Framework
References
- Developers OpenAI Changelog - https://developers.openai.com/changelog/
- GitHub - Pydantic AI Releases - https://github.com/pydantic/pydantic-ai/releases
- Langfuse AI Agent Comparison - https://langfuse.com/blog/2025-03-19-ai-agent-comparison
- OpenAI API pricing - https://platform.openai.com/pricing
- Pydantic AI Changelog - https://pydantic.dev/docs/ai/project/changelog/
- TechCrunch - https://techcrunch.com/2026/04/15/openai-updates-its-agents-sdk-to-help-enterprises-build-safer-more-capable-agents/

