OpenAI's own production agent framework, built around three primitives: Agent, Runner, and function_tool. The same primitives power AgentKit's visual builder underneath — so skills you learn in code transfer when you later compose flows in a UI.
Think of the SDK as opinionated glue: you describe who the agent is (instructions + model), what it can do (tools and optional handoff targets), and the Runner owns the loop — call the model, execute requested tools, feed results back, repeat until the model returns a final answer or hits a guardrail.
Core concepts
- Agent & Runner: An
Agentholds the model, system instructions, tools, and optionalhandoffsto other agents. You never manually alternate “model → tool → model” in application code — you pass the user message toRunner.run_sync()orawait Runner.run(), and the runner executes the full tool-calling loop with consistent error handling. This is the mental model most other Python agent frameworks copy, but OpenAI's version is tightly integrated with OpenAI models, tracing, and hosted tools. Agents SDK Docs · Read more: OpenAI Agents guide @function_tooldecorator: Turns a plain Python function into a callable tool by reading its type hints and docstring — no separate JSON schema file to keep in sync. The model sees a structured tool definition derived from your function signature; when it “calls” the tool, the SDK invokes your Python code and returns the result to the model. Good docstrings on parameters directly improve tool-selection accuracy.- Handoffs: Agents can delegate the conversation to a specialized sub-agent via
handoffs=[]— for example, a triage agent that routes billing vs. technical questions without stuffing every tool into one prompt. Handoffs are OpenAI's built-in answer to multi-agent systems while staying in one SDK (compare to explicit supervisor graphs in LangGraph). - Guardrails: Built-in input/output guardrails can halt a run before an unsafe or off-policy action executes — useful when agents have write access (send email, charge card, delete data). Treat guardrails as product requirements, not an afterthought.
- Tracing: Every model call, tool call, and handoff is automatically recorded and viewable in a tracing dashboard — no manual OpenTelemetry wiring for the default path. Use traces to debug “why did it call that tool?” and to build eval datasets from production failures.
- Sessions & context: Multi-turn agents need explicit session or memory strategy — the SDK can carry conversation state across turns; pair with your own compaction or vector memory when threads get long (see the Memory chapter in this roadmap).
- When to pick this SDK: Best default if you are on OpenAI models, want minimal boilerplate for tool loops, and plan to use OpenAI-hosted tooling (file search, web, etc.) or AgentKit later. For strict typed outputs as first-class citizens, also evaluate Pydantic-AI; for graph-shaped control flow, see LangGraph.
Resources
YouTube learning
Core concepts from video walkthroughs:
The 60-second story
You hire a receptionist and you do not want to stand between them and the phone all day whispering “now ask the calendar, now tell the guest.” You write who they are, what they may touch, and you let a runner handle the loop. That runner is the point of the OpenAI Agents SDK.
An Agent is the job description: model, instructions, tools, and optional coworkers they may hand the guest to. Runner.run is the shift. It calls the model, runs the tool the model asked for, feeds the result back, and repeats until there is a final answer or a guardrail pulls the plug. You stop hand-coding “model, tool, model” in your web handler. That loop is where people forget an error path.
@function_tool turns a normal Python function into something the model can request. The type hints and the docstring become the menu. A vague docstring is a vague menu, and the model will order the wrong thing with confidence. Handoffs are how a front desk sends billing to billing instead of growing forty tools and a headache. Guardrails are the manager who can stop a send-email or a refund before it leaves the building. Tracing is the camera footage of the shift: which tool, which handoff, why.
Sessions are the notepad for the next guest in the same conversation. When the notepad gets huge, that is a memory problem, not a reason to paste the whole year into the next call.
Map the lobby.
Agentis the job description.Runneris the person working the loop.@function_toolis a menu item generated from your function.- A handoff is sending the guest to a specialist.
- A guardrail is a stop before a dangerous action.
- A trace is the footage.
- A session is the notepad for this conversation.
from agents import Agent, Runner, function_tool
@function_tool
def office_hours() -> str:
"""Return the front desk hours."""
return "9 to 5"
agent = Agent(name="Desk", instructions="Be brief.", tools=[office_hours])You still run it with the Runner. The agent does not dial itself just because you constructed it.
This SDK exists so a production loop is not a weekend script you are afraid to touch. It is opinionated glue around OpenAI models, tracing, and hosted tools. If you need a graph you can pause for a human, look at LangGraph. If the thing you worship is a validated Pydantic object, look at Pydantic-AI. If you are already in this ecosystem, the Runner is the least theatrical way to get a tool loop you can debug.
Mental model: a job description plus a runner who works the phones. You describe the person. You do not impersonate the loop.
You can explain the page if you separate “who the agent is” from “who executes the turn,” and if you treat a docstring as part of the product.


