OpenAI Agents SDK

OpenAI's production agent framework: Agent, Runner, and function_tool.

OpenAI's own production agent framework, built around three primitives: Agent, Runner, and function_tool. The same primitives power AgentKit's visual builder underneath — so skills you learn in code transfer when you later compose flows in a UI.

Think of the SDK as opinionated glue: you describe who the agent is (instructions + model), what it can do (tools and optional handoff targets), and the Runner owns the loop — call the model, execute requested tools, feed results back, repeat until the model returns a final answer or hits a guardrail.

Core concepts

  1. Agent & Runner: An Agent holds the model, system instructions, tools, and optional handoffs to other agents. You never manually alternate “model → tool → model” in application code — you pass the user message to Runner.run_sync() or await Runner.run(), and the runner executes the full tool-calling loop with consistent error handling. This is the mental model most other Python agent frameworks copy, but OpenAI's version is tightly integrated with OpenAI models, tracing, and hosted tools. Agents SDK Docs · Read more: OpenAI Agents guide
  2. @function_tool decorator: Turns a plain Python function into a callable tool by reading its type hints and docstring — no separate JSON schema file to keep in sync. The model sees a structured tool definition derived from your function signature; when it “calls” the tool, the SDK invokes your Python code and returns the result to the model. Good docstrings on parameters directly improve tool-selection accuracy.
  3. Handoffs: Agents can delegate the conversation to a specialized sub-agent via handoffs=[] — for example, a triage agent that routes billing vs. technical questions without stuffing every tool into one prompt. Handoffs are OpenAI's built-in answer to multi-agent systems while staying in one SDK (compare to explicit supervisor graphs in LangGraph).
  4. Guardrails: Built-in input/output guardrails can halt a run before an unsafe or off-policy action executes — useful when agents have write access (send email, charge card, delete data). Treat guardrails as product requirements, not an afterthought.
  5. Tracing: Every model call, tool call, and handoff is automatically recorded and viewable in a tracing dashboard — no manual OpenTelemetry wiring for the default path. Use traces to debug “why did it call that tool?” and to build eval datasets from production failures.
  6. Sessions & context: Multi-turn agents need explicit session or memory strategy — the SDK can carry conversation state across turns; pair with your own compaction or vector memory when threads get long (see the Memory chapter in this roadmap).
  7. When to pick this SDK: Best default if you are on OpenAI models, want minimal boilerplate for tool loops, and plan to use OpenAI-hosted tooling (file search, web, etc.) or AgentKit later. For strict typed outputs as first-class citizens, also evaluate Pydantic-AI; for graph-shaped control flow, see LangGraph.

Resources

YouTube learning

Core concepts from video walkthroughs:

The 60-second story

You hire a receptionist and you do not want to stand between them and the phone all day whispering “now ask the calendar, now tell the guest.” You write who they are, what they may touch, and you let a runner handle the loop. That runner is the point of the OpenAI Agents SDK.

An Agent is the job description: model, instructions, tools, and optional coworkers they may hand the guest to. Runner.run is the shift. It calls the model, runs the tool the model asked for, feeds the result back, and repeats until there is a final answer or a guardrail pulls the plug. You stop hand-coding “model, tool, model” in your web handler. That loop is where people forget an error path.

@function_tool turns a normal Python function into something the model can request. The type hints and the docstring become the menu. A vague docstring is a vague menu, and the model will order the wrong thing with confidence. Handoffs are how a front desk sends billing to billing instead of growing forty tools and a headache. Guardrails are the manager who can stop a send-email or a refund before it leaves the building. Tracing is the camera footage of the shift: which tool, which handoff, why.

Sessions are the notepad for the next guest in the same conversation. When the notepad gets huge, that is a memory problem, not a reason to paste the whole year into the next call.

Map the lobby.

  • Agent is the job description.
  • Runner is the person working the loop.
  • @function_tool is a menu item generated from your function.
  • A handoff is sending the guest to a specialist.
  • A guardrail is a stop before a dangerous action.
  • A trace is the footage.
  • A session is the notepad for this conversation.
python
from agents import Agent, Runner, function_tool

@function_tool
def office_hours() -> str:
    """Return the front desk hours."""
    return "9 to 5"

agent = Agent(name="Desk", instructions="Be brief.", tools=[office_hours])

You still run it with the Runner. The agent does not dial itself just because you constructed it.

This SDK exists so a production loop is not a weekend script you are afraid to touch. It is opinionated glue around OpenAI models, tracing, and hosted tools. If you need a graph you can pause for a human, look at LangGraph. If the thing you worship is a validated Pydantic object, look at Pydantic-AI. If you are already in this ecosystem, the Runner is the least theatrical way to get a tool loop you can debug.

Mental model: a job description plus a runner who works the phones. You describe the person. You do not impersonate the loop.

You can explain the page if you separate “who the agent is” from “who executes the turn,” and if you treat a docstring as part of the product.

Threads in the same domain as this chapter — go deeper on iHateReading without leaving the roadmap.

Python AI agent SDKs in 2026: OpenAI vs Pydantic-AI vs smolagents vs Strands cover

Python AI agent SDKs in 2026: OpenAI vs Pydantic-AI vs smolagents vs Strands

Same three-tool agent in four frameworks — pick by execution model and testing needs.

Agent harnessing: the part of building AI agents nobody talks about cover

Agent harnessing: the part of building AI agents nobody talks about

Loop, tools, memory, sandbox, and tracing live in your code — not the model weights.

OpenAI DevDay 2026 for builders cover

OpenAI DevDay 2026 for builders

Dots distribution, Sign in with ChatGPT, and the trigger–decision–action loop.

Smallest AI Agent cover

Smallest AI Agent

Minimal file-reading agent to learn LLM tool loops without framework magic.