Built by the team behind Pydantic — the validation library already running inside FastAPI and most other Python AI tools. The core idea: treat an LLM call like a typed function call, not a string-in-string-out black box.
Core concepts
- Typed, validated output: Pass
Agent()anoutput_typeand get back a validated Pydantic model instead of a raw string, with automatic retries if the model returns something that doesn't fit. Pydantic-AI Docs @agent.tooland@agent.tool_plain: Register tools two ways —tool_plainfor stateless functions,toolwhen a tool needsRunContextfor dependency injection.RunContextdependency injection: Pass a database connection, an authenticated user, or an HTTP client into every tool call without reaching for global state.TestModel: A substitute for a real LLM in unit tests — deterministic, fast, offline test runs without burning API credits on every CI run.- Provider-neutral: Supports OpenAI, Anthropic, Google, Groq, Mistral, Bedrock, and any OpenAI-compatible endpoint — swapping providers is usually one string.
Resources
YouTube learning
Core concepts from video walkthroughs:
The 60-second story
A model will happily write you a charming paragraph when you needed a form with three boxes. Pydantic-AI is the counter that says “the answer is this shape, or you try again.” You pass an output_type. You get a validated model. If the text does not fit, the framework can nudge a retry instead of handing your app a string to split with hope and split(",").
Tools come in two coats. @agent.tool_plain is a pure helper: give it inputs, get an output, no secret backpack. @agent.tool is for when the tool needs the backpack — a database handle, the signed-in user, an HTTP client. That backpack is RunContext. You pass it in when the run starts. Tools pull from it. You do not hide the database in a global because globals are socks on the floor and every request wears the wrong pair.
TestModel is a stand-in who never calls a real lab. Tests stay fast, offline, and free. If your CI needs a credit card to prove a function returns a name, the test is not a test. It is a weather report.
The provider string is a swappable sign on the door. OpenAI, Anthropic, Google, Groq, an OpenAI-compatible endpoint. The form and the tools stay. The brain behind the counter can change without a rewrite of the lobby.
Map the counter.
output_typeis the form the answer must match.- A retry is “that was not a form, fill it again.”
tool_plainis a helper with empty pockets.RunContextis the backpack of real dependencies.TestModelis the stand-in who does not spend money.- The provider name is a sign you can change.
from pydantic import BaseModel
class Reply(BaseModel):
answer: str
confident: boolThat class is the contract. The rest of the agent is how you reach an instance of it without scraping prose.
This library exists because agent demos die at the edge where a string becomes a decision. Typed output makes the edge boring. Dependency injection exists so tests can hand the tool a fake backpack. Provider neutrality exists so a pricing change is not a rewrite.
Mental model: a form, a backpack, and a stand-in for tests. The model is not allowed to freestyle the result type. The tools are not allowed to smuggle secrets through globals.
If you can explain why tool and tool_plain both exist, and why TestModel saves the credit card, you understand why people pick this stack when correctness is the feature.

