Ir para o conteúdo
Análises
Inside

Pydantic AI's You.com Capabilities Bound What Agents Read

Pydantic AI Harness v0.24.0 added YouSearch and YouResearch, wrapping You.com's five APIs as typed capabilities with hard read budgets and citations that travel as metadata. We read the upstream source. Most of the announcement holds up; a few claims do not.

Por Adam Maguire Wilson9 min de leitura
Network cables patched into server hardware with status lights, representing the infrastructure layer behind web-reading agents.
Nesta página

The least interesting way to read Pydantic's announcement on 1 September is as a partnership post. Yes, Pydantic AI Harness now ships two You.com capabilities, YouSearch and YouResearch, and yes, You.com said so too, four days earlier. The more interesting reading is architectural. Web retrieval, historically the least disciplined edge of an agent system, has been turned into a set of typed tools with a constructor-level budget on what enters the context window, and a citation channel the model never touches.

We read the module in the upstream repository rather than taking the walkthrough at face value. Most of it checks out. A few claims do not, and the gaps are as instructive as the features.

Key Takeaways - YouSearch and YouResearch landed in pydantic-ai-harness v0.24.0 on 20 August 2026, wrapping You.com's five APIs as five agent tools. Verified against v0.29.0, the current release. - num_results × max_text_chars is a hard ceiling on how much web text enters context per call, enforced on the request and again on the response. Retrieval volume becomes configuration, not prompt etiquette. - Citations travel in ToolReturn metadata, which the model never sees. Your application renders footnotes from structured data instead of trusting the model to repeat a URL correctly. - Reading the source corrected three details in the announcement: country codes are not validated at construction, the closed-schema rule is enforced by You.com's API rather than the client, and the article's per-run costs are demonstration figures dominated by model tokens, not You.com prices.

Five APIs behind two capabilities

The capabilities shipped in v0.24.0 on 20 August, contributed by a You.com engineer, and sit unchanged through v0.29.0, released 4 September. Install them with pydantic-ai-harness[youdotcom], which pulls the official youdotcom SDK (pinned >=3.1.1,<4; 3.3.0 is current). One list entry, capabilities=[YouSearch(), YouResearch()], wires tools, instructions and defaults into the agent.

YouSearch covers surveying and reading: web_search returns results with excerpts or full pages attached, and get_page fetches clean markdown from a URL. YouResearch covers synthesis: answer for questions one call can settle, research for multi-step passes, and finance_research over a dedicated index of filings, transcripts and fundamentals. Those map one to one onto You.com's five APIs, the product of a company that spent 2023 to 2025 converting itself from a consumer search engine into an API platform. One small sign of how fresh this is: the capabilities index in Pydantic's own docs does not list either capability yet.

A read budget in the constructor

The design decision worth stopping over is num_results × max_text_chars. The defaults are ten results and 10,000 characters of text each, so a stock web_search call can pour 100,000 characters into context. Tighten it to num_results=3, extraction_mode='highlights', max_text_chars=2_000 and the ceiling drops to 6,000 characters per call, whatever the model asks for. The limit is enforced twice, once in the request sent to You.com and again on the response, and truncated pages come back with a marker telling the model it holds part of a document.

This matters because the alternative is what most agents actually run: a raw search endpoint plus instructions begging the model to be economical, or a provider-native search box whose retrieval behaviour you cannot see, let alone cap. A read budget that lives in code is a control. A read budget that lives in a system prompt is a hope. Anyone who has watched an agent cheerfully ingest fifteen full pages to answer a one-line question will recognise the difference immediately.

freshness is validated at construction, accepting day, week, month, year or an explicit date range, and bad values raise before the agent ever runs. Domain scoping follows a deliberate grammar: include_domains is an allowlist that refuses to combine with anything else, while exclude_domains and boost_domains work together. Two instances of the same capability would collide on tool names, so a second, differently-scoped search gets wrapped in core's PrefixTools. One nuance the announcement flattens: the collision fails when you create the agent, not when you construct the capability. The distinction is minor, but it is the difference between an error in your editor and an error at assembly time.

Research that blocks, priced by effort

research does not hand back a job ID to poll. It blocks for the length of the pass, which is why timeout_ms defaults to ten minutes (600,000 milliseconds in the source, with a comment noting that deep and exhaustive passes routinely take minutes). Effort runs lite, standard, deep, exhaustive. You.com's frontier tier exists but only runs as a background job, so the capability simply does not offer it. Refusing to wrap an API mode you cannot operate honestly is the kind of omission we would like to see more of.

Structured output comes with two sharp edges, and they are enforced in different places. Pairing an output_schema with lite effort raises at construction, in the client. But the requirement that the schema close itself to extra keys, additionalProperties: false, appears nowhere in the module. We checked. It is a You.com API constraint, which means the client will happily build the capability and let the model discover the rejection several minutes into a research call, exactly as the walkthrough warns. Pydantic's post presents both rules as things "the capability catches". Only one of them is.

finance_research is the honest exception to the shared control surface: it takes its input and a finance_effort of deep or exhaustive, nothing else. A domain filter you configured will not narrow it. That is not a bug, but it is the sort of asymmetry that belongs in a runbook before someone builds a compliance claim on top of it.

Citations the model never touches

Every tool returns a ToolReturn. The model sees the text plus an appended Sources: block. Your application reads the same sources as structured records from metadata the model never sees:

Because footnotes come from metadata['sources'] rather than from generated prose, a rendered citation cannot be corrupted by the model paraphrasing a URL into a 404. web_search also drops You.com's search_uuid and latency into metadata, which is precisely what you want when a run went sideways and you need to ask the vendor about one specific query. This is the control-layer instinct applied to evidence: keep the receipt on a channel the reasoning engine cannot edit.

The stronger claim sits on the Answer API itself. You.com says every citation is verified against the source text before the answer returns, and that the excerpts in each citation are the verbatim passages used to ground it, reporting 93.48% accuracy on SimpleQA at a median 2.67 seconds when the API launched on 5 August. Treat that as self-asserted. The verification mechanism and its failure rate are unpublished, the figure is You.com grading its own homework, and You.com's own docs advise following each source URL yourself for legal, financial or medical use. A Baseten case study adds a detail worth knowing: the answer pipeline runs on open-weight models chosen for a cost-accuracy frontier, so the benchmark number attaches to a pipeline whose underlying model can change. None of this makes the verification worthless. It makes it one input to your own checks, which is what the metadata design quietly assumes.

An error taxonomy that reads like policy

The failure handling is three lines of frozenset and a decorator, and it encodes a real operational position:

Rate limits, rejected parameters, server errors and network blips become a ModelRetry the model can work around. Authentication, billing and permission failures abort the run, because no amount of rephrasing fixes your invoice. An empty result set is explicitly not an error: the tool returns No results found for {query!r}. and lets the model reword. A missing API key fails at agent assembly with a UserError, not mid-run.

Two things are absent, and both matter at production scale. There is no client-side rate limiting or backoff anywhere in the module; retries are delegated to the model loop. And there is no spend ceiling. The harness library's cost budgets, tool budgets and approval workflows are all still open pull requests at the time of writing. Domain filters, meanwhile, shape what the agent can see but are not a security boundary for what a retrieved page can tell it, a caveat Pydantic's post states plainly and we will happily second.

What the demonstration costs actually measure

The walkthrough's headline experiment runs a lean agent and a thorough agent on the same silver-price question and reports 7,937 versus 65,987 input tokens, $0.01 versus $0.13 per run, with the deep research call consuming 62% of the thorough run's wall time. The honest reading is that these are five runs inside seven minutes on one question, dominated by Anthropic token spend, and Pydantic says as much. The paired ratio, 5.4x to 11.8x with a median of 8.2x, measures the configuration, not the market.

What those figures are not is You.com pricing. The published schedule is per thousand calls: web search at $5, answers at $5, research from $12, and finance research from $110 to $500 depending on effort. At those rates the lean run's single search call costs about half a cent before the model says a word. The real cost driver is still the model reading what retrieval returns, which is, of course, the entire argument for bounding the reads.

The layer this actually moves

Strip the partnership framing away and what remains is retrieval joining the harness as a first-class, typed capability rather than an improvised tool. You.com is the second provider wired this way after Exa, which suggests the pattern, not the vendor, is the story: survey, read and research as separate tools with separate budgets, citations delivered out of band, and failures sorted by who can fix them. It is refreshingly boring engineering, and agent systems could use more of it.

What it does not yet deliver is the other half of the control layer. A capability that can spend money and read the open web, in a library whose budget and approval machinery is still in pull requests, leaves the operator assembling the guardrails themselves. The harness is on 0.x releases and says so, so the API you pin today may shift under a minor version. The question worth carrying into your next design review is not whether to give the agent web access. It is which of these bounds you would have had to build yourself, and which ones, this month, you still do.

Continue lendo

Agent Field Notes

Receba a próxima edição.

Harnesses de agentes, ambientes de execução, segurança e governança, explicados para quem precisa operar esses sistemas.

Enfrentando uma decisão como esta?

Realizamos revisões de arquitetura, avaliações de governança e comparações de frameworks com versões fixadas para equipes que tomam decisões importantes sobre sistemas de agentes.

Sobre o autor

Adam Maguire Wilson

Fundador e consultor independente em sistemas de agentes de IA.

adam.mw