Skip to main content
WritingSpeakingCodeAboutNow

Your endpoint, as an agent sees it

An agent never reads your API. It reads a flattened projection of it, and the flattening loses things you did not know were losable. Here is how to look at yours.

A customer tells you that their agent keeps getting your API wrong, and the suggestions arrive in the usual order. The model is not very good. It needs a better prompt. Then somebody points out that this is what you get for letting a machine call a production endpoint, which is unhelpful but at least honest.

Then you open the endpoint, and there is nothing wrong with it.

Route::put('/users/{username}', UpdateUser::class)->name('users.update');
final class UpdateUserRequest extends FormRequest
{
public function rules(): array
{
return [
'username' => ['sometimes', 'string', 'max:30'],
'email' => ['sometimes', 'email'],
];
}
}

The path says which user and the body says what to change them to, and nobody has ever been confused by this, because to a person the two username values are obviously different things: one identifies and one replaces.

There is a problem in those eleven lines, and you will not find it by reading them, because it is not in the code at all. It appears during a conversion you never see, somewhere between your specification and the model that is eventually going to call you, and by the time it matters the thing producing it is no longer your API.

I want to walk through that conversion, because most of what it does to your contract is invisible from inside the codebase, and once you have seen it happen to one operation you will not need convincing about the rest.

What a model is actually handed

When someone points an agent at your API, the agent does not receive your API. It receives a tool definition: a name, a description, and one JSON Schema describing a single object of arguments. Anthropic’s shape is three fields.

{
"name": "updateUser",
"description": "Update a user.",
"input_schema": { "type": "object", "properties": {} }
}

OpenAI wraps the same three and calls the schema parameters rather than input_schema. The differences between providers are cosmetic. The shape is not.

Everything a caller will ever know about your operation has to survive being squeezed into those three fields, and a surprising amount of it does not. The losses are structural rather than careless, which is the part worth sitting with: no amount of discipline in your controller prevents any of them, because the discipline is being applied on the wrong side of the conversion.

Four namespaces go in, one comes out

HTTP gives you four independent places to put an input: the path, the query string, the headers and the body. Those are separate namespaces, so a username in the path and a username in the body cannot collide with one another. There is nowhere for them to collide.

A tool takes one object, so the converter flattens all four of those namespaces into it, and the independence has nowhere left to go.

Our route becomes this.

{
"name": "updateUser",
"description": "Update a user.",
"input_schema": {
"type": "object",
"properties": {
"username": { "type": "string" },
"email": { "type": "string" }
},
"required": ["username"]
}
}

That leaves one username carrying two meanings, and a model being asked for a single value to cover an argument that your API is going to read as both the identifier and the replacement.

A converter has to resolve that somehow, and there are only three moves available. It can rename one side to something like body_username, which appears nowhere in your documentation and which the model has therefore never seen. It can nest the entire request body under a body key, which most models do not expect and which quietly changes the shape of every other argument while it is at it. Or it can let one of them win and drop the other without saying so.

All three are wrong in different ways, and which one your customer ends up with depends on whose converter they happened to run.

The fix is in the route rather than the converter.

Route::put('/users/{userId}', UpdateUser::class)->name('updateUser');

Then drop username from the body, or rename it to whatever the field actually means now that it has stopped doing two jobs at once.

The shape worth going looking for is any PUT or PATCH to a resource whose body also carries that resource’s identifier. It turns out to be extremely common once you start grepping for it, and common for the same reason it survived review in the first place, which is that it reads perfectly well.

What is not there at all

Look at that tool definition again, and this time look for what is missing.

There are no status codes, no response schema and no examples of what comes back. There is no Retry-After, no rate limit headers, no auth scopes, no deprecated flag and no tags. Your specification has every one of those things written into it, carefully, and the model receives none of them.

The usual assumption is that this is a converter being lazy, and it is worth correcting, because a tool definition has no field for a response in the first place. Look at the shape once more. Name, description, input schema. There is nowhere for a response schema to go, which means nothing that describes a response can travel through it however diligent the converter is.

Anything you expressed as metadata about a response is lost. Anything you expressed in prose survives.

The description comes through intact, because it is the one free-text field in the whole structure, and that makes it the highest-leverage thing in your entire specification. Which is awkward, given that it is usually the field we autogenerate, or fill in with the operation’s own name spelled slightly differently, or leave until the end and then write in a hurry.

Now follow that to its least comfortable conclusion. You send Deprecation and Sunset headers, with real dates and a link to the migration guide, exactly as RFC 9745 and RFC 8594 describe. A person reading your documentation sees the warning and plans their migration.

An agent sees a tool that works, because there is no deprecation flag in a tool definition and no Sunset header reaching it, so it carries on calling the operation with complete confidence right up until the morning you turn the thing off. Header-based deprecation does not degrade gracefully for this consumer. It does nothing at all, and then it breaks.

So put the sunset date in the description, in words, as well as in the header. It is the cheapest change in this article and the one I would make first.

Three limits you have not had to think about

Names are restricted

OpenAI accepts ^[a-zA-Z0-9_-]{1,64}$. Anthropic accepts the same character class up to 128. Target the intersection: letters, digits, underscore, hyphen, and 64 characters.

Now look again at the route we started with.

->name('users.update');

Dots are not in that character class, so a tool called users.update is rejected rather than sanitised or truncated. Dotted route names are the Laravel convention, which means the framework’s own idiom walks you straight into the problem the moment your specification generator starts using route names as operation IDs.

Keep naming routes with dots, because everything else in Laravel expects them, and map them on the way into the specification instead. If an operation has no ID at all the situation is slightly worse, because every generator will invent one and they will not agree. The name of your tool ends up being a property of whichever converter your customer happened to run rather than anything you decided.

Schemas cannot nest deeply

OpenAI’s strict mode rejects anything past five levels. Well before that ceiling, depth is where tool calls come back structurally wrong, with the right values at the wrong level.

This is worth implementing carefully, because it is easy to implement twice and get two different answers out of it. The providers say five levels of nesting without defining whether the root counts, so every plausible reading is off by one from some other plausible reading. Count the scalar leaf as a level and a payload that is five objects deep reports six, raising a false alarm on something that was fine. Correct for that a little too eagerly and the check quietly stops reporting anything at all.

Pick a definition, write it into the code where someone will read it, and test the boundary rather than the middle.

// A flat body is 0: nothing has to be nested to fill it in.
// {"a": {"b": "string"}} is 1. Five nested objects is 5, which is the limit.
// Only objects and arrays count. A scalar sits inside its parent.
// allOf, oneOf and anyOf add nothing: an allOf of two flat objects is flat.

Composition genuinely does not add a level, which is worth knowing before you go flattening things that were never deep to begin with.

Width costs you on every request

There is no hard limit on argument count. Past roughly fifteen, models start dropping optional arguments they should have sent and confusing ones with similar names, and your min_price, price_min and minimum become a lottery.

The second cost is the one people miss, which is that tool definitions are sent on every request and sit in the context window before the conversation has started. A wide tool is therefore not only harder to call correctly, it is more expensive on every single call your customer makes, including all the ones that never touch it.

Read your own

None of this is worth taking on trust, and it does not need to be. Take the operation you would least like a caller to get wrong, convert it, put the specification away somewhere you cannot see it, and read only what came out the other end.

I wrote an Artisan command to do that, and the interesting part is not the command so much as how little there is to the step that causes all the trouble.

foreach ([...$shared, ...$this->toList($operation['parameters'] ?? null)] as $parameter) {
$parameter = $this->resolve($this->toArray($parameter));
$name = $this->toString($parameter['name'] ?? null);
if ($name === null) {
continue;
}
$this->put(
$properties,
$sources,
$findings,
$name,
$this->resolve($this->toArray($parameter['schema'] ?? ['type' => 'string'])),
$this->toString($parameter['in'] ?? null) ?? 'query',
);
}
foreach ($this->toArray($body['properties'] ?? null) as $name => $schema) {
$this->put($properties, $sources, $findings, (string) $name, $this->resolve($this->toArray($schema)), 'body');
}

Path, query and header parameters go into one array. Body properties go into the same array. That is the whole mechanism, and the collision falls out of put noticing that a name has arrived from two different places.

private function put(array &$properties, array &$sources, array &$findings, string $name, array $schema, string $source): void
{
if (isset($sources[$name]) && $sources[$name] !== $source) {
$findings[] = new Finding(
'collision',
"`{$name}` arrives from both {$sources[$name]} and {$source}.",
);
}
$properties[$name] = $schema;
$sources[$name] = $source;
}

The later source wins, which is deliberate, because that is what a naive converter does and the job here is to show you what happens rather than what ought to.

Run it against a specification that is already in good shape and you get a boring answer, which is the one you want. Run it against something that is not, and you get this.

users.update put /users/{username} 2 args collision, name-charset
recalculate... post /orders/{order} 2 args name-length, depth
get_search get /search 16 args name-missing, width

There is a --strict flag that exits non-zero, so the same command runs in CI and fails the build when somebody adds a dotted operation ID or a body that repeats a path parameter. That is the version actually worth having. A check you run when you remember to run it is a check you have already stopped running.

When your API is too big to fit

The width problem has an upper bound the advice above does not reach. If your API has two hundred operations, you cannot put two hundred tool definitions in front of a model. They would fill the context window before anybody said anything, and you would pay for all of them on every request.

The current answer is to stop sending them all at once. You declare your tools with defer_loading set, alongside a tool search tool, and any given definition is only loaded into context once a search has surfaced it.

{ "type": "tool_search_tool_bm25_20251119", "name": "tool_search_tool_bm25" }

Every other tool in the request gets "defer_loading": true, and two constraints come with that: the search tool itself must not be deferred, and at least one tool has to stay loaded, or the whole request is rejected.

This changes what a large API ought to look like, because descriptions stop being documentation and start being the index that decides whether your operation is considered at all. An operation whose description merely restates its own name used to be a missed opportunity. Under deferred loading it is close to invisible.

What to change, in order

Cheapest first, because these are Monday morning jobs rather than projects.

Give every operation an ID that survives the character class, and stop repeating identifiers between the path and the body. Put anything the caller needs to know about your responses into the description, deprecation dates included, because the structured version of that information is not going to arrive. Then measure your widest and deepest operations and split whichever ones are over.

None of that is new advice. Resource naming, error shapes, payloads a person can hold in their head: we have been telling each other this at conferences for years and getting away with ignoring it, because the consumer on the other end could always work around us. Somebody who cannot tell what an operation does will ask, or will guess and then notice that they guessed.

This consumer does neither. It produces a well-formed, confident, wrong call, and the first sign that anything went wrong is the response it gets back.

Share

XLinkedIn

Related

Keep Reading

All posts →