Benjamin Muskalla
← all posts

Blog

OpenAPI specs as RL gyms

aiagentsreinforcement learningevaluation

We turned OpenAPI specs into mock services for RL training and improved a tool-use benchmark by about 28%; most of the work went into turning a simple seed spec into a world complex enough to be worth learning from.

At the beginning of 2026, we pushed hard to use our own models internally, at that time Laguna XS, a mixture-of-experts model with 33B parameters, 3B active per token, for work on recursive self-improvement (RSI), on training, and on incident management. A lot of the failures we saw had one thing in common: the agent misused an external API whose documentation told it exactly how to call it. It also missed call patterns any engineer reaches for without thinking, such as passing a pagination token back to fetch the next page.

Those failures sent me on a side quest. Could an OpenAPI spec become a reinforcement learning (RL) environment, a place where an agent practices on an unfamiliar API and a program checks whether it got the job done?

The mistakes worth practicing on are small and quiet. An agent matches a record by its display name when two records share it. It stops after the first page of results when the answer sits on a later one. Every request is valid, and every response looks reasonable. The failure lives in how the agent connects the evidence, and a program can check that, if someone builds a world for it to check.

That is the kind of task I wanted to build: a small, executable world in which an agent has to discover, reason, and act, and where a program decides whether it succeeded. An OpenAPI specification already describes the operations and the shapes of the data. Could we generate the rest?

On paper, the supply is enormous. Here is a back-of-the-napkin estimate for the APIs.guru directory alone, in round numbers that are illustrative, not measured:

Napkin math: 2,500 API specs times 10 seeded worlds times 10 tasks times 8 interfaces gives 2 million environments

Combining specs multiplies it again, and the range is wide. At one end, a world chains two everyday APIs: a Slack alert that links to a GitHub pull request, with a task that needs both. At the other, it lives in a specialist domain like the NCBI Datasets API, where a task might look up a species’ reference genome and fetch its gene annotations. A 2024 crawl of GitHub found about 660,000 OpenAPI artifacts, so the long tail is long.

Specs are not the bottleneck. Worthwhile tasks are: complex enough to teach the agent something, and trustworthy enough that a pass means what it says.

One agent builds the world, another works in it

The basic loop is two agents that never talk to each other. An authoring agent gets an OpenAPI specification and writes a scenario, a mock service, seed data, a task instruction, a reference solution, and a verifier. A fresh solver agent gets only the running service and the instruction:

The authoring agent turns an OpenAPI spec into a generated world with a mock service, a task and a verifier; the solver agent sees only the service and the task, and the verifier turns its answer and final state into a reward

That shape serves evaluation as it stands and reinforcement learning once you put a training loop around it. The agent acts, observes responses, and receives a reward grounded in an executable check. Its explanation can sound convincing or awkward; the verifier only looks at whether the work was done.

The specification supplies less of this than it seems to. The following table lists what an environment needs and where each piece comes from.

Ingredient What it contributes
API specification Operations, parameters, and data shapes
Seeded world Concrete records and relationships to investigate
Service behavior State changes, errors, pagination, and other consequences of actions
Task A reason to use the API
Verifier An explicit definition of success
Reset mechanism A fresh starting point for each attempt

Only the first row comes straight from the spec. Service behavior is grounded too: the API’s documentation says how pagination, errors, and state changes work, and the authoring agent’s domain knowledge of the API fills in what the docs leave out. The rest is the author’s call. A spec rarely captures every business rule, and a schema reference does not guarantee a meaningful relationship between records. What you get is a mock grounded in a specification and its documentation. Whether it behaves like the real service is a separate claim that needs its own validation.

A spec can train far more than single lookups. More links and more pages mostly add cost; the variations worth generating change what the agent has to do:

  • Fan out: inspect related records before combining their results.
  • Establish absence: find the candidate that lacks a property, which takes a complete search.
  • Recover: recognize a transient failure and retry.
  • Track progress: submit an operation and poll until its result is ready.
  • Reconcile state: make the world satisfy a condition while preserving unrelated data.

Same world, different interfaces

One extension I find promising holds the environment and verifier fixed and changes the interface. The same world could be exposed as raw HTTP, as Model Context Protocol (MCP) tools, as a generated software development kit (SDK), as a command-line client, or as agent skills: compact Markdown docs generated from the spec with openapi-to-skills. That would show how the interface affects an agent’s ability to discover operations, compose requests, recover from errors, and finish the work.

One OpenAPI spec produces one fixed world (mock service, seed data, verifier) that agents can reach through five interfaces: raw HTTP, MCP tools, a generated SDK, a CLI client and agent skills

Sharing a backend doesn’t make those experiences equivalent, though. An SDK that fetches every page automatically removes a challenge that raw HTTP keeps. A tool description can reveal what another agent has to discover. A fair comparison has to control for both capabilities and information.

Chained specs add one more skill worth training: a chat platform and a code host might name the same person with unrelated identifiers, and a good task makes the solver establish that identity from evidence. Chained specs can also mix interfaces, which is closer to how agents meet APIs in practice: one service arrives as a CLI, the other only as raw HTTP, and the data has to bridge from one to the other.

Test the verifier like code

Anyone who has built an RL environment knows the floor: the reference solution passes, doing nothing fails, and both run from a fresh start. Generated environments need more, because a passing reference solution is a witness, not a proof.

Each environment runs as a Harbor task with the mock server in a sidecar, so the solver can call the API but can’t peek at the seed data or the verifier.

The most useful habit was to treat the verifier as code under test: before a task ships, it has to pass the reference solution and fail every shortcut, from doing nothing to a look-alike record with the same name. A fresh solver agent that sometimes succeeds from the public instruction alone catches the rest: an author, service, and verifier that share the same misunderstanding.

What the training run showed

We took the idea through to an RL training experiment. Starting from the same earlier RL checkpoint as the baseline arm, we continued training on approximately 2,000 generated tasks spanning 658 OpenAPI specifications, using REINFORCE with importance weighting and a leave-one-out baseline over groups of 16 task attempts. Agents accessed mock HTTP services through a terminal harness and received verifier-based rewards alongside harness checks. At 300 additional training steps, the OpenAPI-trained model scored 20.4% pass@1 on ToolathlonVerified versus 16.0% for the baseline. The reported paired comparison covered 107 evaluation samples and showed a gain of 4.44 percentage points—about 28% relative—with a 95% confidence interval of +0.23 to +8.64 points.

The question that mattered was whether the model improved outside the generated worlds. The comparison report gives these results on ToolathlonVerified, the variant of the Toolathlon tool-use benchmark we evaluated on. Pass@1 is the success rate on a single attempt.

Measurement at the evaluated checkpoint Reported result
Baseline pass@1 16.0%
OpenAPI training run pass@1 20.4%
Paired improvement, absolute +4.44 percentage points
Improvement, relative to baseline about +28%
Paired 95% confidence interval +0.23 to +8.64 percentage points

The gain showed up on a separate evaluation, not only on the tasks we trained on. My working hypothesis is that practicing in generated API worlds improves tool use more broadly, and this result is consistent with it.

Cheap to generate, so generate for your gaps

Because worlds are cheap to generate, the question shifts from how many to which ones. Start from your model’s gaps and point the generator at them: single-API tasks or chained specs, the interfaces where your agents stumble, and the domains your users work in. Trust is the part that stays expensive: every world still has to pass its verifier checks, and a gain should show up on APIs the model has never seen.

An API spec gives you a vocabulary of actions and the skeleton of a world. The interesting work is giving that world a purpose, making its consequences reliable, and checking that success means what you think it means.

Cite this post

@misc{muskalla2026openapigyms,
  author = {Benjamin Muskalla},
  title  = {OpenAPI Specs as RL Gyms},
  year   = {2026},
  month  = sep,
  url    = {https://bmuskalla.dev/blog/2026-09-26-openapi-specs-as-rl-gyms/}
}

Related posts

Thoughts?

If this post sparked a question or a disagreement, I'd like to hear it.