Benjamin Muskalla

Advising

Between the lab
and the team.

I build AI products for a living, in the gap between what research makes possible and what engineers ship. At poolside I built the managed agents team and product from zero, and run the layer that connects research with customers. Before that, early engineer on GitHub Copilot, from research prototype to market leader. On the side, I help a few founders get from demo to product, and help their senior engineers work out which habits still hold now that agents write the code.

Where I can help · the product

  • ◇ Agent & ML architecture — the loop from customer usage to data, training, evals and back into production, so a better model ships as a routine release, not a heroic one. Getting the few calls that are expensive to reverse right, and keeping everything else cheap to change.
  • ◇ Evals & telemetry — an eval suite the team trusts enough to ship on, and production signal that turns into a roadmap, not a dashboard nobody reads.
  • ◇ Inference cost & latency — knowing what a request really costs and where to spend it. Before the bill or the p95 starts setting your priorities.
  • ◇ Research ↔ product — one roadmap for researchers and product engineers, and a shared definition of better. So a new checkpoint turns into something users feel.

Where I can help · the team

  • ◇ Your first engineering hires — who to hire first, what to build yourself, what to just buy. I've hired for and led teams from a 20-person startup to a frontier lab.
  • ◇ From a few people coding to a team — clear ownership, design reviews, a backlog you can actually triage. The habits that keep you fast once you're more than three people.
  • ◇ Senior engineers in the agent era — which practices were fundamentals and which were workarounds for expensive typing. What to rebuild when review, not implementation, is the bottleneck.
  • ◇ Remote and distributed — leading teams across Europe and North America since 2010. What works async, what needs a room.
  • ◇ What breaks at each size — early at Tasktop as it went from 20 to 150 people, now running the layer between research and customers at poolside. Which decisions you can defer, and which ones you can't.

How this works

  • ◇ On the side — I have a day job, so it's one or two companies at a time. If I'm full, I'll say so and point you to someone.
  • ◇ Async by default — a shared channel, a design doc, a PR to look at. And a call or a room when that's what the problem needs: a hard decision, a tricky hire, a talk for your team.
  • ◇ Investors — if you fund AI companies and want a second technical opinion on one, the same applies.

Right now

2024–now
Nov 2024 to Present · 1 yr 11 mos
Technical Advisor Dosu

Technical advisor to the founders of Dosu, the AI teammate for engineering teams and open-source projects.

  • Agent and retrieval architecture as usage grows
  • Evals and telemetry: production agent behavior into product and reliability decisions

Before all this

Twenty-plus years of engineering, most of it in open source, from the Eclipse platform onward. Java Champion, principal engineer at Tasktop as it grew from 20 to 150 people, technical reviewer of Build a Large Language Model (From Scratch). First company at 18, so I remember what the founder side feels like. The long version is on the vitae page.

Say hi

Tell me what you're building and where it hurts. Worst case, you leave with a straight answer.