PRODUCTION AI AGENT STUDIO · LONDON

The demo was the easy part.

We design, build, and run AI agents for production — the retrieval, the evals, the failure handling, the monitoring. The unglamorous eighty percent that decides whether an agent earns its keep or quietly gets switched off.

Two agents live in production today. One reads thousands of EU tenders for Magellan Circle. One runs our own outreach. Both measured. Both still running.

A small number of engagements at a time. Led hands-on.

trace · tender-intel · prodrun 8,412

01 / THESIS

Most agents die between the demo and the deploy.

The gap between the two has a shape. It is an eval harness that tells you the truth about behaviour. Retrieval that holds up on your real corpus, not a curated sample. Tool calls that are typed and validated instead of hoped for. Tracing that answers what happened in minutes, not meetings.

None of it shows up in a demo — and neither does the invoice. An agent that costs pennies per run in a demo can eat its own business case at ten thousand runs a month. Cost per task is a design decision, not a line item you discover later. Closing that gap is the entire job — and it is the only job we take on.

An agent that isn't measured is a liability with good manners.

02 / THE ENGAGEMENT

Scope. Build. Run. In that order, never skipping the first.

Every engagement is bespoke, but the shape is consistent: we prove an agent is worth building before you commit, ship it into production, and keep it healthy once it is live. A small number of engagements at a time, led hands-on.

  1. S·01

    Scope

    A short, paid discovery on your real data. We define the agent, agree the KPI it must move, and prove feasibility — and if an agent is the wrong answer, we tell you so and stop there.

  2. S·02

    Build

    Retrieval, a typed tool layer, an eval harness, and observability — designed, built, and shipped into your stack. It arrives measured, not asserted.

  3. S·03

    Run

    Agents drift; unwatched ones drift silently. We monitor, evaluate, and improve yours in production — and report against the number, every week it is live.

03 / WHAT SHIPS IN EVERY BUILD

The parts you don't see in the demo.

EVAL HARNESS
Regression-tested behaviour on every change. You see the score before your users see the answer.
RETRIEVAL
Hybrid and vector search, indexed and tested against your document corpus — not toy data.
TOOL LAYER
Typed, validated calls into your real systems and APIs. No free-form guessing at your database.
OBSERVABILITY
Every step traced from day one. When something fails, you know what, where, and why.
FAILURE PATHS
Timeouts, retries, fallbacks, human handoff. Failure is designed in advance — never discovered in production.
COST ENGINEERING
Model routing, caching, and token discipline, designed to a cost-per-task budget — then tracked in production like any other KPI.

05 / WHO BUILDS IT

Fifteen years of production. Not fifteen months of prompts.

Agent Foundry Labs is led by Haroon Latif — fifteen years building enterprise SaaS at production scale. Most recently CTO of Airportr Technologies, the travel-tech platform used by British Airways, Swiss, Lufthansa, Virgin Atlantic, and American Airlines across six airports in four countries.

Before that he architected real-time recommendation and data platforms at dunnhumby — the data-science company behind Tesco Clubcard — personalising experiences for more than twenty million shoppers. Engagements are led hands-on, with a vetted network of senior associates brought in as the work needs it.

0 yrsENTERPRISE SAAS IN PRODUCTION
0M+SHOPPERS PERSONALISED IN REAL TIME
0AIRLINES ON THE PLATFORM LED AS CTO
0AGENTS LIVE IN PRODUCTION TODAY

06 / METHOD

We agree the number first. Then we prove we moved it.

  1. FIND

    the workflow costing your team the most.

  2. BASELINE

    measure it as it stands — time, cost, or output.

  3. IMPROVE

    an agent, an automation, or a simpler process — whatever moves it.

  4. MEASURE

    the same KPI again. Success is a number, not a feeling.

Sometimes the honest answer after Scope is that you don't need an agent at all. We've said so before. That candour is cheaper than a build you didn't need — and it's why clients come back when they do.

07 / STRAIGHT ANSWERS

Before you book.

What does Agent Foundry Labs do?

Agent Foundry Labs is a specialist studio that designs, builds, and runs production AI agents for companies. We work as a bespoke engagement — scope, build, then run — rather than selling a product, and we measure every agent against an outcome agreed up front.

How do your engagements work?

Every engagement follows three stages: Scope — a short, paid discovery to define the agent and agree the measure of success; Build — we design, build, evaluate, and ship it into production; and Run — we monitor and improve it once it is live. We take on a small number of engagements at a time.

How is a production AI agent different from a demo?

A demo proves an agent can work once; a production agent works every day. The distance between them is evaluation, tool-use fidelity, retrieval that holds up on real data, and the operational discipline to run it — and closing that distance is the whole job.

Who does the work — is it a big team?

Engagements are led hands-on by Haroon Latif, a CTO who has shipped at production scale, with a vetted network of senior associates brought in as a build needs them. You work directly with the person building your agent, not a rotating junior team.

What happens to the cost at scale?

It gets designed, not discovered. Most of an agent's running cost is set by architecture — which model handles which step, what gets cached, how much context each call carries — so we agree a cost-per-task budget at Scope and build to it, then track it in production like any other KPI. We ran this discipline on our own outreach engine and cut its running cost by roughly 60–70%.

What does an engagement cost?

It depends on the workflow and the stage. Most engagements start with a short, paid Scope, so you can decide on the build with a clear picture and no large up-front commitment. We give you a firm quote once the work is scoped.

GET STARTED

Bring us the workflow that hurts.