PRODUCTION AI AGENT STUDIO · LONDON
The demo was the easy part.
We design, build, and run AI agents for production — the retrieval, the evals, the failure handling, the monitoring. The unglamorous eighty percent that decides whether an agent earns its keep or quietly gets switched off.
Two agents live in production today. One reads thousands of EU tenders for Magellan Circle. One runs our own outreach. Both measured. Both still running.
A small number of engagements at a time. Led hands-on.
01 / THESIS
Most agents die between the demo and the deploy.
The gap between the two has a shape. It is an eval harness that tells you the truth about behaviour. Retrieval that holds up on your real corpus, not a curated sample. Tool calls that are typed and validated instead of hoped for. Tracing that answers what happened in minutes, not meetings.
None of it shows up in a demo — and neither does the invoice. An agent that costs pennies per run in a demo can eat its own business case at ten thousand runs a month. Cost per task is a design decision, not a line item you discover later. Closing that gap is the entire job — and it is the only job we take on.
An agent that isn't measured is a liability with good manners.
02 / THE ENGAGEMENT
Scope. Build. Run. In that order, never skipping the first.
Every engagement is bespoke, but the shape is consistent: we prove an agent is worth building before you commit, ship it into production, and keep it healthy once it is live. A small number of engagements at a time, led hands-on.
- S·01
Scope
A short, paid discovery on your real data. We define the agent, agree the KPI it must move, and prove feasibility — and if an agent is the wrong answer, we tell you so and stop there.
- S·02
Build
Retrieval, a typed tool layer, an eval harness, and observability — designed, built, and shipped into your stack. It arrives measured, not asserted.
- S·03
Run
Agents drift; unwatched ones drift silently. We monitor, evaluate, and improve yours in production — and report against the number, every week it is live.
03 / WHAT SHIPS IN EVERY BUILD
The parts you don't see in the demo.
- EVAL HARNESS
- Regression-tested behaviour on every change. You see the score before your users see the answer.
- RETRIEVAL
- Hybrid and vector search, indexed and tested against your document corpus — not toy data.
- TOOL LAYER
- Typed, validated calls into your real systems and APIs. No free-form guessing at your database.
- OBSERVABILITY
- Every step traced from day one. When something fails, you know what, where, and why.
- FAILURE PATHS
- Timeouts, retries, fallbacks, human handoff. Failure is designed in advance — never discovered in production.
- COST ENGINEERING
- Model routing, caching, and token discipline, designed to a cost-per-task budget — then tracked in production like any other KPI.
04 / PROOF
Two agents. Two domains. One standard.
Different problems, same discipline — governed, measured, deployed. Proof the craft travels, not just the domain.
Tender intelligence for Magellan Circle
Ingests EU tenders from multiple public sources, retrieves against the client's own material, extracts structured criteria, and drafts fully-cited first-pass bid responses. Every claim cited — the agent never infers.
A compliance-first outreach engine
From one declarative customer profile it researches prospects, scores fit, and drafts governed, on-brand outreach — read-only and daily-capped by design, with hard compliance rules enforced on every message.
05 / WHO BUILDS IT
Fifteen years of production. Not fifteen months of prompts.
Agent Foundry Labs is led by Haroon Latif — fifteen years building enterprise SaaS at production scale. Most recently CTO of Airportr Technologies, the travel-tech platform used by British Airways, Swiss, Lufthansa, Virgin Atlantic, and American Airlines across six airports in four countries.
Before that he architected real-time recommendation and data platforms at dunnhumby — the data-science company behind Tesco Clubcard — personalising experiences for more than twenty million shoppers. Engagements are led hands-on, with a vetted network of senior associates brought in as the work needs it.
06 / METHOD
We agree the number first. Then we prove we moved it.
- FIND
the workflow costing your team the most.
- BASELINE
measure it as it stands — time, cost, or output.
- IMPROVE
an agent, an automation, or a simpler process — whatever moves it.
- MEASURE
the same KPI again. Success is a number, not a feeling.
Sometimes the honest answer after Scope is that you don't need an agent at all. We've said so before. That candour is cheaper than a build you didn't need — and it's why clients come back when they do.
07 / STRAIGHT ANSWERS
Before you book.
What does Agent Foundry Labs do?
Agent Foundry Labs is a specialist studio that designs, builds, and runs production AI agents for companies. We work as a bespoke engagement — scope, build, then run — rather than selling a product, and we measure every agent against an outcome agreed up front.
How do your engagements work?
Every engagement follows three stages: Scope — a short, paid discovery to define the agent and agree the measure of success; Build — we design, build, evaluate, and ship it into production; and Run — we monitor and improve it once it is live. We take on a small number of engagements at a time.
How is a production AI agent different from a demo?
A demo proves an agent can work once; a production agent works every day. The distance between them is evaluation, tool-use fidelity, retrieval that holds up on real data, and the operational discipline to run it — and closing that distance is the whole job.
Who does the work — is it a big team?
Engagements are led hands-on by Haroon Latif, a CTO who has shipped at production scale, with a vetted network of senior associates brought in as a build needs them. You work directly with the person building your agent, not a rotating junior team.
What happens to the cost at scale?
It gets designed, not discovered. Most of an agent's running cost is set by architecture — which model handles which step, what gets cached, how much context each call carries — so we agree a cost-per-task budget at Scope and build to it, then track it in production like any other KPI. We ran this discipline on our own outreach engine and cut its running cost by roughly 60–70%.
What does an engagement cost?
It depends on the workflow and the stage. Most engagements start with a short, paid Scope, so you can decide on the build with a clear picture and no large up-front commitment. We give you a firm quote once the work is scoped.
GET STARTED
Bring us the workflow that hurts.
Prefer email? [email protected]