Revolut · AI-Ops
Revolut AIOps Platform
A single source of truth for model governance and operational safety — rebuilt from the ground up in six months, and adopted by every team shipping AI at Revolut.
- Role
- Senior Product Designer
- Timeline
- 6 months
- Platform
- AI-Ops Platform
- Company
- Revolut

The model playground — define a model, wire up tools and variables, and simulate it, all in one place.
Context
Revolut needed a single source of truth for model governance and operational safety. Over six months we rebuilt the AI-Ops platform from the ground up — fixing the information architecture, introducing a split model definition and a model playground, standardising fallback behaviour and online evaluation, and adding traces and triage tools.
It became the primary model-governance product used by every team — enabling safer releases, faster incident resolution, and consistent observability across every model-driven surface. Stakeholders sat at the top of the company: the CEO office, CEO (Nick), CTO (Vlad), the Head of AI, the Head of Product, and the data-science team.
Problem statements
We rebuilt the product to do three things: unify governance, make model behaviour predictable, and let product teams ship models with confidence.
A fragmented surface
The old AI-Ops UI was hard to use. Teams struggled to find features and to actually operationalise their models.
Testing happened elsewhere
Teams tested models with ad-hoc external tools — inconsistent quality checks and risky rollouts.
Failures with no safety net
Hallucinations, timeouts, and format errors caused user friction and long incident MTTR. There was no unified fallback, observability, or evaluation.
No governance
Versioning, routing, and rollback had no clear process — releases across dozens of teams were risky.

A single model incident, traced span-by-span — the observability teams simply didn’t have before.
Primary goals
One place to govern
A single, discoverable platform for model management and governance.
Fewer failures
Reduce user-facing model failures with standardised fallback strategies.
Safe releases
Ship safely with online evaluation, shadowing, and canaries.
Faster recovery
Cut MTTR for model incidents with traces, samples, and a triage UI.
Faster iteration
Let product teams move quickly with a playground and split model definitions.
My role, team & process
I led product design end-to-end: discovery, IA, interaction design, prototypes, and the ops dashboards. The cross-functional team spanned PM, ML engineers, infra engineers, data scientists, QA, and ops — and executive sponsors were closely involved, which sped up decisions and alignment.
We ran a rapid six-month delivery in iterative releases: discovery (2–3 weeks), IA and prototyping (4 weeks), implementation collaboration (12–14 weeks), then rollout and stabilisation across the rest.
Discovery & key insights
Methods
- Stakeholder interviews — product owners, Head of AI, PMs, ML and data scientists
- Logs analysis of past incidents and support escalations
- Shadowing ops and support during live incidents
- Usability testing of the old platform
- Platform benchmarking
Key insights
- Users repeatedly couldn’t find features — IA was the single biggest blocker to adoption
- Teams tested models outside the platform because definition and testing were disconnected
- Ops needed fast access to example failures (sampled traces) and side-by-side fallback comparison
- Users didn’t understand how versioning or presets worked
- The model component was too basic to be useful, which held back adoption
What we shipped
We rebuilt the platform around three pillars.
Information architecture & UX
Reorganised navigation by user job — Govern → Test → Monitor → Release — not by technical component. Discoverability went up; the learning curve for new teams came down.
Split model definition & playground
Separated model definition (metadata, prompts, I/O schema, scoring) from execution and testing. The playground runs scenarios, swaps tools and providers, tweaks variables, and compares outputs — without touching production.
Operational features
Standardised fallback strategies, online evaluation with shadow traffic and ground-truth capture, traces and sampled responses for triage, and model routing and versioning (canary, traffic split, rollback).
IA redesign
Users looked for workflows — test a model, set a fallback, review an incident — but the old UI forced them to hunt through technical menus. We reorganised the whole structure around those jobs.
LLM model flow

The end-to-end model flow that shaped the new structure.
Designed solution
Everything below is finalised design, currently live in production. Where an approach was tried and dropped, I’ve kept the “what didn’t work” — the iterations mattered as much as the outcome.
Split model & playground
Splitting model definition from execution let teams shape a model — prompts, tools, variables, schema, scoring — and test it in the same UI, without risking production.
The first cut was too rigid
My first version modelled everything as fixed tiers. It looked tidy, but the moment teams started using it the rigidity got in the way — every change fought the structure. I split definition from execution and replaced the tiers with a flexible playground (the “First cut” frames above).
Traces & sessions
For triage, ops needed to see exactly what a model did — every span, tool call, and token step — and to pull sampled failures fast.
Traces grouped by the wrong thing
The first iteration grouped every trace by agent — click a model, see its traces. After launch we found that a single session often uses several models to get the best result, so “by model” hid the real story. We split the IA: traces and sessions became separate views.
Model fallback
A simple tier builder — Primary → Secondary → Deterministic — with clear precedence, so a failing model hands off predictably. Teams drag to reorder a chain, and compare the outputs of fallback models against the same input, side by side.
Online evaluation
A guided wizard with sensible defaults — sample size, minimum confidence, KPI targets — and templates for common evaluations. Models can be shadowed against live traffic with ground-truth capture before they take real load, and the report aggregates every metric and drills down to the trace level, so a regression is easy to explain.
The first evaluation surface was hard to read
The earlier version buried the answer in dense tables with little narrative — you couldn’t tell at a glance what regressed, or why. We rebuilt it around a clear report, with the trace-level data one click behind it.
Versioning & publish
Every model carries a version history with aliases, an activity timeline, and one-click rollback. Releases roll out as canaries with traffic split, so a bad version never hits everyone at once. This was one of the hardest surfaces to get right — and one I’m most proud of.
Versions and presets blurred together
The first versioning model made it hard to tell a version from a preset — exactly the wrong thing to be unsure about before a release. I reworked it into a clear history with aliases, an activity timeline, and one-click rollback.
Results & impact
One platform, adopted across the company.
65 teams migrated
Every team moved onto the single governance product.
Faster incident resolution
Traces, samples, and a triage UI cut the time to find and fix model incidents.
Safer releases
Canary and shadowing caught multiple production regressions before users did — and made rollbacks quick.
Higher developer velocity
Product teams reported shorter time to validate and ship models.
Organisational confidence
Central governance made executives comfortable expanding AI across the company.
“I’d love to take your questions.”


















































