Revolut · AI Platform

LLM model fallback

A safety net for AI-powered chats: when the main model fails or slows down, traffic switches to a backup in under a second — and every switch is logged.

Role
Senior Product Designer
Timeline
1 week
Platform
AI Ops Platform
Company
Revolut

Impact

0%

Workflow uptime

up from 73.84%

−0%

Recovery time

10 → 0.8 min

0h→1m

To set up a fallback

0→95%

Team adoption

Simulator comparing a base model and two fallback models answering the same question side by side

One prompt, three models side by side — the base model and two backups answering the same question, so you can see how each one holds up.

Context

When an AI model failed or slowed down, the whole chat flow could break — broken conversations, timeouts, and a spike in support tickets. Failovers that weren’t thought through also pushed inference costs up.

It made teams cautious, too. Rolling back a bad model was slow and manual, so people held off on shipping model upgrades at all.

Model catalog table listing vendors and models with price, speed, context window, tools support and reasoning

The model catalog — every model compared by price, speed, context window, and reasoning, so picking a backup is a deliberate choice.

The problem

There was no reliable way to keep AI flows running when the main model went down — and no clean record of what happened when it did.

Teams needed a backup plan they could set once and trust: one that switched automatically, stayed within cost and safety limits, and left a clear trail for audit.

A failing model shouldn’t mean a broken conversation.

Prompt configuration with a base model and a single fallback model set up

The starting point: a base model with a single backup, before the fallback chain is built out.

My contributions

01

Strategy & alignment

Ran workshops with ML, SRE, product, and ops to agree on what mattered — the safety rules, the metrics to watch, and when a fallback should kick in.

02

Service & failure mapping

Mapped the whole request path and the ways it breaks — timeouts, wrong answers, quality drops, cost spikes — and set the signals that trigger a switch.

03

Interaction & UI

Designed the fallback dashboard: set the order of backup models, choose what triggers a switch, and add guardrails.

04

Research & validation

10 stakeholder interviews and 6 usability sessions, plus a two-week pilot on live traffic to check the triggers and copy held up.

05

Benchmarking

Looked at how other AI platforms handle failover and safety, then borrowed the good defaults and fixed the gaps.

06

Handoff & rollout

Shipped specs and copy, worked with engineers on logging and audit trails, and rolled out in phases — pilot, then a few teams, then everyone.

Research & validation

Before shipping, I checked how teams were coping today and whether the new flow actually held up under real use.

Ten interviews across ML, SRE, product, and support; six usability sessions on the prototype; and a two-week pilot on two live flows — plus a quick look at how other AI platforms handle failover, to borrow the sensible defaults.

Set up a fallback in minutes

Pick a base model, add a couple of backups in order, and set what counts as a failure. No pull requests, no approvals — it’s a few clicks.

Start empty
01Start emptya blank prompt, ready for its first model
Pick your models
02Pick your modelsthe catalog, filtered to what a flow needs
Set the order
03Set the orderbase model first, then backups by priority

Reorder the chain by dragging

The order is the priority — top to bottom. Drag any model to change which backup runs first.

Open the menu
01Open the menuevery model has a Reorder action
Reorder mode
02Reorder modeeach model gets a drag handle
Drag to move
03Drag to movegrab a model and slide it up or down
Confirm the order
04Confirm the orderthe new priority, ready to save

Compare every model on the same question

Run the base model and each backup against the same input, side by side — an easy way to see how a fallback would actually answer before you rely on it.

One shared setup
01One shared setupmodels, tools, and variables in a single place
Choose what to compare
02Choose what to comparetick the models to run against each other
Answers side by side
03Answers side by sideevery model responds to the same prompt
Focus mode
04Focus modecollapse the config to scan responses full-width

Outcomes & impact

Fewer user-facing incidents

Traffic routed around failing models instead of breaking the chat — so customers hit fewer broken conversations and timeouts.

Faster recovery

On-call teams spent far less time manually diagnosing and fixing model failures.

Lower support cost

Fewer AI-related incidents meant fewer support tickets and manual fixes.

Safer, faster iteration

With an auditable, reversible safety net, product and ML teams shipped model upgrades far more often.

Higher AI adoption

As reliability improved, more product teams turned on AI-driven flows in production.

Before / after

KPIBeforeAfterChange
AI-workflow uptime73.84%99.0%+25 pp
Visible AI failures · per 10k7540−47%
Time-to-recovery (median)10 min0.8 min−92%
Session completion79%93%+14 pp
Support tickets · per week12050−58%
Time to set up a fallback72 hr1 min−99.9%
Adoption · teams using it20%95%+75 pp

From internal pilot metrics across two product teams over the first three months.

It moved Revolut from fragile AI deployments to reliable, auditable ones — fewer customer-facing failures, less manual firefighting, and the confidence to ship model upgrades faster.

Why it mattered