Retry a step and stop on a final error
Some failures go away: the payroll system is down for a minute. Others do not: the employee has no vacation days left, and asking again only delays the answer. In this recipe a step retries the first kind and stops at once on the second.
Save the two files as retry-and-fail.workflow.ts and
retry-and-fail.workflow.case.ts.
The host deploys only a workflow file’s default export, so one file holds one
workflow; the test imports the same definition by name.
import * as v from '@identity-flow/sdk/valibot';import { NonRetryableError, defineBinding, defineWorkflow } from '@identity-flow/sdk';
export interface Payroll { book(employeeId: string, days: number): Promise<{ readonly bookingId: string }>;}
export const Payroll = defineBinding('payroll', (): Payroll => { throw new Error('Replace this factory with one that returns your payroll client');});
/** * `limit` counts attempts, not extra ones: three in total. `backoff` doubles * `delay` per attempt already made, so IdentityFlow waits four seconds after the * first failure and eight after the second, then gives up. * * Two conditions make that useful: the step has to be safe to run twice, and its * name has to be stable, or the second attempt looks like different work in the * event history. * * A refusal the other side will repeat is not worth a second attempt. Rethrowing * it as a `NonRetryableError` ends the instance on the step that failed, with * the original error kept as the cause — no retry is scheduled at all. */export const bookVacation = defineWorkflow( { name: 'book-vacation', version: '1.0.0', draft: false, schema: v.object({ employeeId: v.string(), days: v.number() }), }, async (flow) => { const payroll = flow.use(Payroll);
return await flow.do( 'book vacation days', { retries: { limit: 3, delay: '2 seconds', backoff: 'exponential' } }, async () => { try { return await payroll.book(flow.params.employeeId, flow.params.days); } catch (error) { if (isInsufficientBalance(error)) { throw new NonRetryableError('Vacation balance is exhausted', { cause: error }); } throw error; } }, ); },);
// The host deploys only the default export, so one file holds one workflow. The// test imports the named export; both are the same definition.export default bookVacation;
function isInsufficientBalance(error: unknown): boolean { return error instanceof Error && error.message.includes('INSUFFICIENT_BALANCE');}Two parts of this file are placeholders. The Payroll factory throws until you
return your real payroll client from it. isInsufficientBalance matches on the
error message only because the example has no real error type; test for your
system’s status or error code instead.
limit counts total attempts, not extra ones: limit: 3 means the step runs at
most three times, so at most two of them are retries.
backoff: 'exponential' doubles delay per attempt already made, which is one
step further along than it reads. With delay: '2 seconds' the first retry is
four seconds out and the second eight — not two and four. Work the first wait
out before you pick the number, because it is also the floor for how long a test
of this step has to be allowed to run. The other forms of retries and the
defaults are listed in Errors and retries.
A step also accepts a timeout option, but 0.3.0 does not enforce it: an
attempt runs until the callback returns or throws. To bound a call, give your
client its own timeout, for example by passing
AbortSignal.any([signal, AbortSignal.timeout(30_000)]), where signal is the
step’s signal.
NonRetryableError separates the two cases. Throwing it ends the step as
ERRORED without another attempt and keeps the original error as cause, so
the event history still shows what the other system returned. The workflow does
not catch the error, so the instance ends as ERRORED as well. Other errors
that stop retries are listed under
Errors that stop retries.
import { workflowTest } from '@identity-flow/testing';import { expect, vi } from 'vitest';
import { Payroll, bookVacation } from './retry-and-fail.workflow';
const test = workflowTest({ workflow: bookVacation, accounts: { operator: { providerId: 'acme', principals: ['role:payroll-operator'] } },});
test('retries a failure that can go away', async ({ flow, accounts, fixture }) => { const book = vi .fn<Payroll['book']>() .mockRejectedValueOnce(new Error('PAYROLL_UNAVAILABLE')) .mockResolvedValue({ bookingId: 'bk-4711' }); fixture(Payroll, { book });
const instance = await flow.actAs(accounts.operator).start({ employeeId: 'e-2041', days: 5 });
// The second attempt is four seconds out, so the wait has to outlast it. const completed = await flow.waitFor(instance, { timeout: 30_000 }).toBeCompleted();
expect(completed.data).toEqual({ bookingId: 'bk-4711' }); expect(book).toHaveBeenCalledTimes(2);});
test('stops at a refusal that will not', async ({ flow, accounts, fixture }) => { const book = vi.fn<Payroll['book']>().mockRejectedValue(new Error('INSUFFICIENT_BALANCE')); fixture(Payroll, { book });
const instance = await flow.actAs(accounts.operator).start({ employeeId: 'e-2041', days: 90 });
await expect.poll(async () => (await flow.observe(instance)).instance.status).toBe('ERRORED');
const observed = await flow.observe(instance); expect(observed.steps).toContainEqual( expect.objectContaining({ name: 'book vacation days', status: 'ERRORED' }), ); // A NonRetryableError schedules no retry, so the first attempt is the last. expect(book).toHaveBeenCalledOnce();});fixture(Payroll, …) replaces the binding with a test implementation, so the
production factory is never called. The first test shows the retry by counting
calls. It gives waitFor a timeout of 30 seconds, because the default of one
second is shorter than the four-second wait before the second attempt. The
second test shows that a NonRetryableError schedules no retry: one call, then
ERRORED.
Run the test with the Workflow Test distribution, as described in Before you start. The first test takes at least four seconds.
Next: Errors and retries for why an external call can still arrive twice, and Workflow API for the exact shapes.