FOSS · MIT—Open-source agent-swarm, the operating system for all your AI agents→
How-To/Set Timeout Budgets for Multi-Agent Workflows | desplega.ai

Set Timeout Budgets for Multi-Agent Workflows | desplega.ai

Learn how to configure timeout budgets for long-running multi-agent workflows so tests stay reliable, cost-controlled, and free from silent hangs. If you are also wiring agents into your editor, our guide on setting up MCP in Cursor pairs well with this workflow-level reliability work.

Prerequisites

Map Your Workflow's Critical Paths

Before setting any numbers, list every agent step in your workflow and estimate its expected wall-clock duration. This gives you a realistic baseline so your budgets reflect actual behavior rather than guesswork.

# Example: document your agent timing expectations
# agent: web-scraper       expected: 8s   max-acceptable: 20s
# agent: llm-summarizer    expected: 12s  max-acceptable: 30s
# agent: db-writer         expected: 3s   max-acceptable: 10s
# total workflow budget    expected: 23s  hard-limit: 60s

Define a Global Workflow Timeout

Set a top-level timeout that caps the entire multi-agent run, preventing runaway workflows from consuming infrastructure indefinitely. This is your outermost safety net and should equal the sum of per-agent maximums plus a small coordination buffer.

// desplega.config.ts
export default {
  workflows: {
    'multi-agent-pipeline': {
      timeout: 90_000,        // 90 s global hard limit (ms)
      timeoutMessage: 'Workflow exceeded 90 s budget — failing fast'
    }
  }
};

Assign Per-Agent Step Timeouts

Nest individual timeouts inside each agent step so a single slow agent cannot exhaust the global budget for everyone else. Granular budgets also make failure messages pinpoint which agent is the bottleneck.

// desplega.config.ts — per-step budgets
export default {
  workflows: {
    'multi-agent-pipeline': {
      timeout: 90_000,
      steps: {
        'web-scraper':    { timeout: 20_000 },
        'llm-summarizer': { timeout: 30_000 },
        'db-writer':      { timeout: 10_000 }
      }
    }
  }
};

Configure Retry Budgets Separately from Timeouts

Retries extend elapsed time, so make sure your retry count multiplied by the per-step timeout still fits inside the global budget. Set maxRetries and a backoff strategy explicitly to avoid accidental budget overruns on transient failures.

steps: {
  'llm-summarizer': {
    timeout: 30_000,
    retry: {
      maxAttempts: 2,
      backoff: 'linear',   // 0 s, 5 s — total worst-case: 65 s
      backoffMs: 5_000
    }
  }
}

Add Timeout Assertions in Your Test Suite

Write explicit desplega test assertions that verify the workflow completes within its declared budget during every CI run. This turns timeout compliance into a first-class, automatically enforced contract rather than a manual check.

// tests/multi-agent-pipeline.spec.ts
import { runWorkflow } from '@desplega/sdk';

test('pipeline completes within 90 s budget', async () => {
  const start = Date.now();
  const result = await runWorkflow('multi-agent-pipeline');
  const elapsed = Date.now() - start;

  expect(result.status).toBe('success');
  expect(elapsed).toBeLessThan(90_000);
});

Emit Timeout Telemetry for Observability

Instrument each agent step to emit a structured timing event so you can track p95 durations in your dashboard and tighten budgets over time as the workflow matures. Without telemetry, budgets are set once and never revisited.

// Inside your agent handler
import { desplega } from '@desplega/sdk';

async function runSummarizer(input: string) {
  const span = desplega.startSpan('llm-summarizer');
  try {
    const output = await callLLM(input);
    span.end({ status: 'ok' });
    return output;
  } catch (err) {
    span.end({ status: 'error', error: err });
    throw err;
  }
}

Review and Tighten Budgets After Each Release

After your workflow runs stabilize in production, revisit the timeout values and shrink them toward your observed p95 plus a 20% safety margin. Tighter budgets catch regressions faster and reduce wasted compute on stalled runs.

# Recommended review cadence script (run in CI on schedule)
# Query desplega analytics for p95 step durations, then validate
# budgets are no more than 1.2× the observed p95.
#
# desplega analyze workflow multi-agent-pipeline \
#   --metric p95_duration \
#   --recommend-budgets \
#   --safety-factor 1.2

Why Timeout Budgets Matter

Multi-agent workflows are uniquely vulnerable to silent hangs because each agent can wait on an external API, an LLM provider, or a shared resource lock. Without explicit budgets, a single stalled agent can hold compute for hours while appearing healthy in your logs. Timeout budgets convert that invisible risk into a fast, actionable failure that surfaces the exact step responsible.

Budgets also protect your CI pipeline. A workflow that normally finishes in 30 seconds but occasionally drifts to five minutes will erode developer trust and inflate infrastructure spend. By enforcing a hard ceiling at both the workflow and step level, you guarantee predictable feedback loops and keep cost per run flat even as you add more agents.

Common Pitfalls to Avoid

The most frequent mistake is setting a generous global timeout while leaving individual steps unbounded. This lets one slow agent consume the entire budget, starving downstream steps and producing confusing failure messages. Always pair a global cap with per-step limits so blame is assigned precisely.

Another trap is forgetting that retries multiply elapsed time. A step with a 30-second timeout and three retry attempts can quietly burn 90 seconds plus backoff, blowing past your global budget even though each individual value looked reasonable in isolation. Do the arithmetic up front and encode it in your config so CI catches violations before production does.

Finally, resist the urge to set budgets once and forget them. Agent latency drifts as models change, prompts grow, and data volumes increase. Schedule a recurring review — monthly or after every major release — to compare declared budgets against observed p95 durations and tighten accordingly. For a broader look at agent reliability patterns, see our MCP best practices guide.