DocumentationMonitor, retry and recover deliveries

Monitor, retry and recover deliveries

Operate the Webhook Relay Outbound Gateway: delivery history, retrying one delivery, recovering failed deliveries, replaying skipped messages, endpoint health, payload Functions, limits and MCP tools.

The Outbound Gateway records every message, delivery and attempt for 35 days. You can see what a customer's endpoint answered, retry one delivery, re-send everything that failed during an outage, or replay messages skipped while an endpoint was paused. None of this requires publishing the event again, and every copy keeps the original webhook-id. Everything on this page is also available in the dashboard's Outbound webhooks page and through MCP.

Inspect delivery history

QuestionCall
What did we publish?GET /v1/outbound/messages?consumer=&event_type=&limit=&offset= (newest first, without payloads)
Where did one message go?GET /v1/outbound/messages/{id}: the payload plus one delivery per endpoint, with attempts
What did one endpoint receive?GET /v1/outbound/endpoints/{id}/deliveries

Filter an endpoint's deliveries with status (sent, failed, stalled, received or rejected), event_type or message_id, and page with limit (up to 100) and offset. Each attempt records its status code, duration, response headers (up to 8 KiB) and the start of the response body (4 KiB per attempt; the latest response is kept up to 120 KiB). Custom endpoint headers are left out of the recorded request because they can hold credentials.

Delivery statusMeaning
receivedQueued; the first attempt hasn't been made yet.
stalledAn attempt failed and a retry is scheduled.
sentThe endpoint answered 2xx.
failedRetries ran out, the endpoint answered 410, or the endpoint's Function failed.
rejectedSkipped because the endpoint was paused or disabled, or dropped by a Function.
curl --user "$RELAY_KEY:$RELAY_SECRET" \
  "https://my.webhookrelay.com/v1/outbound/endpoints/6f0c2c0e-5a3b-4f43-9d0a-4c4f7d7a2b11/deliveries?status=failed&limit=20"

Retry one delivery

POST /v1/outbound/endpoints/{id}/retry with a message_id re-sends that message to that endpoint in the background and returns a recovery task. If the delivery is waiting for its next scheduled retry, it is sent now. A delivery whose first attempt is still queued is refused with 409.

curl --user "$RELAY_KEY:$RELAY_SECRET" \
  -H 'Content-Type: application/json' \
  -X POST https://my.webhookrelay.com/v1/outbound/endpoints/6f0c2c0e-5a3b-4f43-9d0a-4c4f7d7a2b11/retry \
  -d '{ "message_id": "0199cf5e-7a3c-7d2e-9b1a-3f4e5d6c7b8a" }'

Recover failed deliveries and replay missed ones

Two recovery tasks walk an endpoint's messages from since (RFC 3339, defaulting to 24 hours ago, at most 35 days back) until now:

  • Recover (POST /v1/outbound/endpoints/{id}/recover) re-sends deliveries that ran out of retries, typically after the customer fixes an outage.
  • Replay missing (POST /v1/outbound/endpoints/{id}/replay-missing) sends deliveries that were skipped while the endpoint was paused or disabled. Resume the endpoint first.
curl --user "$RELAY_KEY:$RELAY_SECRET" \
  -H 'Content-Type: application/json' \
  -X POST https://my.webhookrelay.com/v1/outbound/endpoints/6f0c2c0e-5a3b-4f43-9d0a-4c4f7d7a2b11/recover \
  -d '{ "since": "2026-10-08T00:00:00Z" }'

Each call answers 202 with a task. Follow its progress with GET /v1/outbound/recovery-tasks/{id}. status moves from pending to running to completed (or failed with an error), and processed counts the deliveries handled so far. A re-sent delivery starts a fresh retry schedule and keeps its earlier attempts in the history. An account can have up to 10 active recovery tasks.

The SDK equivalents:

const task = await relay.outbound.endpoints.recover(endpointId, { since: "2026-10-08T00:00:00Z" });
await relay.outbound.endpoints.resume(endpointId);
await relay.outbound.endpoints.replayMissing(endpointId);
await relay.outbound.endpoints.retry(endpointId, messageId);
const progress = await relay.outbound.recoveryTasks.get(task.id);
task, err := api.RecoverOutboundDeliveries(endpointID, time.Now().Add(-48*time.Hour))
_, err = api.ResumeOutboundEndpoint(endpointID)
_, err = api.ReplayMissingOutboundDeliveries(endpointID, time.Time{}) // zero time: last 24 hours
_, err = api.RetryOutboundDelivery(endpointID, messageID)
progress, err := api.GetOutboundRecoveryTask(task.ID)

Watch endpoint health

With many customers, some endpoints will always be failing. The gateway tracks this on each endpoint: state, consecutive_failures, failing_since and stats (attempts and failures over the last 24 hours). Failing endpoints never open incidents or send email to your account.

  • GET /v1/outbound/health counts endpoints by state and reports the account's attempts and failures over the last 24 hours (updated within about 15 seconds).
  • GET /v1/outbound/endpoints?state=failing lists failing endpoints across all consumers, longest failing first. Add consumer to narrow it down.
const { endpoints, stats } = await relay.outbound.health();
const failing = await relay.outbound.endpoints.listAll({ state: "failing", limit: 20 });

The dashboard shows the same data on the Endpoint health tab. Use POST /v1/outbound/endpoints/{id}/pause to stop deliveries to an endpoint your customer is migrating, and .../resume to start them again. See endpoint states for how endpoints move between active, failing, paused and disabled.

Transform or drop payloads with Functions

Attach a Function to an endpoint with function_id to tailor what one customer receives. The Function gets the published JSON payload as r.body, and it can:

  • replace the body with r.setBody(...). The result must be valid JSON up to 256 KiB.
  • drop the message with r.stopForwarding(). The delivery is recorded as rejected, "Message dropped by Function".
const invoice = JSON.parse(r.body)

if (invoice.amount === 0) {
  r.stopForwarding()
} else {
  r.setBody(JSON.stringify({ id: invoice.invoice_id, total: invoice.amount / 100, currency: "USD" }))
}

Outbound Functions can change only the body. The method, URL and headers belong to the endpoint, so a Function that changes them fails the delivery with an explanatory error. The Function runs once per message and endpoint, and its result is reused for every retry and replay. The Function attached when a message is published is the one that runs for it. If a worker crashes before the result is saved, the Function can run again, so avoid external side effects.

Limits

LimitValue
PayloadJSON, 256 KiB
Consumers1,000 per account
Event types1,000 per account; 100 per endpoint (or *)
Endpoints100 per consumer, one per URL
Custom headers32 per endpoint
Consumer IDs, event type names, idempotency keys, event_idUp to 128 characters
Rate50 deliveries per second per endpoint by default, up to 1,000; up to 16 in flight at once (or the rate, if lower)
Attempt timeout15 seconds by default, up to 60
Retry windowUp to 8 attempts within 48 hours
Messages awaiting preparation10,000 per account (429 beyond that)
Active recovery tasks10 per account
Recorded responseHeaders up to 8 KiB per attempt; body 4 KiB per attempt, 120 KiB for the latest response
History35 days

MCP tools and the dashboard agent

Every REST operation has an MCP tool with the same name. The dashboard agent uses the same tools, so you can ask it things like "which endpoints are failing?" or "recover customer_42's failed deliveries since Monday". Tools follow your role and any outbound token restriction, and configuration, secret and recovery changes are audited the same way as REST calls.

AreaTools
Accessget_outbound_access, request_outbound_access
Consumerslist_outbound_consumers, get_outbound_consumer, upsert_outbound_consumer, delete_outbound_consumer
Event typeslist_outbound_event_types, save_outbound_event_type, delete_outbound_event_type
Endpointslist_outbound_endpoints, get_outbound_endpoint, create_outbound_endpoint, update_outbound_endpoint, delete_outbound_endpoint, pause_outbound_endpoint, resume_outbound_endpoint, rotate_outbound_endpoint_secret, reveal_outbound_endpoint_secret, get_outbound_health
Messagespublish_outbound_message, list_outbound_messages, get_outbound_message
Deliveries and recoverylist_outbound_deliveries, retry_outbound_delivery, recover_outbound_deliveries, replay_missing_outbound_deliveries, get_outbound_recovery_task

Frequently asked questions

How do I resend outbound webhooks after a customer's outage?

Once the endpoint is reachable again, call POST /v1/outbound/endpoints/{id}/recover with an optional since timestamp (the default is the last 24 hours, the maximum is 35 days). Webhook Relay re-sends that endpoint's deliveries that ran out of retries, in the background, with the original webhook-id. If the endpoint was paused or disabled instead, resume it and call replay-missing.

How long is outbound delivery history kept?

Messages, deliveries with their attempts, and recovery tasks are kept for 35 days. Retries, recovery and replay work within that window.

Can I change the payload for one customer?

Yes. Attach a Function to that customer's endpoint with function_id. The Function receives the published JSON payload and can return a different JSON body (up to 256 KiB) or drop the message. It runs once per message and endpoint, and the result is reused for every retry.

Can an AI agent manage outbound webhooks?

Yes. Every Outbound Gateway REST operation has an MCP tool with the same name, such as publish_outbound_message, list_outbound_deliveries or recover_outbound_deliveries. The tools are available to MCP clients and to the agent in the Webhook Relay dashboard, under the same permissions and audit log as the API.

Did this page help you?