Monitor, retry and recover deliveries
Operate the Webhook Relay Outbound Gateway: delivery history, retrying one delivery, recovering failed deliveries, replaying skipped messages, endpoint health, payload Functions, limits and MCP tools.
The Outbound Gateway records every message, delivery and attempt for 35 days. You can see what a customer's endpoint answered, retry one delivery, re-send everything that failed during an outage, or replay messages skipped while an endpoint was paused. None of this requires publishing the event again, and every copy keeps the original webhook-id. Everything on this page is also available in the dashboard's Outbound webhooks page and through MCP.
Inspect delivery history
| Question | Call |
|---|---|
| What did we publish? | GET /v1/outbound/messages?consumer=&event_type=&limit=&offset= (newest first, without payloads) |
| Where did one message go? | GET /v1/outbound/messages/{id}: the payload plus one delivery per endpoint, with attempts |
| What did one endpoint receive? | GET /v1/outbound/endpoints/{id}/deliveries |
Filter an endpoint's deliveries with status (sent, failed, stalled, received or rejected), event_type or message_id, and page with limit (up to 100) and offset. Each attempt records its status code, duration, response headers (up to 8 KiB) and the start of the response body (4 KiB per attempt; the latest response is kept up to 120 KiB). Custom endpoint headers are left out of the recorded request because they can hold credentials.
| Delivery status | Meaning |
|---|---|
received | Queued; the first attempt hasn't been made yet. |
stalled | An attempt failed and a retry is scheduled. |
sent | The endpoint answered 2xx. |
failed | Retries ran out, the endpoint answered 410, or the endpoint's Function failed. |
rejected | Skipped because the endpoint was paused or disabled, or dropped by a Function. |
curl --user "$RELAY_KEY:$RELAY_SECRET" \
"https://my.webhookrelay.com/v1/outbound/endpoints/6f0c2c0e-5a3b-4f43-9d0a-4c4f7d7a2b11/deliveries?status=failed&limit=20"
Retry one delivery
POST /v1/outbound/endpoints/{id}/retry with a message_id re-sends that message to that endpoint in the background and returns a recovery task. If the delivery is waiting for its next scheduled retry, it is sent now. A delivery whose first attempt is still queued is refused with 409.
curl --user "$RELAY_KEY:$RELAY_SECRET" \
-H 'Content-Type: application/json' \
-X POST https://my.webhookrelay.com/v1/outbound/endpoints/6f0c2c0e-5a3b-4f43-9d0a-4c4f7d7a2b11/retry \
-d '{ "message_id": "0199cf5e-7a3c-7d2e-9b1a-3f4e5d6c7b8a" }'
Recover failed deliveries and replay missed ones
Two recovery tasks walk an endpoint's messages from since (RFC 3339, defaulting to 24 hours ago, at most 35 days back) until now:
- Recover (
POST /v1/outbound/endpoints/{id}/recover) re-sends deliveries that ran out of retries, typically after the customer fixes an outage. - Replay missing (
POST /v1/outbound/endpoints/{id}/replay-missing) sends deliveries that were skipped while the endpoint was paused or disabled. Resume the endpoint first.
curl --user "$RELAY_KEY:$RELAY_SECRET" \
-H 'Content-Type: application/json' \
-X POST https://my.webhookrelay.com/v1/outbound/endpoints/6f0c2c0e-5a3b-4f43-9d0a-4c4f7d7a2b11/recover \
-d '{ "since": "2026-10-08T00:00:00Z" }'
Each call answers 202 with a task. Follow its progress with GET /v1/outbound/recovery-tasks/{id}. status moves from pending to running to completed (or failed with an error), and processed counts the deliveries handled so far. A re-sent delivery starts a fresh retry schedule and keeps its earlier attempts in the history. An account can have up to 10 active recovery tasks.
The SDK equivalents:
const task = await relay.outbound.endpoints.recover(endpointId, { since: "2026-10-08T00:00:00Z" });
await relay.outbound.endpoints.resume(endpointId);
await relay.outbound.endpoints.replayMissing(endpointId);
await relay.outbound.endpoints.retry(endpointId, messageId);
const progress = await relay.outbound.recoveryTasks.get(task.id);
task, err := api.RecoverOutboundDeliveries(endpointID, time.Now().Add(-48*time.Hour))
_, err = api.ResumeOutboundEndpoint(endpointID)
_, err = api.ReplayMissingOutboundDeliveries(endpointID, time.Time{}) // zero time: last 24 hours
_, err = api.RetryOutboundDelivery(endpointID, messageID)
progress, err := api.GetOutboundRecoveryTask(task.ID)
Watch endpoint health
With many customers, some endpoints will always be failing. The gateway tracks this on each endpoint: state, consecutive_failures, failing_since and stats (attempts and failures over the last 24 hours). Failing endpoints never open incidents or send email to your account.
GET /v1/outbound/healthcounts endpoints by state and reports the account's attempts and failures over the last 24 hours (updated within about 15 seconds).GET /v1/outbound/endpoints?state=failinglists failing endpoints across all consumers, longest failing first. Addconsumerto narrow it down.
const { endpoints, stats } = await relay.outbound.health();
const failing = await relay.outbound.endpoints.listAll({ state: "failing", limit: 20 });
The dashboard shows the same data on the Endpoint health tab. Use POST /v1/outbound/endpoints/{id}/pause to stop deliveries to an endpoint your customer is migrating, and .../resume to start them again. See endpoint states for how endpoints move between active, failing, paused and disabled.
Transform or drop payloads with Functions
Attach a Function to an endpoint with function_id to tailor what one customer receives. The Function gets the published JSON payload as r.body, and it can:
- replace the body with
r.setBody(...). The result must be valid JSON up to 256 KiB. - drop the message with
r.stopForwarding(). The delivery is recorded asrejected, "Message dropped by Function".
const invoice = JSON.parse(r.body)
if (invoice.amount === 0) {
r.stopForwarding()
} else {
r.setBody(JSON.stringify({ id: invoice.invoice_id, total: invoice.amount / 100, currency: "USD" }))
}
Outbound Functions can change only the body. The method, URL and headers belong to the endpoint, so a Function that changes them fails the delivery with an explanatory error. The Function runs once per message and endpoint, and its result is reused for every retry and replay. The Function attached when a message is published is the one that runs for it. If a worker crashes before the result is saved, the Function can run again, so avoid external side effects.
Limits
| Limit | Value |
|---|---|
| Payload | JSON, 256 KiB |
| Consumers | 1,000 per account |
| Event types | 1,000 per account; 100 per endpoint (or *) |
| Endpoints | 100 per consumer, one per URL |
| Custom headers | 32 per endpoint |
Consumer IDs, event type names, idempotency keys, event_id | Up to 128 characters |
| Rate | 50 deliveries per second per endpoint by default, up to 1,000; up to 16 in flight at once (or the rate, if lower) |
| Attempt timeout | 15 seconds by default, up to 60 |
| Retry window | Up to 8 attempts within 48 hours |
| Messages awaiting preparation | 10,000 per account (429 beyond that) |
| Active recovery tasks | 10 per account |
| Recorded response | Headers up to 8 KiB per attempt; body 4 KiB per attempt, 120 KiB for the latest response |
| History | 35 days |
MCP tools and the dashboard agent
Every REST operation has an MCP tool with the same name. The dashboard agent uses the same tools, so you can ask it things like "which endpoints are failing?" or "recover customer_42's failed deliveries since Monday". Tools follow your role and any outbound token restriction, and configuration, secret and recovery changes are audited the same way as REST calls.
| Area | Tools |
|---|---|
| Access | get_outbound_access, request_outbound_access |
| Consumers | list_outbound_consumers, get_outbound_consumer, upsert_outbound_consumer, delete_outbound_consumer |
| Event types | list_outbound_event_types, save_outbound_event_type, delete_outbound_event_type |
| Endpoints | list_outbound_endpoints, get_outbound_endpoint, create_outbound_endpoint, update_outbound_endpoint, delete_outbound_endpoint, pause_outbound_endpoint, resume_outbound_endpoint, rotate_outbound_endpoint_secret, reveal_outbound_endpoint_secret, get_outbound_health |
| Messages | publish_outbound_message, list_outbound_messages, get_outbound_message |
| Deliveries and recovery | list_outbound_deliveries, retry_outbound_delivery, recover_outbound_deliveries, replay_missing_outbound_deliveries, get_outbound_recovery_task |
Frequently asked questions
How do I resend outbound webhooks after a customer's outage?
Once the endpoint is reachable again, call POST /v1/outbound/endpoints/{id}/recover with an optional since timestamp (the default is the last 24 hours, the maximum is 35 days). Webhook Relay re-sends that endpoint's deliveries that ran out of retries, in the background, with the original webhook-id. If the endpoint was paused or disabled instead, resume it and call replay-missing.
How long is outbound delivery history kept?
Messages, deliveries with their attempts, and recovery tasks are kept for 35 days. Retries, recovery and replay work within that window.
Can I change the payload for one customer?
Yes. Attach a Function to that customer's endpoint with function_id. The Function receives the published JSON payload and can return a different JSON body (up to 256 KiB) or drop the message. It runs once per message and endpoint, and the result is reused for every retry.
Can an AI agent manage outbound webhooks?
Yes. Every Outbound Gateway REST operation has an MCP tool with the same name, such as publish_outbound_message, list_outbound_deliveries or recover_outbound_deliveries. The tools are available to MCP clients and to the agent in the Webhook Relay dashboard, under the same permissions and audit log as the API.
