SIP retransmissions and retransmission storms
Short version: retransmissions are not the fault. They are UDP's transaction layer doing its job while something else is broken. What matters is which message is repeating and what is failing to come back — those two facts identify the fault almost on their own.
What counts as a retransmission
A retransmission is the same message, sent again by the same transaction. Concretely:
identical branch parameter on the top Via, identical CSeq,
identical request line. Three things get mistaken for it:
- A re-INVITE — same Call-ID and same dialog, but a new
CSeqnumber and a new branch. That is a fresh transaction asking for a change, not a repeat. - A fork — same request going to several contacts, each with its own branch.
- A loop — the same request coming back around the path with
Max-Forwardsdecremented. If the counter is falling, you have a routing loop, and it will end in 483 Too Many Hops rather than a timeout.
The timers
RFC 3261 defaults: T1 = 500 ms (the assumed round trip), T2 = 4 s (the ceiling for most retransmit intervals), T4 = 5 s.
| What is repeating | Interval | Gives up after |
|---|---|---|
| INVITE request (Timer A) | T1, doubling — no cap | Timer B, 64 × T1 = 32 s |
| Non-INVITE request (Timer E) | T1, doubling, capped at T2 | Timer F, 64 × T1 = 32 s |
| Final non-2xx response (Timer G) | T1, doubling, capped at T2 | Timer H, 64 × T1, or an ACK |
| 2xx response to an INVITE | T1, doubling, capped at T2 | 64 × T1, or an ACK |
Two consequences fall straight out of that table. Over a reliable transport such as TCP there are no transaction-layer retransmissions at all — but the 32-second timeout still applies. And ordinary provisional responses are never retransmitted, so a lost 180 Ringing is simply lost; only reliable provisionals (RFC 3262, the ones that demand a PRACK) get repeated.
Five patterns and what each one means
INVITE repeating, nothing coming back
The far end never saw it, or cannot answer. Six retransmissions and a local 408 — see 408 Request Timeout for the full list of reasons, of which oversized fragmented UDP is the one people miss most often.
INVITE repeating, then a late 100 Trying
Not a loss problem: the far end is slow. Something in its path — a database lookup, an ENUM query, a DNS resolution it does synchronously — is taking longer than 500 ms, so it has not yet produced a provisional. The retransmissions stop the moment the 100 arrives, and the call usually completes. Worth fixing, rarely worth alarming about.
200 OK repeating, no ACK
The most diagnostic pattern in SIP, and always worth stopping for. The callee answered
and cannot get the acknowledgement back, so it repeats the 200 OK for 64 × T1 and then
tears down the call with a BYE — roughly 32 seconds of ringback-then-nothing that
users report as "it answers and then hangs up". The ACK for a 2xx is a new transaction
addressed to the Contact URI from the 200 OK, routed through the Record-Route
set, so the usual causes are a Contact containing a private or otherwise
unroutable address, a Record-Route set that a B2BUA mangled, or a firewall permitting
the INVITE path but not the ACK's.
A final 4xx/5xx/6xx repeating
Same idea, opposite direction: the server transaction is retransmitting its final response because no ACK arrived. Note that the ACK for a non-2xx is generated by the transaction layer and carries the same branch as the INVITE — unlike the 2xx case — so if you cannot find one, look for the branch rather than the dialog.
Everything on the box repeating at once
When unrelated calls all start retransmitting inside the same few seconds, the call you were asked about is a bystander. The tell is correlation in time rather than in dialog: an interface flap, a CPU or queue overload, a downstream peer that went away, or a DNS server that stopped answering. Fixing the "call" in this situation achieves nothing, and the storm can become self-sustaining, because every retransmission adds load to the box that is already too busy to answer.
What to check in the trace
- Group by Via branch, not by Call-ID. That single grouping separates retransmissions from re-INVITEs and forks without any guesswork.
- Check the intervals against the table above. A ladder that does not double is not RFC 3261's transaction layer — it is an application-level retry, and those come from somewhere much more interesting.
- Establish what should have answered and did not: a provisional, a final response, or an ACK. That is the actual fault.
- Measure message size and look for IP fragmentation on anything approaching the path MTU.
- Widen the window to the whole capture and ask whether other dialogs were retransmitting at the same moment.
- Resist tuning T1 downwards to "fix" it. Shorter timers put more copies of the same message on a path that is already failing to carry one.
hiccup collapses a retransmission ladder into a single readable row and classifies it — never arrived, slow far end, UDP fragmentation, ACK not landing, or a box-wide storm that was never about your call. Self-hosted, free for individual users.
upload a trace