try hiccup

SIP retransmissions and retransmission storms

Short version: retransmissions are not the fault. They are UDP's transaction layer doing its job while something else is broken. What matters is which message is repeating and what is failing to come back — those two facts identify the fault almost on their own.

What counts as a retransmission

A retransmission is the same message, sent again by the same transaction. Concretely: identical branch parameter on the top Via, identical CSeq, identical request line. Three things get mistaken for it:

The timers

RFC 3261 defaults: T1 = 500 ms (the assumed round trip), T2 = 4 s (the ceiling for most retransmit intervals), T4 = 5 s.

What is repeatingIntervalGives up after
INVITE request (Timer A)T1, doubling — no capTimer B, 64 × T1 = 32 s
Non-INVITE request (Timer E)T1, doubling, capped at T2Timer F, 64 × T1 = 32 s
Final non-2xx response (Timer G)T1, doubling, capped at T2Timer H, 64 × T1, or an ACK
2xx response to an INVITET1, doubling, capped at T264 × T1, or an ACK

Two consequences fall straight out of that table. Over a reliable transport such as TCP there are no transaction-layer retransmissions at all — but the 32-second timeout still applies. And ordinary provisional responses are never retransmitted, so a lost 180 Ringing is simply lost; only reliable provisionals (RFC 3262, the ones that demand a PRACK) get repeated.

Five patterns and what each one means

INVITE repeating, nothing coming back

The far end never saw it, or cannot answer. Six retransmissions and a local 408 — see 408 Request Timeout for the full list of reasons, of which oversized fragmented UDP is the one people miss most often.

INVITE repeating, then a late 100 Trying

Not a loss problem: the far end is slow. Something in its path — a database lookup, an ENUM query, a DNS resolution it does synchronously — is taking longer than 500 ms, so it has not yet produced a provisional. The retransmissions stop the moment the 100 arrives, and the call usually completes. Worth fixing, rarely worth alarming about.

200 OK repeating, no ACK

The most diagnostic pattern in SIP, and always worth stopping for. The callee answered and cannot get the acknowledgement back, so it repeats the 200 OK for 64 × T1 and then tears down the call with a BYE — roughly 32 seconds of ringback-then-nothing that users report as "it answers and then hangs up". The ACK for a 2xx is a new transaction addressed to the Contact URI from the 200 OK, routed through the Record-Route set, so the usual causes are a Contact containing a private or otherwise unroutable address, a Record-Route set that a B2BUA mangled, or a firewall permitting the INVITE path but not the ACK's.

A final 4xx/5xx/6xx repeating

Same idea, opposite direction: the server transaction is retransmitting its final response because no ACK arrived. Note that the ACK for a non-2xx is generated by the transaction layer and carries the same branch as the INVITE — unlike the 2xx case — so if you cannot find one, look for the branch rather than the dialog.

Everything on the box repeating at once

When unrelated calls all start retransmitting inside the same few seconds, the call you were asked about is a bystander. The tell is correlation in time rather than in dialog: an interface flap, a CPU or queue overload, a downstream peer that went away, or a DNS server that stopped answering. Fixing the "call" in this situation achieves nothing, and the storm can become self-sustaining, because every retransmission adds load to the box that is already too busy to answer.

What to check in the trace

  1. Group by Via branch, not by Call-ID. That single grouping separates retransmissions from re-INVITEs and forks without any guesswork.
  2. Check the intervals against the table above. A ladder that does not double is not RFC 3261's transaction layer — it is an application-level retry, and those come from somewhere much more interesting.
  3. Establish what should have answered and did not: a provisional, a final response, or an ACK. That is the actual fault.
  4. Measure message size and look for IP fragmentation on anything approaching the path MTU.
  5. Widen the window to the whole capture and ask whether other dialogs were retransmitting at the same moment.
  6. Resist tuning T1 downwards to "fix" it. Shorter timers put more copies of the same message on a path that is already failing to carry one.

hiccup collapses a retransmission ladder into a single readable row and classifies it — never arrived, slow far end, UDP fragmentation, ACK not landing, or a box-wide storm that was never about your call. Self-hosted, free for individual users.

upload a trace