try hiccup

SIP 408 Request Timeout

Short version: most 408s you see in a CDR were never sent by anybody. Your own stack minted one when its transaction timer expired after 32 seconds of silence. The fault is in the packets that are missing, not in the 408.

Where a 408 comes from

There are two completely different animals wearing the same number.

A received 408 is a real response from a real server that could not produce an answer in time — classically a proxy that could not locate the user quickly enough. These are comparatively rare, and you can see them in the capture as an actual packet with a Via stack.

A synthesised 408 is what your UA, proxy or SBC writes into its own logs and CDRs when a transaction times out with no response at all. Nothing crossed the wire. Signalling-layer tools will show it, packet captures will not, and that discrepancy is itself the clue: if the CDR says 408 and the pcap contains no such packet, you are looking at a local timeout.

Where 32 seconds comes from

RFC 3261 sets T1 to 500 ms as the default round-trip estimate. For an INVITE over UDP, Timer A starts at T1 and doubles after every retransmission, while Timer B — the transaction timeout — is fixed at 64 × T1 = 32 seconds. That produces the ladder everyone recognises:

t=0.0s   INVITE
t=0.5s   INVITE (retransmit 1)
t=1.5s   INVITE (retransmit 2)
t=3.5s   INVITE (retransmit 3)
t=7.5s   INVITE (retransmit 4)
t=15.5s  INVITE (retransmit 5)
t=31.5s  INVITE (retransmit 6)
t=32.0s  Timer B expires -> local 408

Seven INVITEs and then a timeout is not a symptom of anything specific; it is simply what "no reply" looks like. Non-INVITE requests behave the same way except that their retransmit interval caps at T2 = 4 seconds, and Timer F still ends things at 32.

One detail worth committing to memory: a provisional response stops the clock. Once a 100 Trying arrives, the client transaction leaves the Calling state, retransmissions stop, and Timer B no longer applies. A call that hangs for three minutes before failing therefore did not die of Timer B — that is a proxy's Timer C (RFC 3261 requires more than 3 minutes) giving up on a transaction that was acknowledged and then abandoned.

What makes the packets disappear

The request never arrived

Firewall or SBC signalling ACL not updated after the carrier changed a media/SIP address, a route pointing at a decommissioned node, or an SRV record that resolves to a host nobody has been maintaining. Capture at the egress interface: if you can see your own retransmissions leaving, the loss is downstream of you.

The response could not get back

Everything about SIP response routing depends on the Via header. If the far end ignores rport (RFC 3581) and answers to the address in sent-by rather than the source it actually received the packet from, responses go to a private address and die. NAT bindings that expire faster than the call takes to be answered produce the same picture, as does asymmetric routing where the return path skips the stateful device that is holding the pinhole open.

The request was too big for the path

A UDP INVITE larger than the path MTU fragments, and plenty of firewalls and NATs quietly discard non-initial fragments. The far end never reassembles a complete datagram, so it never answers, so you retransmit the same oversized message six more times. RFC 3261 is explicit about this: a request within 200 bytes of the path MTU, or over 1300 bytes when the MTU is unknown, must be sent over a congestion-controlled transport such as TCP. Large SDP bodies, long Route sets, P-Asserted-Identity plus History-Info plus a certificate-bearing header — they add up faster than people expect.

Name resolution

RFC 3263 resolution walks NAPTR, then SRV, then A. A DNS server that blackholes NAPTR queries instead of answering "no such record" adds seconds before the request even leaves. Where DNS is the culprit the capture shows the timeout with no outbound INVITE at all, which is a very distinctive shape.

What to check in the trace

  1. Did any provisional response arrive? A 100 Trying changes the diagnosis completely — the far end saw the request, and the problem is beyond it.
  2. Count the retransmissions and check their spacing. A clean 0.5/1/2/4/8/16 ladder means your stack behaved normally and nothing came back. An irregular ladder means something else is going on.
  3. Measure the INVITE's size on the wire, and look for IP fragmentation. Any request near or over ~1300 bytes on UDP is a suspect.
  4. Compare the Via sent-by, received and rport values against the actual source address of the packet.
  5. Check whether the same peer's OPTIONS or REGISTER keepalives were also timing out at the same moment. If they were, the call is a bystander and the peer is the story.
  6. Capture on both sides of every NAT or SBC in the path. A 408 is the one failure where a single vantage point is most likely to mislead you.

hiccup collapses a retransmission ladder into one row, tells you whether it ended in a timeout or a late reply, and flags oversized UDP requests and fragmented datagrams as findings rather than leaving them for you to spot. Self-hosted, free for individual users.

upload a trace