try hiccup

NAT traversal for voice: STUN, TURN and ICE

Short version: SIP breaks across NAT because it writes IP addresses inside the messages, and NAT only rewrites the ones on the outside. Every traversal mechanism is a different way of getting a usable address into the SDP — ask the network (STUN), rent a relay (TURN), try everything and keep what works (ICE), or let the carrier's SBC ignore the SDP entirely (latching).

The problem, precisely

A phone on 192.168.1.20 puts that address in its Via, its Contact, and — fatally — in the SDP c= line that tells the far end where to send audio. The NAT rewrites the IP header on the way out; the SIP body sails through untouched. The far end then obediently sends RTP to 192.168.1.20, which from the internet is nowhere. Result: the most common media failure in existence — signalling fine, audio missing in one direction, covered from the symptom side in one-way audio.

Two protocol-level patches help the signalling itself: rport (RFC 3581) makes responses return to the source port the request actually came from, and symmetric behaviour (sending and receiving on the same port) keeps the NAT mapping alive and predictable. But something still has to fix the addresses in the SDP.

The mechanisms, and what each one actually contributes

MechanismWhat it doesWhere it falls short
STUNA server on the internet tells you what your address looks like from outside ("server-reflexive" address). Cheap, stateless.Useless alone behind NATs whose mapping depends on the destination (the behaviour loosely called symmetric NAT): the address STUN discovered is only valid towards the STUN server.
TURNRents you a relay address on a server; all media flows through it. Works behind anything.Cost and latency — every packet takes the detour, and someone pays for the relay's bandwidth. The fallback, not the plan.
ICEThe system that uses both: gather every candidate address (local, STUN-derived, TURN relay), exchange them in SDP, connectivity-check all pairs with STUN packets, pick the best that works.Complexity and setup time — a full check round takes hundreds of milliseconds, and both ends must speak it. Universal in WebRTC; still patchy in desk-phone SIP.
SBC latchingThe carrier-side answer: the SBC ignores the SDP address and locks onto wherever the endpoint's RTP actually arrives from, then sends its own media back to that observed source.Requires the endpoint to send first, and to send symmetrically. Which is why "no audio until the callee speaks" is the classic latching-adjacent symptom.

ICE in ninety seconds

Each side gathers candidates — host (its own interfaces), srflx (what STUN saw), relay (what TURN rented) — and ships them as a=candidate lines with priorities: host preferred, relay last. Both sides then pair their candidates with the peer's and send STUN binding requests over every pair, in priority order. The pairs that produce answers are usable; the controlling side nominates one, and media settles onto it. Meanwhile — and this is the detail that saves real calls — media can start flowing on the first working pair before checking finishes. In a capture, a working ICE call shows a burst of STUN on the media ports before and during early RTP; a failing one shows checks in one direction only, which localises the broken NAT instantly.

The SIP ALG, and why the advice is always "turn it off"

Consumer routers ship an application-layer gateway that tries to fix SIP by rewriting the addresses inside the messages as they pass. The idea is sound; the implementations are, almost universally, not. Real ALGs rewrite the Via but not the SDP, or the SDP but not the Contact; they mangle one message in a dialog and miss its retransmission; they fight with the phone's own STUN or the carrier's latching, each undoing the other's correction; and they rewrite REGISTER Contacts so registrations point at addresses that stop existing when the mapping expires. The resulting faults are gloriously inconsistent — calls that work until they re-INVITE, audio that dies after exactly the NAT's UDP timeout, registrations that flap. It is the first thing to disable on any site with inexplicable VoIP behaviour, and the trace fingerprint is an address in SIP that neither endpoint claims to have written.

Reading a NAT problem in a trace

  1. Compare the SDP c= address with the actual source of the RTP. A mismatch is not automatically a fault (latching tolerates it), but it names the mechanism in play.
  2. Private addresses (RFC 1918) in c= or Contact on a public-facing leg mean no traversal mechanism ran — or an ALG stripped its work.
  3. Look for the STUN checks of ICE (a=candidate in SDP, STUN on the media ports). Checks in one direction only point at the side whose NAT is eating inbound packets.
  4. Audio that dies mid-call after a suspiciously round number of seconds is a NAT mapping expiring: look for missing keepalives (empty RTP, STUN, or RTCP) during silence.
  5. Mangled-but-almost-right addresses — right subnet, wrong port; rewritten Via with untouched SDP — are the ALG's signature.

hiccup decodes the STUN/ICE traffic in your capture alongside SIP and RTP, flags private addresses in public SDP, and tells you whether media actually followed the addresses the signalling promised. Self-hosted, free for individual users.

upload a trace