try hiccup

RTP and RTCP: how the audio actually gets there

Short version: RTP puts a sequence number and a media clock on every packet so the receiver can rebuild timing that the network destroyed. RTCP is the receiver telling the sender how badly that went. Between the two of them, a capture contains everything needed to say how a call sounded — no listening required.

The RTP header: five fields do all the work

FieldWhat it isWhat you use it for
SSRCRandom 32-bit stream idDefines the stream. One SSRC = one sender; it changes on restart, and a mid-call SSRC change explains many "audio came back different" reports.
Sequence number16-bit, +1 per packet, random startLoss (gaps), reordering (backwards steps), duplication. Wraps every 65 536 packets — about 22 minutes at 20 ms — so tools track cycles.
TimestampMedia clock, not wall clock — 8000 Hz for narrowband, so +160 per 20 ms packetJitter calculation, silence-suppression gaps (timestamp jumps while sequence stays contiguous), and clock-rate mismatches (chipmunk audio).
Payload type7-bit codec label; static below 96, dynamic 96–127 mapped by SDP a=rtpmapWhat the packet claims to contain. PT changing mid-stream without a re-INVITE is a bug worth chasing; a separate PT (often 101) is DTMF events.
Marker bit1 bit, codec-specific meaningFor voice: first packet after silence — talkspurt boundaries, useful when reading comfort-noise behaviour.

Jitter, precisely

Jitter is not "variance in ping". RFC 3550 defines it exactly: for each pair of packets, compare the spacing at which they arrived with the spacing at which the timestamps say they were sent; the jitter value is a running smoothed average of that difference (in timestamp units — divide by 8 for milliseconds on an 8 kHz codec). The receiver hides jitter with a buffer that trades delay for smoothness: an adaptive buffer grows during bad spells, and every time a packet arrives later than the buffer's current depth, it is late loss — discarded as uselessly late even though it arrived. This is why "the capture shows only 0.4% loss but the call sounded terrible" is not a contradiction: the network delivered the packets, the deadline did not.

RTCP: the feedback channel

RTCP runs beside RTP (traditionally on the next odd port, or muxed on the same one with a=rtcp-mux), throttled to a few percent of the media bandwidth. Two packet types matter daily:

Read RRs as a time series, not a single sample: fraction-lost spiking in one report then recovering is a burst event; cumulative loss climbing steadily is sustained congestion; jitter trending up ahead of loss is a queue filling somewhere. And the absence of RTCP is itself a finding — plenty of middleboxes forward RTP but drop RTCP, blinding both ends' adaptive machinery.

From counters to "how it sounded"

The E-model (ITU-T G.107) turns delay, loss and codec into a single R score, which maps to the familiar MOS scale. The useful engineering intuitions: a narrowband call starts from a ceiling around MOS 4.4 before the network takes anything; one-way delay starts to hurt conversation quality past roughly 150 ms and becomes obvious past 250 ms; random 1% loss is audible on G.711 and worse on compressed codecs, while bursty loss at the same average is far more damaging than the average suggests. Any tool's MOS from a capture — hiccup's included — is an estimate of network-inflicted damage, not a measurement at the ear: it cannot see the microphone, the acoustics, or an analogue leg beyond the last RTP hop.

Reading a media capture

  1. Group by SSRC and direction first. Two streams means two verdicts; averaging them hides the fault. (A missing direction entirely is one-way audio, a different investigation.)
  2. Walk the sequence numbers: gaps, reorder distance, duplicates. Then walk timestamps against arrival times for jitter and for silence-suppression gaps.
  3. Line the RTCP RRs up against the RTP: does the far end's reported loss match what you can see leaving? Loss reported that you did not see departing means the damage is downstream of your capture point — that single comparison localises the fault to a half of the path.
  4. Check the first and last packet times against the SIP ladder: media that starts 3 seconds after the 200 OK, or stops before the BYE, tells its own story.

hiccup does this pass automatically: per-stream loss, jitter, late-loss estimates and an honest quality verdict, lined up against the SIP timeline of the same capture. Self-hosted, free for individual users.

upload a trace