RTP and RTCP: how the audio actually gets there
Short version: RTP puts a sequence number and a media clock on every packet so the receiver can rebuild timing that the network destroyed. RTCP is the receiver telling the sender how badly that went. Between the two of them, a capture contains everything needed to say how a call sounded — no listening required.
The RTP header: five fields do all the work
| Field | What it is | What you use it for |
|---|---|---|
| SSRC | Random 32-bit stream id | Defines the stream. One SSRC = one sender; it changes on restart, and a mid-call SSRC change explains many "audio came back different" reports. |
| Sequence number | 16-bit, +1 per packet, random start | Loss (gaps), reordering (backwards steps), duplication. Wraps every 65 536 packets — about 22 minutes at 20 ms — so tools track cycles. |
| Timestamp | Media clock, not wall clock — 8000 Hz for narrowband, so +160 per 20 ms packet | Jitter calculation, silence-suppression gaps (timestamp jumps while sequence stays contiguous), and clock-rate mismatches (chipmunk audio). |
| Payload type | 7-bit codec label; static below 96, dynamic 96–127 mapped by SDP a=rtpmap | What the packet claims to contain. PT changing mid-stream without a re-INVITE is a bug worth chasing; a separate PT (often 101) is DTMF events. |
| Marker bit | 1 bit, codec-specific meaning | For voice: first packet after silence — talkspurt boundaries, useful when reading comfort-noise behaviour. |
Jitter, precisely
Jitter is not "variance in ping". RFC 3550 defines it exactly: for each pair of packets, compare the spacing at which they arrived with the spacing at which the timestamps say they were sent; the jitter value is a running smoothed average of that difference (in timestamp units — divide by 8 for milliseconds on an 8 kHz codec). The receiver hides jitter with a buffer that trades delay for smoothness: an adaptive buffer grows during bad spells, and every time a packet arrives later than the buffer's current depth, it is late loss — discarded as uselessly late even though it arrived. This is why "the capture shows only 0.4% loss but the call sounded terrible" is not a contradiction: the network delivered the packets, the deadline did not.
RTCP: the feedback channel
RTCP runs beside RTP (traditionally on the next odd port, or muxed on the same one
with a=rtcp-mux), throttled to a few percent of the media bandwidth.
Two packet types matter daily:
- Sender Report (SR): "I have sent N packets, M octets, and my RTP timestamp T corresponds to this NTP wall-clock time." That last pairing is the bridge between media time and real time — it is how receivers lip-sync audio to video, and how tools convert timestamps to seconds.
- Receiver Report (RR): the far end's experience, per SSRC: fraction lost since the last report, cumulative packets lost, highest sequence received, current interarrival jitter, and two fields (LSR/DLSR) that let the original sender compute round-trip time without any clock agreement.
Read RRs as a time series, not a single sample: fraction-lost spiking in one report then recovering is a burst event; cumulative loss climbing steadily is sustained congestion; jitter trending up ahead of loss is a queue filling somewhere. And the absence of RTCP is itself a finding — plenty of middleboxes forward RTP but drop RTCP, blinding both ends' adaptive machinery.
From counters to "how it sounded"
The E-model (ITU-T G.107) turns delay, loss and codec into a single R score, which maps to the familiar MOS scale. The useful engineering intuitions: a narrowband call starts from a ceiling around MOS 4.4 before the network takes anything; one-way delay starts to hurt conversation quality past roughly 150 ms and becomes obvious past 250 ms; random 1% loss is audible on G.711 and worse on compressed codecs, while bursty loss at the same average is far more damaging than the average suggests. Any tool's MOS from a capture — hiccup's included — is an estimate of network-inflicted damage, not a measurement at the ear: it cannot see the microphone, the acoustics, or an analogue leg beyond the last RTP hop.
Reading a media capture
- Group by SSRC and direction first. Two streams means two verdicts; averaging them hides the fault. (A missing direction entirely is one-way audio, a different investigation.)
- Walk the sequence numbers: gaps, reorder distance, duplicates. Then walk timestamps against arrival times for jitter and for silence-suppression gaps.
- Line the RTCP RRs up against the RTP: does the far end's reported loss match what you can see leaving? Loss reported that you did not see departing means the damage is downstream of your capture point — that single comparison localises the fault to a half of the path.
- Check the first and last packet times against the SIP ladder: media that starts 3 seconds after the 200 OK, or stops before the BYE, tells its own story.
hiccup does this pass automatically: per-stream loss, jitter, late-loss estimates and an honest quality verdict, lined up against the SIP timeline of the same capture. Self-hosted, free for individual users.
upload a trace