Skip to main content
TCP / IP
Detailed for the masses
Hi
 Julien PAULI
 Programming in PHP since early 2000s
 PHP Internals hacker and trainer
 PHP 5.5/5.6 Release Manager
 Working at SensioLabs in Paris - Blackfire
 Writing PHP tech articles and books
 http://phpinternalsbook.com
 @julienpauli - github.com/jpauli - jpauli@php.net
 Like working on OSS such as PHP :-)
TOC
 We can't explain "the network" , in an hour or
so (sorry)
TOC
 We can't explain the network , in an hour
 Even if that latter is not that hard to understand
 In fact, it is easy to spot it in its whole, but long
TOC
 Go further by yourself with highly
recommanded books =>
TOC (the real one)
 Reminder on OSI
 Brief presentation of IP : Internet Protocol
 Deep dive into TCP
 Conclusion
TOC (detailed)
 Deep dive into TCP
 The need of reliability on top of IP
 ARQ: Automatic Repeat reQuest and acknowledgments
 TCP header overview
 Pipelined reliable Service model
 Connected model and FSM
 Data Flow, segmentation, MSS and sliding Window
mechanism
 Congestion control and Avoidance algorithms
 Tuning TCP on Linux
 Wireshark examples
Don't be scared
Everything is under control
 Every concept is
 Logical (common sense)
 Maths oriented
 Fully understandable (engineer capabilities assumed)
 TCP is not that hard, it is purely logical
 TCP is not that huge, it mixes several concepts
 TCP is a protocol, a binary language talked
between several hosts
 TCP is technically really beautiful
Mantra
 TCP/IP runs the Internet since 1970, aka 90% of it
(very huge part of it)
 TCP runs over another protocol, usually over IP
 TCP/IP used to be one layer, and one protocol
 They got separated to respect OSI
 IP is layer 3 , TCP is layer 4 : they are used together
 TCP/IP understanding helps a lot in designing other
(Layer 7) protocols
 A simple TCP implementation requires only several
hundreds LOC (C lang)
OSI
OSI encapsulation
IP
TCP
TCP/IP model
 Just a conceptual simplification of OSI
 OSI 5-6-7 and OSI 1-2 are commonly merged
TCP/IP model
Vocabulary (only)
 Every layer adds its infos
 Every layer carries some data
 But we used to give different wordings to them
Frame
Datagram
Segment (also packet)
payload, packet, … ?
IP
IP
 Internet protocol
 Sends a datagram from place A to place B
 A could be several thousands Km far from B
 Datagram will be routed, through routers (packet-
switched network)
 from 1 router, to several dozens, hundreds
 Datagram may or may not arrive at destination
 Datagrams may arrive in different order
 "Best effort" protocol
 Do what you can to route info, as fast as possible
 Do not bother to "check" anything, just route, and go
IP header (IPV4)
IP Forward
 Routers analyze the destination address field
 They route the IP datagram on connected
networks
 Datagram may get destroyed (router overload)
 Datagram may get lost (iterconnexion problem)
 Datagram may get altered, falsified (noise)
 Datagram may arrive good, but the host is
down, or cant process it
Typical routing scenario
Let's go
Typical routing scenario
1
2
3
4
Typical routing scenario
1
2
3
4
Typical routing scenario
1
2
3
4
Typical routing scenario
1
34
2
Typical routing scenario
1
3
4
Typical routing scenario
1
3
4
Typical routing scenario
1
3
4
?
IP notes
 IP is a best effort protocol
 Many scenarios may be part of a failure
 Datagrams can be destroyed / lost
 Datagrams can be misordered
1234 ? 1 34
Remember
 IP is not reliable
 IP is a network protocol : it does what it can to route
packets from A to B , that's all
 IP offers no delivery garanty
 IP route traffic but doesnt route service
 When a datagram reaches a host, which
process should handle it ?
Layer 4 protocols
UDP
 Basically, UDP adds ports to IP
 It then allows multiplexing data streams
 When a datagram arrives, UDP layer processes
it and distributes it to the right host process by
reading the destination port
UDP
 UDP is light : only 8 bytes overhead per
message
 UDP does not provide
 Reliability
 Error correction mechanisms
 Connected service
TCP
TCP is more complex than UDP
 Transmission Control Protocol
 That means : control the transmission :-p
 TCP is order of magnitudes more complex than
UDP
 That's not hard to do given the simplicity of UDP
 TCP brings a lot of intelligence
 UDP has none at all
TCP overview
 TCP runs on top of IP
 TCP provides
 Full duplex connection oriented data exchanges
 Bytes can go in the 2 ways - Senders may be receivers
 PointToPoint connection
 No broadcast , no multicast , only 2 points
 Service multiplexing
 Through ports
 Stream data transfer
 No need to prepare the data to a specific form before sending
TCP overview
 … TCP provides
 Reliability
 data sent will be checked and ACKed
 Order
 data will be treated in the same order its been sent
 Flow control
 data will arrive whatever the processing speed at either
side*
 Congestion avoidance
 data will arrive whatever the middleware network load*
* : assuming L3 is not down
TCP header
 Variable length :
 20 bytes minimum
 60 bytes maximum (with all options field)
TCP header
Transmission control
Congestion controlProtocol control bits
Service multiplexingService multiplexing
TCP header overview
 Very verbose, embeds many informations
 Min 20 bytes, min 2.5x > UDP header
 Carries ports, like UDP (con multiplexing)
 Carries a checksum, like UDP
 Carries Sequence number and ACK number
 The most important things in TCP
Reliability
The concept of reliability
 You are host "sender"
 You send a message to host "receiver"
 How can you be sure message arrived at
destination ?
ACK makes reliability
 The answer is to have receiver ACK sender's
message
1 1
ack 1ack 1
Reliability
 What happens if B doesn't ACK ?
 What happens if B's ACK is lost in network ?
1 1
ack 1
ARQ
 Automatic Repeat reQuest
 Sender re-sends the exact same message
 … and still waits until ACK
1 1
ack 1
1 1
ack 1ack 1
TCP :-)
 You just understood the very main baseline
behind TCP
 TCP is a connected protocol
 sender and receiver keep communicating
with each other
 receiver ACK messages received from sender
It all starts by a TCP
connection
TCP connection
 Before exchanging data, hosts must handshake
 Like in phone conversation
 Sen ->Rec "Hello ?"
 Rec ->Sen "Hello !"
 Sen ->Rec "Hi, right, let's start talking now"
 This is called the 3 way handshake
 Because it effectively requires 3 message exchanges
 Before any data exchange
TCP 3-way handshake
SYN
SYN
ACK
ACK
ACK bit
SYN bit
active opener
connect()
active opener
connect()
passive opener
listen()
Communication negotiation
 When in SYN, hosts will also exchange protocol
details
 Those are stuck in the options of the header
 "Do you support feature XXX" ?
 "Yes I do, we will use it"
 "No I dont, never use it for this communication"
options
TCP options negociation
 MSS is mandatory (IP MTU negociation)
 Prevents hawful IP fragmentation (end-to-end only)
 Window Scale factor is used as of 2017
 "SACK" is really commonly used
Let's disconnect
 We are now connected
 Disconnection is also a matter of message
exchanges
 But remember the full-duplex bi-directionnal
feature of TCP
 Connection must be closed in both ways
 From A to B
 And from B to A
 (Or the opposite : it doesn't matter)
TCP full closing
 One host (A or B) wishes to disconnect
 FIN - ACK and FIN - ACK
FIN
ACK
FIN
active closer
passive closer
ACK
FIN bitFIN bit
TCP half-closing
 Connection could be half closed
 Host A cant send any data to host B anymore
 But host B still can send data to host A
 Until host B initiates a FIN from its side
 This is uncommon , but still can happen
 close() : full close
 shutdown() : half close
 The host closing the connection will free
resources on its side
 It may then be interesting in some scenarios
TCP half-closing
 One host (A or B) wishes to disconnect
 Sends a FIN
 Waits for the corresponding ACK
 A can't send anymore data to B now
 But will still ACK it !
 But B still can send data to A
 Until itself sends a FIN , and receives the ACK
FIN
ACK
A B
TCP FSM
Remember
 A TCP connection is uniquely identified by the
quadruplet
 Source IP address
 Source Port
 Destination IP address
 Destination Port
 Those 4 infos together, are unique quadruplet
 A TCP connection needs a preamble handshake
 A TCP connection needs a preamble close dialogue
to close
 And it may be "half closed"
Remember
 A TCP connection is between 2 and only 2 hosts
 One can be seen as the "client"
 The other as the "server"
 This is purely conceptual
 I prefer talking about "source" and "destination"
 Or "host A" and "host B"
 client/server , is what will be done at layer 7, with TCP
 Connection will run in both directions (full duplex)
 Any part, can initiate a close, whenever it wants
Data exchange
Reminder : ACK
 Ok with that ?
1
ack 1
2
ack 2
...
This is highly problematic !
 Guess the problem ?
Bandwidth inefficient
RTT
RTT
 Round Trip Time - The time for a packet
 To be generated
 Create the segment
 Compute the checksum
 Add the header to the segment
 To be sent - routed
 Router1 → Router 2 → Router 3 - - - - - - → Router 42 - - - - ?
 To be received by destination
 NIC will interrupt CPU
 To be treated by destination
 Destination will extract header
 Check the checksum
 Done
RTT
1
ack 1
...
waiting for ACK
 If you wait during RTT for an ACK …
 You highly underuse the network bandwidth
1
ack 1
t
What do you suggest ?
Sliding Window
 Don't send just one packet at a time , but
several !
 Wait until ACKed
 Re-send if not
 Go further
 This is the crucial concept of TCP sliding
Window
Multiple packets on the way
 TCP sliding window
1234
ack 1 ack 2 ack 3 ack 4
TCP Sliding Windows Demo
 The best link about Sliding Window :
 http://www2.rad.com/networks/2004/sliding_win
dow/
TCP's heart in just one picture
SEQ numbers and ACK dialog
Seq numbers and ACK
 TCP is a stream oriented protocol
 Layer 7 send() to a TCP socket some bytes
 TCP shrinks those bytes to some segment sizes
 The current segment size depends on
 MSS announced in SYN phase
 Total data volume to transfer
 TCP starts sending segments
Seq numbers
 The Sequence number indicates where we are
in the sent stream from our sender view of
the connection
 The ACK number indicates what we expect
next from the receiver view of the
connection
 We always ACK bytes, not segment
numbers
 We can use cumulative ACK
Communication example
Example
 A got 750 bytes to send to B
 We'll assume A and B have connected
successfully
 Let's say segment size is 100 bytes
 Let's say window size is 500 bytes
A B
750 bytes
Seq and ACK in byte stream
 A receives from B
 ACK = 101 meaning B received bytes 1-100 and now expects 101
 ACK = 201 meaning B received bytes 1-200 and now expects 201
 ACK = 301 meaning B received bytes 1-300 and now expects 301
 A received from B ack = 301 so far
 That means that A received contiguously up to byte 300
 That means that A is expecting byte 301 now
seq = 1
len = 100
seq = 101
len = 100
seq = 201
len = 100
seq = 301
len = 100
seq = 401
len = 100
ack = 101ack = 201 ack = 301
Seq and ACK in byte stream
ack = 101
seq = 301
len = 100
seq = 401
len = 100
ack = 401
seq = 501
len = 100
seq = 601
len = 100
ack = 201 ack = 301
 Host A retransmits missing ACK segments
 Host A got as last ACK=601 from B
ack = 501 ack = 601
Seq ACK
 Remember ACK show what host expects next
 Here, B received up until 601, but not 601
ack = 601ack = 401ack = 501
Seq ACK
seq = 601
len = 100
seq = 701
len = 50
ack = 701
seq = 701
len = 50
ack = 751
ack = 701
Done !
 We just transfered bytes with TCP
 We managed to cope with lost segments
 Remember TCP pairs keep communicating
together
 Sender retransmit what receiver missed
SACK
Selective ACK (SACK)
 ACK system is nice but got a drawback
 By design, TCP resends every segments from
the point which has been lost
 TCP does not ACK each received segment
 But TCP ACKS the most recent contiguous
received bytes
 SACK is about ACKing segments
 To better detect holes, and only resend lost segments
TCP default ACK drawbacks
 The answer from B to A is 101 followed by 3 times 201
 That means that B has received the first 2 segments OK
 And that B received a total of 4 segments
 Where are the holes ? Which segments have been lost ?
 Segment 201 ? 301 ? 401 ?
 By default ; TCP must resend from 201 to 501 (3 segments)
seq = 1
len = 100
seq = 101
len = 100
seq = 201
len = 100
seq = 301
len = 100
seq = 401
len = 100
ack = 101ack = 201 ack = 201ack = 201
SACK in a word
 SACK is an option that allows ACKs to signal a
range of noncontiguous data received so
far
 The sender then has a better idea of which
segments have been received by the receiver,
or not
 SACK is a TCP Option (RFC 2018)
 That must be negociated in the handshake
 Both A and B must support it to use it
SACK example
seq = 1
len = 100
seq = 101
len = 100
seq = 201
len = 100
seq = 301
len = 100
seq = 401
len = 100
ack = 101ack = 201 ack = 201
sack = 301
ack = 201
sack = 301-401
seq = 201
len = 100
ack = 501
End-to-end Congestion control
Window size adjustment
Sliding Window size
 The Window size is how many bytes the
receiving part is able to proceed
 This is typically the side of its buffer
 It is also how maximum bytes can be in flight
seq = 645
len = 1000
seq = 1645
len = 1000
2000 bytes in flight
Sliding Window size
 When host B receives those, it sticks the
segments together in a buffered stream
 It effectively recreates the sent stream on its side
 TCP pushes up the bytes to layer 7
 Now, B layer 7 app must receive that buffer
recv()
tcp buffertcp buffer
layer 7 buffer
Sliding Window size
 What if B processes the bytes too slowly ?
 By reading f.e 10 bytes only from the received stream
 B's TCP receive buffer will saturate
 Because B doesn't process it fast enough at layer 7
 If A keeps sending, B won't be able to handle the
traffic
 A congestion will appear
 Resources will be wasted, as well as network bandwidth
 Packets will be lost and need re-send
recv(10)
tcp buffertcp buffer
layer 7 buffer
Window size advertisement
 The receiver advertises its window size to
the sender depending on its saturation
seq = 42
len = 1000
seq = 1042
len = 1500
seq = 2542
len = 1500
ack = 1042
win = 500
seq = 1042
len = 100
seq = 1142
len = 100
seq = 1242
len = 100
seq = 1342
len = 100
seq = 1442
len = 100
4000
bytes
500
bytes
Window size advertisement
 The receiver advertizes its receiving capabilities
to the sender
ack = 1042
win = 500
ouch, please don't send me more than 500 bytes as of now !
Window size
Window 0
 Eventually, the receiver will completely saturate
 It will advsertise a 0 Window
ack = 1042
win = 0
Please, stop any transfer
Wireshark Demo
Network congestion detection
and avoidance
What do you do when you
detect congestion ?
You ask to slow things down
Types of congestions
 So far, we've seen how hosts can advertize their
capabilities
 The receiving host advertises the sending host on
its receiving capabilities
 By adjusting the sliding window size
 That is, the number of bytes allowed in flight
 This way, the sender will never saturate the receiver
 But how to detect network (middle) congestion,
and what to do in such cases ?
The need to detect congestions
 Remember that TCP is blind about IP traffic*
 TCP must continuously guess the IP layer state from
PTP
 TCP must measure and evaluate continuously
 RTT
 Dup ACKs, meaning packet lost
 Advertized window size
 And TCP must run specific algorithms in case of
congestion detection
*: congestion signals exist (IP ECN)
Network Congestion
123 2
 Detecting network congestion is easy
 Lost packets (dup ACKs)
 TTL increasing
 But getting congestion to smoothly resorb it,
without improving it, is a challenge
 TCP must do its best to smoothly use the network
 TCP must guess, and act on the unknown network
state