Compare commits

...
Author SHA1 Message Date
Leency 658a3bc5dd kernel/net: keep out-of-order TCP segments in their own NET_BUFFs
Check kernel codestyle / Check kernel codestyle (pull_request) Successful in 26s
Test PR / Build (en_US) (pull_request) Successful in 3m34s
Test PR / Build (ru_RU) (pull_request) Successful in 3m41s
Test PR / Build (es_ES) (pull_request) Successful in 3m44s
Summary: the reassembly queue no longer copies every out-of-order segment
into a malloc'd node; it queues the NET_BUFF the segment arrived in, as
suggested in review. Also fixes a stuck full queue and FIN processing in
the FIN_WAIT states.

Details:
malloc is the kernel's small heap: 128 KB with no way to grow. One socket
was allowed 64 queued segments, about 94 KB of payload, so a single
download over a lossy link could take almost the whole heap away from
every other kernel user. The buffer the segment arrived in is owned by
nobody else, so it now becomes the queue node itself: NextPtr links the
queue, offset and length are narrowed down to the TCP payload, type is
set to the new NET_BUFF_TCP_REASM, and the sequence number and FIN flag
are kept right behind the NET_BUFF header (TCP_reasm_buff), over link
and IP headers that are no longer needed. Nothing is allocated or copied.

tcp_reassemble takes ownership of the buffer: it either queues or frees
it, and tcp_input marks this with TCP_BIT_BUFF_QUEUED so that its exit
path skips net_buff_free. Queued buffers come from the pool the network
drivers receive into, so besides the per-socket limit of 64 there is a
limit for all sockets together (TCP_reasm_maxbuffs, a quarter of the
pool), counted in TCP_reasm_buffs.

Neither limit applies to a segment that starts at RCV_NXT once the
connection is established. That segment is the one which drains the
queue; with the limit reached it used to be dropped like any other, and
a full queue could then never move again. Before ESTABLISHED nothing is
delivered, so there the limits stay in force.

The queue also records a FIN and tcp_reassemble returns TH_FIN once the
FIN is next in sequence, so tcp_input processes it without touching the
freed segment. This matters beyond out-of-order segments: the in-order
path is only taken in ESTABLISHED, so in FIN_WAIT_1 and FIN_WAIT_2 every
segment, in order or not, goes through the queue. Before, the FIN of a
peer answering our close was never processed there and the connection
never reached TIME_WAIT. Data is not delivered before the connection is
established; tcp_reassemble_drain delivers it once it is.

Tested in QEMU with rtl8139 and a relay that drops and reorders the data
segments of a 2 MB transfer checked byte by byte in the guest.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-28 14:42:02 +03:00
LeencyandClaude Opus 5 134f6c1936 kernel/net: implement the TCP reassembly queue for out-of-order segments
Check kernel codestyle / Check kernel codestyle (pull_request) Successful in 38s
Test PR / Build (en_US) (pull_request) Successful in 4m23s
Test PR / Build (es_ES) (pull_request) Successful in 3m54s
Test PR / Build (ru_RU) (pull_request) Successful in 4m18s
Summary: tcp_reassemble was a stub that returned immediately, so every
segment arriving out of order was discarded and a single lost packet threw
away the whole window in flight behind it.

Подробно:
The receive path had two exits: data landing exactly at RCV_NXT was copied
into the socket buffer, and anything else fell into .out_of_order, which
called the empty tcp_reassemble and jumped to .final_processing with a
comment reading "HACK because of unimplemented reassembly queue". The peer
then had to retransmit not only the lost segment but every segment sent
after it, and the connection collapsed to stop-and-wait for the duration of
the recovery. On a lossy link this is the difference between a download
running at line rate and one crawling.

The queue is a singly linked list of tcp_reasm_seg nodes hanging off
TCP_SOCKET.seg_next, sorted by sequence number, with the segment payload
copied into the node so that the caller can free the packet buffer as usual.
tcp_reassemble inserts and then delivers; tcp_reassemble_drain only delivers
and is used where a connection reaches ESTABLISHED. Delivery walks the list
from the head, trimming each node against RCV_NXT and stopping at the first
hole, so overlapping retransmissions never deliver a byte twice. If
socket_ring_write cannot take everything -- the application is not reading
fast enough -- RCV_NXT advances only by what was actually stored and the
remainder stays queued, keeping the advertised window honest.

All sequence arithmetic subtracts and tests the sign, so it is correct
across the 32-bit wraparound. Segments below RCV_NXT are trimmed, complete
duplicates and data more than a receive buffer ahead of RCV_NXT are dropped,
and the queue is bounded at TCP_reasm_maxsegs nodes per socket by a counter
in the socket structure -- counting only the nodes walked during insertion
would bound nothing, since segments arriving in descending sequence order
insert at the head every time. socket_free releases anything still queued.

Two callers in tcp_input already tested seg_next before taking their fast
paths, so both the header-prediction path and the in-order slow path now
correctly divert into the queue while it is non-empty, and the retransmission
that closes a hole drains everything contiguous behind it in one go. A FIN
riding on an out-of-order segment is deliberately left unprocessed; the peer
retransmits it and it is handled when it arrives in sequence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 13:32:06 +03:00
5 changed files with 228 additions and 9 deletions

No files matched your search

+8
View File
@@ -142,6 +142,7 @@ struct TCP_SOCKET IP_SOCKET
ts_val dd ?
seg_next dd ? ; re-assembly queue
seg_count dd ? ; number of segments queued in it
ends
@@ -2132,6 +2133,13 @@ socket_free:
je @f
stdcall free_kernel_space, [ebx + STREAM_SOCKET.snd.start_ptr]
@@:
; free the TCP out-of-order reassembly queue, if any
.free_reasm:
cmp [ebx + TCP_SOCKET.seg_next], 0
je .reasm_done
call tcp_reasm_free_head
jmp .free_reasm
.reasm_done:
mov eax, ebx
.no_stream:
+1
View File
@@ -167,6 +167,7 @@ NET_HWACC_TCP_IPv4_OUT = 1 shl 1
; Network frame types
NET_BUFF_LOOPBACK = 0
NET_BUFF_ETH = 1
NET_BUFF_TCP_REASM = 2 ; held in a TCP reassembly queue, offset/length = TCP payload
struct NET_DEVICE
+12
View File
@@ -99,6 +99,17 @@ TCP_QUEUE_SIZE = 50
TCP_ISSINCR = 128000
; A NET_BUFF queued for reassembly (TCP_SOCKET.seg_next): offset/length = TCP payload
struct TCP_reasm_buff NET_BUFF
Seq dd ? ; sequence number of the first payload byte
Flags dd ? ; TH_FIN if the segment carried a FIN
ends
TCP_reasm_maxsegs = 64 ; max queued segments per socket
TCP_reasm_maxbuffs = NET_BUFFERS / 4 ; max queued segments, all sockets together
struct TCP_header
SourcePort dw ?
@@ -137,6 +148,7 @@ align 4
TCP_queue rd (TCP_QUEUE_SIZE*sizeof.TCP_queue_entry + sizeof.queue)/4
TCP_input_event dd ?
TCP_timer1_event dd ?
TCP_reasm_buffs dd ? ; NET_BUFFs held by all reassembly queues
endg
uglobal
+19 -4
View File
@@ -19,6 +19,7 @@ TCP_BIT_NEEDOUTPUT = 1 shl 0
TCP_BIT_TIMESTAMP = 1 shl 1
TCP_BIT_DROPSOCKET = 1 shl 2
TCP_BIT_FIN_IS_ACKED = 1 shl 3
TCP_BIT_BUFF_QUEUED = 1 shl 4 ; the reassembly queue owns the buffer
;-----------------------------------------------------------------;
; ;
@@ -927,7 +928,7 @@ endl
pop word[ebx + TCP_SOCKET.SND_SCALE]
@@:
call tcp_reassemble
call tcp_reassemble_drain
mov eax, [edx + TCP_header.SequenceNumber]
dec eax
@@ -1691,16 +1692,22 @@ endl
DEBUGF DEBUG_NETWORK_VERBOSE, "TCP data is out of order!\nSequencenumber is %u, we expected %u.\n", \
[edx + TCP_header.SequenceNumber], [ebx + TCP_SOCKET.RCV_NXT]
; Uh-oh, some data is out of order, let's call TCP reassemble for help
; The reassembly queue takes over the buffer: edx is invalid from here on
call tcp_reassemble ;;; TODO!
mov esi, [dataoffset]
add esi, edx
mov eax, [esp] ; the NET_BUFF
call tcp_reassemble
or [temp_bits], TCP_BIT_BUFF_QUEUED
; Generate ACK immediately, to let the other end know that a segment was received out of order,
; and to tell it what sequence number is expected. This aids the fast-retransmit algorithm.
or [ebx + TCP_SOCKET.t_flags], TF_ACKNOW
jmp .final_processing ;;; HACK because of unimplemented reassembly queue!
test eax, TH_FIN
jnz .process_fin
jmp .final_processing
.data_done:
;-----------------------------------------------------------------------------------
@@ -1711,6 +1718,7 @@ endl
test [edx + TCP_header.Flags], TH_FIN
jz .final_processing
.process_fin: ; must not use edx from here on
DEBUGF DEBUG_NETWORK_VERBOSE, "TCP_input: Processing FIN\n"
@@ -1797,11 +1805,18 @@ endl
call tcp_output
.done:
test [temp_bits], TCP_BIT_BUFF_QUEUED
jnz .buff_queued
DEBUGF DEBUG_NETWORK_VERBOSE, "TCP_input: dumping\n"
call net_buff_free
jmp .loop
.buff_queued:
add esp, 4 ; the queue has the buffer
jmp .loop
;-----------------------------------------------------------------------------------
;
; Drop segment, reply with an RST segment when needed
+188 -5
View File
@@ -601,17 +601,200 @@ tcp_mss:
;-----------------------------------------------------------------;
; ;
; tcp_reassemble ;
; tcp_reassemble: queue an out-of-order segment, deliver what is ;
; contiguous with RCV_NXT. The queue owns the NET_BUFF from here. ;
; ;
; IN: ebx = socket ptr ;
; edx = segment ptr ;
; IN: eax = NET_BUFF ptr ;
; ebx = socket ptr (locked) ;
; edx = TCP header ptr ;
; esi = ptr to segment data ;
; ecx = segment data length ;
; ;
; OUT: / ;
; OUT: eax = TH_FIN if the FIN is next in sequence, else 0 ;
; other registers preserved ;
; ;
; tcp_reassemble_drain: delivery only ;
; ;
;-----------------------------------------------------------------;
align 4
tcp_reassemble:
;;;;; TODO
pushad
call tcp_reasm_insert
call tcp_reasm_deliver
mov [esp + 28], eax ; eax slot of pushad
popad
ret
align 4
tcp_reassemble_drain:
pushad
call tcp_reasm_deliver
popad
ret
; IN: eax = NET_BUFF, ebx = socket, edx = TCP header, esi = data, ecx = length
align 4
tcp_reasm_insert:
mov edi, eax ; edi = buffer
mov eax, [edx + TCP_header.SequenceNumber]
movzx ebp, [edx + TCP_header.Flags]
and ebp, TH_FIN
test ecx, ecx
jnz @f
test ebp, ebp
jz .drop ; neither data nor FIN
@@:
; trim data below RCV_NXT
mov edx, eax
sub edx, [ebx + TCP_SOCKET.RCV_NXT]
jns .no_head_trim
neg edx
cmp edx, ecx
ja .drop ; complete duplicate
jb @f
test ebp, ebp ; only the FIN is new?
jz .drop
@@:
add esi, edx
add eax, edx ; seq = RCV_NXT now
sub ecx, edx
xor edx, edx
.no_head_trim:
cmp edx, SOCKET_BUFFER_SIZE
ja .drop
; limits do not apply to the segment at RCV_NXT: it drains the queue
test edx, edx
jnz .check_limits
cmp [ebx + TCP_SOCKET.t_state], TCPS_ESTABLISHED
jae .limits_ok
.check_limits:
cmp [ebx + TCP_SOCKET.seg_count], TCP_reasm_maxsegs
jae .drop
cmp [TCP_reasm_buffs], TCP_reasm_maxbuffs
jae .drop
.limits_ok:
mov [edi + NET_BUFF.type], NET_BUFF_TCP_REASM
mov [edi + NET_BUFF.length], ecx
sub esi, edi
mov [edi + NET_BUFF.offset], esi
mov [edi + TCP_reasm_buff.Seq], eax
mov [edi + TCP_reasm_buff.Flags], ebp
; find the insertion point, esi = link to update
lea esi, [ebx + TCP_SOCKET.seg_next]
.find:
mov edx, [esi] ; edx = successor candidate
test edx, edx
jz .found
mov ebp, [edx + TCP_reasm_buff.Seq]
sub ebp, eax
jg .found ; successor starts past us
mov esi, edx
jmp .find
.found:
; drop the segment if the previous node already covers it
lea ebp, [ebx + TCP_SOCKET.seg_next]
cmp esi, ebp
je .link ; inserting at the head
mov ebp, [esi + TCP_reasm_buff.Seq]
add ebp, [esi + NET_BUFF.length]
sub ebp, eax
sub ebp, ecx ; prev end - our end
js .link
jnz .drop ; covered, and then some
cmp [edi + TCP_reasm_buff.Flags], 0
je .drop ; covered exactly, no FIN to add
.link:
mov [edi + NET_BUFF.NextPtr], edx
mov [esi + NET_BUFF.NextPtr], edi
inc [ebx + TCP_SOCKET.seg_count]
lock inc [TCP_reasm_buffs]
inc [TCPS_rcvoopack]
add [TCPS_rcvoobyte], ecx
ret
.drop:
stdcall net_buff_free, edi
ret
; IN: ebx = socket, OUT: eax = TH_FIN or 0
align 4
tcp_reasm_deliver:
xor ebp, ebp ; 'delivered anything' flag
cmp [ebx + TCP_SOCKET.t_state], TCPS_ESTABLISHED
jb .done
.loop:
mov edx, [ebx + TCP_SOCKET.seg_next]
test edx, edx
jz .done
mov eax, [edx + TCP_reasm_buff.Seq]
sub eax, [ebx + TCP_SOCKET.RCV_NXT]
jg .done ; hole before this node
neg eax ; node bytes below RCV_NXT
mov ecx, [edx + NET_BUFF.length]
sub ecx, eax ; fresh bytes in the node
jg .deliver
jl .unlink ; ends before RCV_NXT
cmp [edx + TCP_reasm_buff.Flags], 0
jne .fin
.unlink:
call tcp_reasm_free_head
jmp .loop
.deliver:
mov esi, [edx + NET_BUFF.offset]
add esi, edx
add esi, eax
lea eax, [ebx + STREAM_SOCKET.rcv]
call socket_ring_write ; -> ecx = bytes stored
test ecx, ecx
jz .done ; receive buffer full
add [ebx + TCP_SOCKET.RCV_NXT], ecx
mov ebp, 1
jmp .loop
.fin:
call tcp_reasm_free_head
cmp [ebx + TCP_SOCKET.seg_next], 0
jne .fin
mov eax, TH_FIN
jmp .notify
.done:
xor eax, eax
.notify:
test ebp, ebp
jz .nothing
push eax
mov eax, ebx
call socket_notify
pop eax
.nothing:
ret
; IN: ebx = socket (locked), OUT: eax destroyed
align 4
tcp_reasm_free_head:
mov eax, [ebx + TCP_SOCKET.seg_next]
push [eax + NET_BUFF.NextPtr]
pop [ebx + TCP_SOCKET.seg_next]
dec [ebx + TCP_SOCKET.seg_count]
lock dec [TCP_reasm_buffs]
stdcall net_buff_free, eax
ret