Lab: networking

In this lab you will write an xv6 device driver for a network interface card (NIC), and then write the receive half of an ethernet/IP/UDP protocol processing stack.

Fetch the xv6 source for the lab and check out the net branch:

  $ git fetch
  $ git checkout net
  $ make clean

Background

Before writing code, you may find it helpful to review "Chapter 6: Interrupts and device drivers" in the xv6 book.

You'll use a network device called the Tulip, or DEC 21143, to handle network communication. To xv6 (and the driver you write), the Tulip looks like a real piece of hardware connected to a real Ethernet local area network (LAN). In fact, the Tulip NIC your driver will talk to is an emulation provided by qemu, connected to a LAN that is also emulated by qemu. On this emulated LAN, xv6 (the "guest") has an IP address of 10.0.2.15. Qemu arranges for the computer running qemu (the "host") to appear on the LAN with IP address 10.0.2.2. When xv6 uses the Tulip NIC to send a packet to 10.0.2.2, qemu delivers the packet to the appropriate application on the host.

You will use QEMU's "user-mode network stack". QEMU's documentation has more about the user-mode stack here. We've updated the Makefile to enable QEMU's user-mode network stack and Tulip network card emulation.

The Makefile configures QEMU to record all incoming and outgoing packets to the file packets.pcap in your lab directory. It may be helpful to review these recordings to confirm that xv6 is transmitting and receiving the packets you expect. To display the recorded packets:

tcpdump -XXnr packets.pcap

We've added files to the xv6 repository for this lab. You should add your Tulip driver code to kernel/tulip.c. kernel/net.c and kernel/net.h contain a simple network stack that implements the IP, UDP, and ARP protocols; net.c has complete code for user processes to send UDP packets, but lacks most of the code to receive packets and deliver them to user space. Finally, kernel/pci.c contains code that searches for a Tulip card on the PCI bus when xv6 boots.

Part One: NIC

Your job is to implement a Tulip driver in kernel/tulip.c that can cause the Tulip to transmit and receive packets. You are done with this part when make grade says your solution passes the "txone" and "rxone" tests.

While writing your code, you'll refer to the Tulip Hardware Reference Manual. The following sections are particularly important, and you should read (or at least skim) them before writing code:

You should ignore Sections 2, 3.1, 5, 6, 7, 8, and the appendices.

The Tulip has control and status registers (CSRs) which software uses to initialize the Tulip and control its behavior; these are described in the manual's Section 3.2. The pci.c code we give you maps the Tulip's CSRs to addresses in kernel memory, and passes the address of CSR0 to tulip_init(). Your driver will need to use CSR1, CSR2, CSR3, CSR4, CSR5, CSR6, and CSR7, but not the others.

The Tulip uses DMA (direct memory access) to read outgoing packets from RAM, and to copy incoming packets into RAM so that software can process them. For efficient batching, the Tulip allows the driver to provide it with multiple packets to send, and multiple memory areas into which the Tulip can copy received packets. The driver describes each of these packet buffers with a transmit or receive descriptor holding the buffer's memory address, length, and status, as detailed in the Tulip manual's Section 4.2. The Tulip supports using descriptors as a ring, so that after sending or receiving into the last entry in a fixed-size array of descriptors, the Tulip goes back to the first descriptor.

Step One: Initialization

Modify tulip_init() in tulip.c to initialize the Tulip hardware. Here's a rough outline of what you'll need to do; the manual sections indicated above contain the details.

Define two C struct types that follow the layout of the transmit and receive descriptors in the manual's Figure 4-7 and Figure 4-2. The easiest approach is to have each struct contain four unsigned int's (for RDES0..3 and TDES0..3). Then define, as global variables, an array of transmit descriptors, and an array of receive descriptors. Arrays of size 4 will work. You won't need to use TDES3 or RDES3.

Initialize each descriptor in the receive descriptor array. Set RDES2 to a buffer returned by kalloc(). Set RDES1 to 1020 (the length). Set the OWN bit to 1.

Set CSR3 to the address of the receive descriptor array, and CSR4 to the address of the transmit descriptor array.

Set up CSR7 to generate Normal, Receive, and Transmit interrupts.

Set bit 25 in CSR6, as well as ST, PR, and SR.

Step Two: Transmit

When the network stack in net.c needs to send a packet, it calls tulip_xmit() with a pointer to a buffer that holds the packet to be sent; net.c allocates this buffer with kalloc(). Your transmit code must place a pointer to the packet data in a descriptor in the transmit ring. Read the manual's Section 4.2.2 to see how to place a packet in a transmit descriptor and indicate that the Tulip should send it.

tulip_xmit() should not block: it should return after placing the packet in the transmit DMA ring.

You must ensure that each buffer is eventually passed to kfree(), but only after the Tulip has finished transmitting the packet.

When you're done, you can test your transmit code as follows. Run python3 nettest.py txone in one window (outside of xv6), and in another window run nettest txone in xv6, which sends a single packet. If all goes well the two windows should look like:

$ nettest txone
txone: sending one packet
$ python3 nettest.py txone
tx: listening for a UDP packet
txone: OK

tcpdump -XXnr packets.pcap should then produce output much like this:

reading from file packets.pcap, link-type EN10MB (Ethernet)
21:27:31.688123 IP 10.0.2.15.2000 > 10.0.2.2.25603: UDP, length 5
        0x0000:  5255 0a00 0202 5254 0012 3456 0800 4500  RU....RT..4V..E.
        0x0010:  0021 0000 0000 6411 3ebc 0a00 020f 0a00  .!....d.>.......
        0x0020:  0202 07d0 6403 000d 0000 7478 6f6e 65    ....d.....txone

You should also test your driver's ability to send more than one packet by running python3 nettest.py tx32 on one window, and then running nettest tx32 inside xv6. In the first window you should see:

$ python3 nettest.py tx32
tx: waiting for 32 UDP packets...
tx: OK

Step Three: Receive

Your initialization code should have configured the Tulip to generate interrupts when it receives packets. These interrupts arrive on IRQ 33. Define an interrupt-handling function in tulip.c, and modify devintr() in trap.c to call your handler if the irq is 33.

Your interrupt handler should clean the interrupt status flags in CSR5. Because the Tulip may generate a single interrupt for multiple packets that arrive in a batch, your handler should scan your receive descriptor array looking for packets that the Tulip has DMA'd into memory. Section 4.2 describes what you'll find in the receive descriptors.

Your handler should pass each arrived packet to the network stack (in net.c) by calling net_rx(). net_rx() will kfree() the packet. The handler should then allocate a new buffer with kalloc() and place it into the descriptor, so that when the Tulip reaches that point in the receive ring again it finds a fresh buffer into which to DMA a new packet.

Your handler should set CSR2 to 1 at the end; see Section 3.2.2.3.

To test receive, start xv6 in one window, and then run python3 nettest.py rxone in another window. nettest.py rxone sends two packets to xv6: an ARP request, and then a UDP/IP packet. net.c contains the code to detect the ARP request and call tulip_xmit() to send an ARP reply. In the xv6 window you should see:

init: starting sh
$ arp_rx: received an ARP packet
ip_rx: received an IP packet
.
And in the outside window:
$ python3 nettest.py rxone
txone: sending one UDP packet

If all went well, tcpdump -XXnr packets.pcap should produce output like this:

reading from file packets.pcap, link-type EN10MB (Ethernet)
21:29:16.893600 ARP, Request who-has 10.0.2.15 tell 10.0.2.2, length 28
        0x0000:  ffff ffff ffff 5255 0a00 0202 0806 0001  ......RU........
        0x0010:  0800 0604 0001 5255 0a00 0202 0a00 0202  ......RU........
        0x0020:  0000 0000 0000 0a00 020f                 ..........
21:29:16.894543 ARP, Reply 10.0.2.15 is-at 52:54:00:12:34:56, length 28
        0x0000:  5255 0a00 0202 5254 0012 3456 0806 0001  RU....RT..4V....
        0x0010:  0800 0604 0002 5254 0012 3456 0a00 020f  ......RT..4V....
        0x0020:  5255 0a00 0202 0a00 0202                 RU........
21:29:16.902656 IP 10.0.2.2.61350 > 10.0.2.15.2000: UDP, length 3
        0x0000:  5254 0012 3456 5255 0a00 0202 0800 4500  RT..4VRU......E.
        0x0010:  001f 0000 0000 4011 62be 0a00 0202 0a00  ......@.b.......
        0x0020:  020f efa6 07d0 000b fdd6 7879 7a         ..........xyz
To test that your driver can receive many packets in a row, boot xv6 in one window, and in another window run python3 nettest.py rx32. In the xv6 window you should see:
init: starting sh
$ arp_rx: received an ARP packet
ip_rx: received an IP packet
................................
net.c prints a dot for each IP packet that the Tulip driver hands to net_rx(); there should be 32 of them.

Tulip Tips

You can ask qemu's Tulip emulator to print tracing (debugging) messages with the qemu --trace option. An easy way to do this is to add one or more of these lines to xv6's Makefile:

QEMUOPTS += --trace 'tulip_reg_write'
QEMUOPTS += --trace 'tulip_reg_read'
QEMUOPTS += --trace 'tulip_descriptor'
QEMUOPTS += --trace 'tulip_receive'
QEMUOPTS += --trace 'tulip_tx_state'
QEMUOPTS += --trace 'tulip_rx_state'
QEMUOPTS += --trace 'tulip_irq'
QEMUOPTS += --trace 'tulip_*'

You can move on to the next part of the lab once make grade shows that the first three tests pass:

$ make grade
...
== Test   nettest: txone == 
  nettest: txone: OK 
== Test   nettest: arp_rx == 
  nettest: arp_rx: OK 
== Test   nettest: ip_rx == 
  nettest: ip_rx: OK 

Part Two: UDP Receive

UDP, the User Datagram Protocol, allows user processes on different Internet hosts to exchange individual packets (datagrams). UDP is layered on top of IP. A user process indicates which host it wants to send a packet to by specifying a 32-bit IP address. Each UDP packet contains a source port number and a destination port number; processes can request to receive packets that arrive addressed to particular port numbers, and can specify the destination port number when sending. Thus two processes on different hosts can communicate with UDP if they know each other's IP addresses and the port numbers each is listening for. For example, Google operates a DNS name server on the host with IP address 8.8.8.8, listening on UDP port 53.

In this task, you'll add code to kernel/net.c to receive UDP packets, queue them, and allow user processes to read them. net.c already contains the code required for user processes to transmit UDP packets (with the exception of tulip_xmit(), which you provide).

Your job is to fill in the implementations of ip_rx(), sys_recv(), and sys_bind() in kernel/net.c. You are done when make grade says your solution passes all of the tests.

You can run the same tests that make grade runs by starting python3 nettest.py grade in one window, and then running nettest grade inside xv6 in another window. If all goes well, you should see this in the first (outside of xv6) window:

$ python3 nettest.py grade
txone: OK
rxone: sending one UDP packet
and this in the xv6 window:

$ nettest grade
txone: sending one packet
arp_rx: received an ARP packet
ip_rx: received an IP packet
ping0: starting
ping0: OK
ping1: starting
ping1: OK
ping2: starting
ping2: OK
ping3: starting
ping3: OK
dns: starting
DNS arecord for pdos.csail.mit.edu. is 128.52.129.126
dns: OK
free: OK

This lab's system-call API specification for UDP looks like this:

All the addresses and port numbers passed as arguments to these system calls, and returned by them, must be in host byte order (see below).

You'll need to provide the kernel implementations of the system calls, with the exception of send(). The program user/nettest.c uses this API.

To make recv() work, you'll need to add code to ip_rx(), which net_rx() calls for each received IP packet. ip_rx() should decide if the arriving packet is UDP, and whether its destination port has been passed to bind(); if both are true, it should save the packet where recv() can find it. However, for any given port, no more than 16 packets should be saved; if 16 are already waiting for recv(), an incoming packet for that port should be dropped. The point of this rule is to prevent a fast or abusive sender from forcing xv6 to run out of memory. Furthermore, if packets are being dropped for one port because it already has 16 packets waiting, that should not affect packets arriving for other ports.

The packet buffers that ip_rx() looks at contain a 14-byte ethernet header, followed by a 20-byte IP header, followed by an 8-byte UDP header, followed by the UDP payload. You'll find C struct definitions for each of these in kernel/net.h. Wikipedia has a description of the IP header here, and UDP here.

Production IP/UDP implementations are complex, handling protocol options and validating invariants. You only need to do enough to pass make grade. Your code needs to look at ip_p and ip_src in the IP header, and dport, sport, and ulen in the UDP header.

Pay attention to byte order. Ethernet, IP, and UDP header fields that contain multi-byte integers place the most significant byte first in the packet. The RISC-V CPU, when it lays out a multi-byte integer in memory, places the least-significant byte first. This means that, when code extracts a multi-byte integer from a packet, it must re-arrange the bytes. This applies to short (2-byte) and int (4-byte) fields. You can use the ntohs() and ntohl() functions for 2-byte and 4-byte fields, respectively. Look at net_rx() for an example of this when looking at the 2-byte ethernet type field.

If there are errors or omissions in your Tulip driver, they may only start to cause problems during the ping tests. For example, the ping tests send and receive enough packets that the descriptor ring indices will wrap around.

Some hints:

Submit the lab

Time spent

Create a new file, time.txt, and put in a single integer, the number of hours you spent on the lab. git add and git commit the file.

Answers

If this lab had questions, write up your answers in answers-*.txt. git add and git commit these files.

Submit

Assignment submissions are handled by Gradescope. You will need an MIT gradescope account. See Piazza for the entry code to join the class. Use this link if you need more help joining.

When you're ready to submit, run make zipball, which will generate lab.zip. Upload this zip file to the corresponding Gradescope assignment.

If you run make zipball and you have either uncommitted changes or untracked files, you will see output similar to the following:

 M hello.c
?? bar.c
?? foo.pyc
Untracked files will not be handed in.  Continue? [y/N]
Inspect the above lines and make sure all files that your lab solution needs are tracked, i.e., not listed in a line that begins with ??. You can cause git to track a new file that you create using git add {filename}.

Optional Challenges:

Some of these challenges are intended to increase performance in ways that may not be apparent or measurable under QEMU.

If you pursue a challenge problem, whether it is related to networking or not, please let the course staff know!


Questions or comments regarding 6.1810? Send e-mail to the course staff at 61810-staff@lists.csail.mit.edu.

Creative Commons License