Proxmox VE now ships an arm64 build. We measured how it performs on an inexpensive ARM box, and how fast the box can go. NVMe/TCP reads make a good test because they lean on the NIC, PCIe, the I/O memory management unit (IOMMU) and the CPUs at the same time. Getting Proxmox to boot and stay up on this box had a few wrinkles of its own, which we wrote up on the Proxmox forum: the black screen after initrd, the hard reset every 20 minutes, and the firmware power cap. All measurements were taken on the Proxmox host itself, with no VMs involved. Guest performance is a separate question that we did not test.

The setup was simple: one GX10, one NVMe/TCP target, and two 200G cables between them.

  • Host: ASUS GX10, NVIDIA GB10 (the DGX Spark SoC), 10 Cortex-X925 + 10 Cortex-A725 cores, one NUMA node
  • OS: Proxmox VE 9.2 arm64, stock kernel 7.0.14-6-pve
  • NIC: ConnectX-7, dual-port 200G, in-tree mlx5 driver, MTU 9000
  • Target: Blockbridge NVME20 (AMD EPYC Zen 5), dual-port 200G

With the kernel NVMe/TCP initiator the box reads 28.1 GB/s of 1 MiB random IO. With SPDK, the userspace initiator from the Storage Performance Development Kit, it reads 28.9 GB/s. By our math the two Gen5 x4 PCIe links behind the ConnectX-7 top out at about 29.9 GB/s of NVMe payload, so those results are 94% and 97% of what the hardware can do.

Getting there took two kernel parameters. On a vanilla Proxmox installation the kernel initiator read 6.3 GB/s. The first parameter took it to 24.4 GB/s and the second to 26.5 GB/s.

The charts below show the climb, first for the kernel initiator and then for SPDK. Each bar keeps the changes above it and adds one more. The blue bar is where we ended up.

The kernel initiator gets almost all of its gain from the first change. SPDK starts much higher, and the second change does a little more for it than the first. The last row of the table is the calculated limit of the two PCIe links.

Step Change Kernel initiator SPDK
0 Proxmox defaults: strict IOMMU, MaxPayload 128 6.3 GB/s 21.3 GB/s
1 iommu.strict=0 24.4 GB/s 24.7 GB/s
2 pci=pcie_bus_perf (MaxPayload 512) 26.5 GB/s 28.8 GB/s
3 LRO on 26.8 GB/s 28.9 GB/s
4 Fewer reads in flight, IRQs pinned 28.1 GB/s 28.9 GB/s
  Best single run in the set (SPDK, step 2)   29.2 GB/s
  Both PCIe links, calculated 29.9 GB/s 29.9 GB/s

Reads in flight means the total across all four namespaces. Steps 0 to 3 ran at 160 in flight for the kernel initiator and 40 for SPDK, and step 4 drops those to 40 and 20. Step 4 also pins the NIC interrupts to the X925 cores, which on its own is worth less than 1%.

The table shows averages. On a good run, SPDK on 12 cores (ten X925 and two A725) with pause on and two reads in flight per core reads 29.3 GB/s of NVMe payload, which is 29.5 GB/s on the PCIe links. We say “on a good run” because it does not happen every time. The NIC spreads the TCP connections across its receive queues by hashing them, and some spreads are better than others, so you may need a few tries to land there. That is also what we think costs the 2% in the link table below.

Here is one of those runs, caught on the dashboard we kept up while testing. It reads the NIC counters, the PCIe link math and the SPDK output in one place, and it is why we could see exactly where the ceiling was. Click through for the full-size version.

Dashboard of the GX10 during a 29.3 GB/s SPDK run: 237 Gbit/s on the network, 29.5 GB/s on PCIe, 29.3 GB/s of NVMe payload

Summary

  1. Take the IOMMU out of strict mode with iommu.strict=0. Strict is the kernel default on arm64. On this box it holds the kernel NVMe/TCP initiator to about 6 GB/s, and nothing else you tune will change that.
  2. Add pci=pcie_bus_perf to the kernel command line. The firmware leaves the NIC’s PCIe links at a 128-byte MaxPayload, and that costs about 10% of each link.
  3. The ConnectX-7 hangs off two independent Gen5 x4 links, and each physical function (PF) sits on one of them. Put an IP on all four PFs, each on its own subnet, and connect a controller through each, so that both links and both ports carry traffic.
  4. Set arp_ignore=1 and arp_announce=2. Two PFs share each physical port, and with the Linux defaults one of them can quietly take the other’s traffic.
  5. Keep the number of reads in flight low for 1 MiB IO. About 40 suits the kernel initiator and 20 suits SPDK. Going deeper costs the kernel initiator 3 to 5%.
  6. Turn large receive offload (LRO) on. It is worth about 2% to the kernel initiator. Leave flow-control pause at its default, which is on.
  7. Check your work: lspci -vv should show MaxPayload 512 on the NIC, and rx_bytes in ethtool -S should climb on every PF under load.

How We Measured

All of the numbers come from one set of 352 scripted runs, taken one change at a time across six boots. Unless we say otherwise, a figure is the average of at least two runs, and runs of the same configuration usually agreed within 1%. The script was designed to refuse to start a run if the kernel command line, IOMMU mode, MaxPayload or NIC settings were not what that run expected.

We only measured reads. Reads are the expensive direction for the host, since every byte is copied out of socket buffers on the way in. We did not measure from inside a VM.

A single PCIe link ran at 99% of its calculated capacity, and the pair ran at 97%, meaning that the target was not the limit.

How the ConnectX-7 Is Attached

You might expect the ConnectX-7 dual-port 200G NIC to sit on one x16 link. On the GB10 it sits on two independent Gen5 x4 root ports, and those links are what limit it, not the ports. Each PF has its own MAC, and lspci -tv shows the card as two dual-port PCI devices:

-[0000:00]---00.0-[01-0f]--+-00.0  Mellanox MT2910 [ConnectX-7]   <- nic2, physical port 0
                           \-00.1  Mellanox MT2910 [ConnectX-7]   <- nic4, physical port 1
-[0002:00]---00.0-[01-0f]--+-00.0  Mellanox MT2910 [ConnectX-7]   <- nic5, physical port 0
                           \-00.1  Mellanox MT2910 [ConnectX-7]   <- nic6, physical port 1

0000:00:00.0 PCI bridge: NVIDIA GB10 GEN5 X4 PCIe host
0002:00:00.0 PCI bridge: NVIDIA GB10 GEN5 X4 PCIe host

From here on we call the two root ports half A and half B. Half A is PCI domain 0000. It is one Gen5 x4 link and carries 0000:01:00.0 (nic2) and 0000:01:00.1 (nic4). Half B is PCI domain 0002. It is the other Gen5 x4 link and carries 0002:01:00.0 (nic5) and 0002:01:00.1 (nic6). Each half has its own bandwidth into memory. Function .0 on each half is QSFP port 0 and function .1 is port 1, so every physical port has one PF on each half. Every PF has its own MAC, so a packet’s destination MAC picks its half.

  QSFP port   PF (netdev)            PCIe half  root port     link              host

  port 0 -->  nic2  0000:01:00.0 --+
  port 1 -->  nic4  0000:01:00.1 --+--> half A  0000:00:00.0  Gen5 x4  ~15 GB/s --+
                                                                                  +--> GB10 SoC
  port 0 -->  nic5  0002:01:00.0 --+                                              |    memory
  port 1 -->  nic6  0002:01:00.1 --+--> half B  0002:00:00.0  Gen5 x4  ~15 GB/s --+

  Each physical port appears twice: it has one PF on each half.

To see what this means in practice, we connected namespaces through different PFs and measured each layout with SPDK. In the chart, the first two bars are the same length: a second PF on the same half has no more PCIe to use. The third bar is longer because the second PF is on the other half, and it stops where it does because both PFs are on one 200G port. The last bar uses both ports and both links, and at 28.9 GB/s it is within 4% of what the two Gen5 x4 links can carry.

The table below adds the kernel initiator.

Namespaces connected through Limit Kernel initiator SPDK
nic2 only (half A) one Gen5 x4 link 12.7 GB/s 14.8 GB/s
nic2 and nic4 (both on half A) the same Gen5 x4 link 14.5 GB/s 14.8 GB/s
nic2 and nic5 (one per half, both on port 0) one 200G port 17.9 GB/s 24.7 GB/s
All four PFs both Gen5 x4 links 28.1 GB/s 28.9 GB/s

The SPDK column is the one that shows what the links do; the kernel rows with fewer namespaces also run fewer fio jobs. One link on its own reaches 14.8 GB/s, 99% of what it can carry, and for SPDK a second PF on the same half adds nothing. To use both links you need a PF from each half, and to use both ports as well you need all four.

Our final layout is four subsystems with one namespace each. Every subsystem has exactly one NVMe/TCP controller, and there is one controller per PF, so two controllers enter through each PCIe half. Nothing in this technote uses multipathing. We connect each controller from the address of the netdev it should use, for example -w 10.0.0.2 for nic2 and -w 10.0.1.2 for nic5. Each PF is on its own subnet, so routing picks the interface.

A single namespace with a controller on each half and NVMe native multipath should work as well, though we did not test it in these runs. If you try it, change the iopolicy. By default it is numa, which on an SoC with one NUMA node picks one path and stays there:

echo round-robin > /sys/class/nvme-subsystem/nvme-subsysN/iopolicy

or make it permanent with options nvme_core iopolicy=round-robin in /etc/modprobe.d/.

Two PFs on One Port Answer Each Other’s ARP

nic2 and nic5 share physical port 0, so both of them see every broadcast that arrives on it. With the Linux default of arp_ignore=0, an interface will answer ARP for any address the host owns. Both PFs answer a request for nic5’s address, and the target keeps whichever reply gets there first.

On one boot the target learned nic2’s MAC for nic5’s address. Nothing looked wrong. All four controllers were connected and all four namespaces were reading. But nic5 was receiving nothing, and both of the port 0 namespaces were coming in through half A. The kernel accepts the misdirected packets because the host owns the address, and does not care which interface they arrived on. Whether you hit this depends on which PF answered first, so it can come and go from one boot to the next.

You can only see it in the per-PF rx_bytes counter from ethtool -S. The rx_bytes_phy counter is per physical port, so it looks fine. The fix is two sysctls that make each interface answer only for its own address:

# /etc/sysctl.d/90-arp-per-interface.conf
net.ipv4.conf.all.arp_ignore = 1
net.ipv4.conf.all.arp_announce = 2

The target’s ARP cache will still hold the wrong entry after you apply them. Clear it on the target, or send a gratuitous ARP from each PF. Then check that all four PFs are receiving.

A Gen5 lane runs at 32 GT/s. Four lanes give 128 Gbit/s. After 128b/130b encoding, that is 15.75 GB/s. The NIC writes received data into host memory as PCIe packets that hold up to MaxPayload bytes each, with 20 bytes of header and framing around every one. At the firmware’s 128-byte MaxPayload that is 86% efficient. At 512 bytes it is 96%. Completion records, interrupts and link maintenance take roughly another 1%. What remains is about 15.0 GB/s of Ethernet frames per link. Not all of those frame bytes are data. Each frame carries Ethernet, TCP and NVMe/TCP headers, and with large receive offload (LRO) on they add up to about 0.4%. That leaves 14.95 GB/s of NVMe payload per link, or about 29.9 GB/s for the two.

The table below puts those calculated limits next to what we measured. The measured column is NVMe payload as the initiator counts it: fio’s bandwidth for the kernel initiator, and the total that spdk_nvme_perf prints for SPDK.

  Calculated Measured Measured / calculated
One link, SPDK (nic2 only) 14.95 GB/s 14.8 GB/s 99%
Both links, SPDK 29.9 GB/s 28.9 GB/s 97%
Both links, kernel initiator 29.9 GB/s 28.1 GB/s 94%

One link on its own runs at 99% of the math. With both links loaded, each carries about 2% less than it does alone. The likely cause is uneven receive-queue hashing: with 80 connections spread over four PFs, some queues get more flows than others, and the busiest one paces the rest. We did not chase it.

The Kernel Command Line

You need one parameter for the IOMMU and one for the PCIe links. Add them to the kernel command line and reboot. Afterwards, cat /proc/cmdline should contain:

quiet console=tty0 iommu.strict=0 pci=pcie_bus_perf

iommu.strict=0

On arm64 the kernel defaults to strict IOMMU mode, and the Proxmox kernel keeps that default (CONFIG_IOMMU_DEFAULT_DMA_STRICT=y). In strict mode every DMA unmap waits for the IOTLB, the IOMMU’s translation cache, to be invalidated.

This one setting matters more than anything else we changed. The table below shows both IOMMU modes at both MaxPayload sizes, so you can see each one’s effect on its own:

IOMMU mode MaxPayload Kernel initiator SPDK
strict (default) 128 (default) 6.2 GB/s 22.8 GB/s
strict 512 6.2 GB/s 28.4 GB/s
lazy (iommu.strict=0) 128 25.1 GB/s 26.2 GB/s
lazy 512 28.1 GB/s 28.9 GB/s

Every row here was run with LRO on and the reads in flight already tuned, so the strict and MaxPayload 128 rows come out a little higher than the same steps in the intro table, which had neither.

In strict mode the kernel initiator reads 6.2 GB/s whatever you do with MaxPayload or pause. Its 99th percentile latency is above 60 ms, where lazy mode gives about 3 ms. A profile shows where the time goes: arm_smmu_cmdq_issue_cmdlist takes 57% of all cycles on the X925 cores. The SMMU is ARM’s system MMU, the IOMMU on this SoC, and it has a single command queue. Every invalidation goes through that queue, and the cores line up behind it.

Strict mode costs SPDK much less: under 2% at MaxPayload 512. Note the strict rows of the SPDK column, though: raising MaxPayload alone takes it from 22.8 to 28.4 GB/s. Strict mode and small PCIe packets compound, and we did not work out why. If you run SPDK, MaxPayload is the bigger of the two parameters. As for why SPDK suffers less than the kernel initiator overall, one difference we can measure is on the transmit side. In the lazy, MaxPayload 512 runs SPDK sends about 8 thousand packets per GB read, and the kernel initiator sends about 34 thousand.

We also tried an identity mapping, iommu.passthrough=1, which skips translation entirely. It was no better than lazy mode: about 1% lower for both initiators, on a different boot.

pci=pcie_bus_perf

The firmware leaves both of the NIC’s root ports at a 128-byte MaxPayload. Linux’s default policy only matches an endpoint to its parent, so the ConnectX-7 runs at 128 as well, even though both ends support 512:

# before
0000:00:00.0  DevCtl: MaxPayload 128 bytes   (root port)
0000:01:00.0  DevCtl: MaxPayload 128 bytes   (ConnectX-7 PF)
# after pci=pcie_bus_perf
0000:00:00.0  DevCtl: MaxPayload 512 bytes
0000:01:00.0  DevCtl: MaxPayload 512 bytes

Every PCIe packet carries the same 20 bytes of header and framing, so a 128-byte packet spends a much bigger share of the link on framing than a 512-byte one does. That is the difference between 13.5 GB/s and 15.0 GB/s per link. In the table above, going to 512 took SPDK from 26.2 to 28.9 GB/s in lazy mode, and the kernel initiator from 25.1 to 28.1 GB/s.

NIC Settings

LRO

Hardware LRO is worth about 2% with the kernel initiator and nothing measurable with SPDK. We turn it on for all four PFs, but it is not a big deal if you skip it. The setting is lost on reboot and on driver reload, so put it in the interface configuration. ifupdown2 has a keyword for it, and ifreload -a reapplies it after a manual driver rebind.

auto nic2
iface nic2 inet static
    address 10.0.0.2/24
    mtu 9000
    lro-offload on

On Proxmox, check that it really applied. The kernel force-disables LRO on any netdev that is enslaved to a bridge, and whenever net.ipv4.ip_forward=1. In mlx5, LRO and rx-gro-hw are mutually exclusive.

Pause

Leave flow-control pause on, which is the default. We expected pause off to be faster, so we ran a full depth sweep both ways. For the kernel initiator, pause on was as fast or faster at every depth. For SPDK the two are within noise of each other. We did not run SPDK deeper than 160.

Reads in flight Kernel, pause on Kernel, pause off SPDK, pause on SPDK, pause off
20 26.4 GB/s 26.1 GB/s 28.9 GB/s 28.3 GB/s
40 28.1 GB/s 27.3 GB/s 28.9 GB/s 28.9 GB/s
80 27.5 GB/s 26.5 GB/s 28.8 GB/s 28.4 GB/s
160 26.8 GB/s 25.3 GB/s 27.8 GB/s 28.2 GB/s
320 27.0 GB/s 24.6 GB/s    
640 27.2 GB/s 24.4 GB/s    

With pause off the NIC drops frames when it runs out of room, and every dropped frame waits for TCP to retransmit it. The slowest 1 MiB read in a run was in the hundreds of milliseconds with pause off and in the tens with pause on.

How many reads you keep in flight matters more than either NIC setting. More is not better here. With pause on, the kernel initiator does best with 40 reads outstanding across the four namespaces, and from 160 up it is 3 to 5% below that. SPDK does best with 20.

Kernel Initiator vs SPDK

The kernel NVMe/TCP initiator reaches 28.1 GB/s on all 20 cores, driven by fio with libaio, 20 jobs at iodepth 2. SPDK’s spdk_nvme_perf reaches 28.9 GB/s.

SPDK’s edge does not come from avoiding the data copy. Both initiators copy received data once, from socket buffers into the application’s pages, and in the profiles that copy is the biggest single item for both: 29% of X925 cycles with the kernel initiator and 23% with SPDK. The NIC still interrupts and TCP still runs in softirq under both. What SPDK does differently is poll its sockets from user space instead of sleeping on completions, and submit straight from the application to the NVMe queue pair, which skips the block layer. We think that is where the difference comes from, but we did not measure it.

The GB10 has two kinds of cores, and for this work they are not equal. The ten Cortex-X925 cores are the fast ones, and the ten Cortex-A725 cores are the efficient ones. To see how much each kind contributes, we confined the initiator to a set of cores with fio’s cpus_allowed or SPDK’s core mask, and measured each set.

The first five bars add X925 cores two at a time. The sixth is the ten A725 cores on their own, and the last is all 20.

The table adds a per-core figure:

Cores running the initiator Kernel initiator GB/s per core SPDK GB/s per core
2 X925 8.2 GB/s 4.08 10.9 GB/s 5.47
4 X925 13.4 GB/s 3.34 17.2 GB/s 4.30
6 X925 18.1 GB/s 3.02 22.6 GB/s 3.77
8 X925 21.4 GB/s 2.67 26.5 GB/s 3.31
10 X925 24.4 GB/s 2.44 28.3 GB/s 2.83
10 A725 18.0 GB/s 1.80 26.0 GB/s 2.60
All 20 28.1 GB/s 1.41 28.9 GB/s 1.44

These rows count only the cores the application runs on, not the total CPU cost, because NIC interrupts and TCP receive work stay spread over all ten X925 cores in every row, the A725 rows included. Fewer cores need more reads outstanding per core: the SPDK rows below 20 cores keep two per core, and with one per core the ten X925 cores read only 24.3 GB/s.

The ten X925 cores get SPDK within 2% of all 20. The kernel initiator on the same cores is 13% short.

Reproducing the Result

Here is what we ran for the two headline numbers. The host is the GX10 on Proxmox VE 9.2 arm64 with the stock kernel, plus the command line, sysctls and NIC settings described above. The target can be any NVMe/TCP target able to source 29 GB/s of reads. Ours exported four subsystems with one namespace each. They listened on 10.0.0.3, 10.0.1.3, 10.0.2.3 and 10.0.3.3, port 4420, and accepted our host NQN (NVMe Qualified Name) from all four subnets.

Connecting the Namespaces

Each PF has its own subnet, as in the interface stanza earlier. nic2 is 10.0.0.2, nic5 is 10.0.1.2, nic4 is 10.0.2.2 and nic6 is 10.0.3.2. We connect one controller per PF, from that PF’s address, to the portal on the same subnet. By default a controller gets one IO queue per online CPU, so each one opens 20 TCP connections for IO and the four together open 80. The connections do not survive a reboot.

for n in 0 1 2 3; do
  nvme discover -t tcp -a 10.0.$n.3 -s 4420 -w 10.0.$n.2
done
nvme connect -t tcp -a 10.0.0.3 -s 4420 -w 10.0.0.2 -n <NQN on 10.0.0.3>
nvme connect -t tcp -a 10.0.1.3 -s 4420 -w 10.0.1.2 -n <NQN on 10.0.1.3>
nvme connect -t tcp -a 10.0.2.3 -s 4420 -w 10.0.2.2 -n <NQN on 10.0.2.3>
nvme connect -t tcp -a 10.0.3.3 -s 4420 -w 10.0.3.2 -n <NQN on 10.0.3.3>
nvme list-subsys

Start a read against all four namespaces and confirm that rx_bytes is climbing on all four PFs:

for n in nic2 nic4 nic5 nic6; do ethtool -S $n | grep -w rx_bytes; done

The Kernel Run

This job file produced the 28.1 GB/s. It floats 20 jobs over all 20 cores, five per namespace at iodepth 2, which makes 40 reads in flight. Do not put the four devices in one filename list. If you do, fio round-robins reads across them equally, and one slow namespace sets the pace for all four. Use nvme list-subsys to map the device names to your controllers.

[global]
ioengine=libaio
direct=1
rw=randread
bs=1M
iodepth=2
runtime=30
time_based
group_reporting
norandommap=1
randrepeat=0
cpus_allowed=0-19
numjobs=5

[nic2]
filename=/dev/nvme1n1
[nic4]
filename=/dev/nvme2n1
[nic5]
filename=/dev/nvme3n1
[nic6]
filename=/dev/nvme4n1

The SPDK Run

We built SPDK v25.05 from source with its default configuration. SPDK needs hugepages, and on this box there are none after a boot.

echo 2048 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages

This command produced the 28.9 GB/s. -c 0xFFFFF uses all 20 cores, and spdk_nvme_perf pairs each core with one namespace in turn. -q 1 keeps one read outstanding per core, which makes 20 in flight. -P 2 spreads each core’s reads over two queue pairs. List the namespaces alternating halves, A, B, A, B, as below. perf assigns them to cores round-robin, and with ten cores an A, A, B, B order puts six cores on one half. To run on the ten X925 cores alone, use -c 0xF83E0 and -q 2. Multiply the total MiB/s that perf prints by 1.048576 to get the payload figure in MB/s.

HOSTNQN=$(cat /etc/nvme/hostnqn)
/root/spdk/build/bin/spdk_nvme_perf -q 1 -P 2 -o 1048576 -w randread -t 25 -c 0xFFFFF \
  -r "trtype:TCP adrfam:IPv4 traddr:10.0.0.3 trsvcid:4420 subnqn:<NQN 0> hostnqn:$HOSTNQN" \
  -r "trtype:TCP adrfam:IPv4 traddr:10.0.1.3 trsvcid:4420 subnqn:<NQN 1> hostnqn:$HOSTNQN" \
  -r "trtype:TCP adrfam:IPv4 traddr:10.0.2.3 trsvcid:4420 subnqn:<NQN 2> hostnqn:$HOSTNQN" \
  -r "trtype:TCP adrfam:IPv4 traddr:10.0.3.3 trsvcid:4420 subnqn:<NQN 3> hostnqn:$HOSTNQN"

Do not raise -q above 8 with 1 MiB reads. spdk_nvme_perf splits large reads into child requests, and each queue pair has a fixed pool of requests. When the pool runs dry, perf prints starting I/O failed: -12 and the result drops.

CONCLUSION

Out of the box, the kernel initiator read 6 GB/s, and the IOMMU gated it entirely. One kernel parameter took it to 24 GB/s. A second one raised the PCIe ceiling. The rest was working out how the NIC is wired. The GB10 presents a ConnectX-7 connected by two x4 PCIe links. These links are effectively shared across the two QSFP ports. With traffic through all four PFs, the kernel initiator reads 28.1 GB/s and SPDK reads 28.9 GB/s, against a calculated 29.9 GB/s for the two links.

All of this ran on the stock Proxmox VE kernel with the in-tree mlx5 driver. We did not build or patch anything. No vendor driver stack was installed. The two settings that mattered were kernel command-line parameters. The kernel’s own NVMe/TCP initiator, driven by fio, came within 6% of what the PCIe links can carry. Proxmox on ARM works and is looking pretty awesome.

Not bad for a compact desktop built to run AI models.

ADDITIONAL RESOURCES