IO pressure stall – unexplained is the most annoying kind because the system looks fine, but containers just freeze for seconds at a time. There are only four things it usually is: ZFS, an LXC I/O limit, the storage backend, or the disk itself. Here’s how to tell which one without guessing.

Start with PSI, before anything else
Pressure Stall Information tells you whether the host is actually hurting, or whether the problem is inside one container. Run this on the Proxmox node:
cat /proc/pressure/io
You get three numbers — avg10, avg60, avg300. If avg10 is above 40-50%, the host itself is stalling on I/O and the cause is somewhere below: ZFS, the backend, or the hardware. If PSI is low and a container is still slow, the limit is inside that container and you can skip straight to the LXC section.
This is the one check that splits the problem in half, so do it first.
Check if ZFS is the culprit
ZFS loves to stall IO when it’s doing a scrub, resilver, or just handling a ton of sync writes. Run zpool status and look for any scrub or resilver in progress. If there is one, that’s probably your stall. You can pause it with zpool scrub -p yourpool and see if the stalls stop.
Sync writes are another classic. If you have a database or something doing a lot of fsync, ZFS can choke if your SLOG device is slow or nonexistent. Check zpool iostat -v 1 and look for high latency on the pool. If you see huge numbers, you might need a faster SLOG or to set sync=disabled on datasets where you don’t care about data loss on power failure (like a media library).
LXC IO limits
If you’re running containers, check if you set any IO limits on them. In Proxmox, you can set IO limits on the container’s disk. If a container hits that limit, it’ll stall IO for everything else on that storage. The fastest way to look:
grep -E '^(lxc.cgroup2.io|limits)' /etc/pve/lxc/CTID.conf
If you see lxc.cgroup2.io.max with a small number, that’s your culprit — the default is unlimited, so if it’s there, someone set it. On older configs the same thing shows up as iops or mbps lines in /etc/pve/lxc/<CTID>.conf. Raise it or remove it temporarily and see if the stalls go away.
Also check the container’s cgroup IO stats. Run cat /sys/fs/cgroup/blkio/lxc/<CTID>/blkio.throttle.io_service_bytes and see if any container is hitting a limit. If so, that’s your stall.

Is it the NFS or TrueNAS backend?
If host PSI is high and there are no LXC limits, the storage underneath is the next suspect. On the Proxmox node:
iostat -x 1
Watch the %util column for the disk backing your NFS mount. If it’s pegged at 100% and await is high, the storage is saturated. Then go and look at the NAS itself: is the pool healthy, are you on SMR drives, and is sync set to always on the dataset? NFS with sync writes on a slow pool will absolutely tank performance, and it looks exactly like a host problem from the Proxmox side.
Check the network between the two as well. A bad cable or a flapping switch port causes retransmits that present as I/O stalls and nothing else. iftop or nload on both ends will show packet loss quickly. And if you are running unprivileged containers against a NAS share, UID/GID mapping can produce its own weird I/O behaviour on top of all this.
Other things to check
Sometimes it’s none of the above. Check dmesg for any storage errors, like SATA link resets or NVMe controller issues. If you see those, it’s a hardware problem. Also check iostat -x 1 on the host and look for devices with high await or %util. If one disk is maxed out, that’s your bottleneck.
If you’re using a single SSD for everything, like in a Jellyfin setup with one SSD, you might just be hitting the disk’s limits. IO pressure stalls are often just the disk being too slow for what you’re asking it to do.
My take
Read PSI first — it tells you whether to look at the host or inside a container, and it takes one command. After that, most of the time it’s a stupid LXC limit someone set and forgot, or sync writes on NFS. Fix those before you start blaming the hardware. If you’re on a single disk, consider adding a second one or moving some workloads to a different storage. Don’t overthink it — just check the obvious stuff first.
Part of Proxmox in a homelab — the backup and storage section.