FreeBSD / bhyve / PCI passthrough

Passthrough reset lifecycle

proposed design and investigation readiness poll submitted upstream, pending review guardrails and config restore proposed, on a local branch - not landed 2026-09

Handing a physical PCI function to a guest and taking it back are the two moments where an unreset or half-configured device can corrupt the host. bhyve's ppt driver and the generic PCI layer already reset a device around assignment, but the reset is only as safe as the state it leaves behind. This page proposes four guardrails that tighten that path, and shows the one place where the missing guardrail is not just an attach failure but a host killer.

None of this has landed. The post-reset readiness poll is submitted upstream and awaiting review; the config-space restore and the ppt checks live on a local branch. The drawings describe the intended behaviour, not the current tree.

The assign path, guardrail by guardrail

host / vmm vm_assign_pptdev vmm / ppt ppt_assign_device device function config 1 map domain create + install gpa -> hpa 2 save config pci_save_state BARs, command 3 quiesce clear busmaster wait idle <= 1s 4 FLR issue + await config cleared 5 restore pci_restore_state BARs, command 6 add + arm iommu_add_device busmaster LAST The invariant: bus mastering is enabled only after step 6. Translations are installed before the device is assigned (step 1, before assign), and the reset (step 4) happens while mastering is off. So the device cannot issue a transaction into a domain that has no mappings yet, nor with descriptor rings left over from its previous owner. Teal steps are the IOMMU-mapping guardrail; the plain steps are ppt's save and restore; the rust step is the reset itself, the only point where config space is destroyed.
Order is the whole guardrail. The domain is created and populated before ppt_assign_device runs, the device is quiesced and reset before its config is restored, and PCIM_CMD_BUSMASTEREN is set last of all. Unassign runs the mirror image: clear the decode and mastering bits, save, reset, restore, then remove the device from the domain, so a device handed back to the host carries none of the guest's state.

The reset escalation ladder

wider blast radius, harder recovery 1 Function-Level Reset (FLR) pcie_flr(): clear busmaster, wait idle, set PCIER_DEVICE_CTL FLR bit 0x8000 needs PCIEM_CAP_FLR - scope: one function 2 Secondary bus reset bridge ctl PCIR_BRIDGECTL_1 (0x3e), bit PCIB_BCR_SECBUS_RESET (0x40); or link disable/retrain - scope: the whole bus 3 Parent-bridge / fabric reset detach bridge, pci_power_reset it, re-enumerate the subtree scope: every device below the bridge FLR absent, or issued but ignored bus reset did not recover On assign, if FLR is unavailable ppt tries a D3/D0 power reset - a no-op on a No_Soft_Reset (0x0008) function - so it refuses (EOPNOTSUPP) rather than hand a guest an unreset device.
You climb the ladder only when the rung below fails. pcie_flr() returns false when a function has no FLR capability, and a function can also advertise FLR yet leave itself unchanged after one - capability is not efficacy. Either way the narrow reset did not take, so recovery escalates to the bridge's secondary bus reset and, failing that, to detaching and power-cycling the parent bridge and re-enumerating its subtree. Each step resets more of the machine, which is why the cheapest reset that actually works is the one to reach for.
rungmechanismscopeescalate when
Function-Level Resetpcie_flr(): quiesce, then set the FLR bit in the PCIe Device Control registerone functionno FLR capability, or the FLR is issued but the function does not reset
Secondary bus resetbridge control bit PCIB_BCR_SECBUS_RESET, or downstream link disable and retraineverything on the secondary busthe function is still wedged after a function reset
Parent-bridge / fabric resetdetach the bridge, pci_power_reset it, re-enumerate the subtreethe whole bridge subtreelast resort before a host reboot; recovers a subtree without one

The config-restore defect on the detach path

device_detach -> pci_child_detached saves config: BARs into cfg.maps pcie_flr() resets the function BARs read 0x00000004, decode bits cleared as shipped: no restore proposed fix device_probe_and_attach runs against zeroed config space outcome depends on the card: aq 10G NIC: attach fails, "unknown RBL status 0xffff" Intel I350: fatal PCIe error, host wedged, power cycle to recover pci_restore_state(child) rewrites BARs + command register first device_probe_and_attach BARs preserved, decode restored, device reattaches cleanly and passes traffic - not merely enumerable
Detaching saves the state the reset then clears, but nothing writes it back. On the DEVF_RESET_DETACH path of pci_reset_child the FLR-success branch reattaches the child against config space the reset just zeroed. The saved bases survive in cfg.maps, so the fix is one call - pci_restore_state(child) before device_probe_and_attach. The plain devctl reset (suspend path) is unaffected, because pci_resume_child already restores. Proven on real hardware across two device classes, with opposite severities.

What is proposed, and how well it is proven

These are separate changes that share a code path, not one patch. They are stated here at the honesty the evidence supports: two are proven on hardware or by fault injection, one rests on code reading alone, and none has landed. The config restore is the one with a hardware A/B behind it and the one that turns a corrupted host into a recoverable one; recovery itself needs no reboot, either by escalating the reset up the ladder to the parent bridge or by writing the saved BARs and command register back and reattaching.

guardrailstatusevidence
Map the IOMMU domain before assigning the deviceproposed, not landedreasoned from a recorded IO_PAGE_FAULT trace; the fault it prevents was not reproduced on hardware here
Quiesce the device before resetting it, extending the wait to 1000 msproposed, not landedproven by fail(9) injection against a control; no device on hand held transactions pending naturally
Poll for the function to answer config reads after a resetsubmitted upstream, pending reviewproven by fail(9) injection; matches an independent upstream bug report against real hardware
Restore config state after resetting a detached childproposed, on a local branchproven on hardware: failed attach on an aq 10G NIC, host-killing fatal PCIe error on an Intel I350, both fixed
Refuse to assign a device that cannot be resetproposed, not landedrefusal path proven on a card with no FLR that sets No_Soft_Reset; over-refusal checked against an FLR-capable card

Source: sys/dev/pci/pci.c (pcie_flr, pci_reset_child, pci_power_reset), sys/dev/pci/pci_pci.c (pcib_reset_child), sys/amd64/vmm/io/ppt.c (ppt_pci_reset, ppt_assign_device), sys/amd64/vmm/vmm.c (vm_assign_pptdev, vm_iommu_map).