raft / docs
Guides

Troubleshooting

Diagnostic procedures and solutions for common deployment, configuration, and runtime issues.

Controller diagnostics

Preflight checks with raft doctor

Run raft doctor on your controller to inspect host health across configured locations:

CONTROLLER: Run environment health check
raft doctor

raft doctor connects to each host over SSH and reports:

  • Incus version running in project raft
  • Active status of host systemd services (incus, raft-expire.timer, raft-network.service)
  • Host root filesystem disk usage (df -h /)
  • Status and usage of the raft-data Btrfs pool
  • Count of existing saved containers

Doctor scope

raft doctor reports runtime daemon and storage status. It does not validate image fingerprints, inspect nftables bridge rules, enforce admission caps, or emit 5 GiB low-disk warnings.

Listing active and stopped containers

Inspect workspace states across locations:

CONTROLLER: List workspaces across all hosts
raft list

# Filter by a single location
raft list --location lab

If a host is unreachable over SSH, the CLI reports a connection failure. An unreachable host does not mean containers are absent or deleted.

Host-side diagnostics

Connect to the host over SSH to inspect underlying services when commands fail.

Inspecting Incus project status

Verify container states and storage pools in project raft:

HOST: Inspect containers and storage pools
sudo incus --project raft list
sudo incus --project raft storage info raft-data

Checking host systemd units

Raft installs two systemd units on the host:

  • raft-network.service: Applies firewall and bridge isolation rules before Incus starts.
  • raft-expire.timer / raft-expire.service: Runs the expiration worker every 15 seconds to stop expired workspaces.

Inspect unit logs with journalctl:

HOST: Check Raft systemd unit logs
sudo systemctl status raft-network.service
sudo systemctl status raft-expire.timer
sudo journalctl -u raft-expire.service -n 50 --no-pager

Host journals and guest secrets

Host system journals can capture command lines, arguments, and environment variables executed in guest containers. Treat host logs as sensitive records that may contain guest credentials.

Manually triggering garbage collection

The host timer stops expired containers every 15 seconds. Trigger an immediate expiration run from the controller:

CONTROLLER: Run expiration workers manually
raft gc

Expired containers transition from Running to Stopped. Expiration never deletes container files or snapshots.

Common error scenarios

Configuration errors

Missing configuration file error

If ~/.config/raft/incus.json does not exist, the CLI raises FileNotFoundError. Create the configuration directory and file. Set mode 0600 on the file only.

"Configure at least one host in ~/.config/raft/incus.json"

The configuration file contains valid JSON that is {} or is not an object (an empty file raises a JSON decoding error instead). Ensure the file defines at least one location with a non-empty ssh string.

"Build an image and configure its full immutable fingerprint before creating boxes"

The image field is missing or is not a 64-character hexadecimal SHA-256 fingerprint. Supply the full fingerprint from sudo incus --project raft image list raft-dev --format json.

Admission and capacity errors

"Location already holds four boxes"

Raft enforces a hard limit of four saved boxes per location. Both running and stopped containers count toward this limit. Destroy unused boxes with raft destroy <box> before creating, forking, or recovering a box.

Low disk warnings

raft limits warns when free space on host disk or the raft-data pool falls below 5 GiB. This warning is diagnostic and does not halt container operations. Clean up dangling images, old snapshots, or unused containers to free space.

Network and forwarding issues

Port forward connects but returns connection refused

Verify that the guest service binds to 0.0.0.0 or the container's eth0 IP address. Services bound strictly to 127.0.0.1 inside the container cannot receive traffic routed through the bridge.

Workspaces cannot reach other workspaces on the same host

This is intended behavior. Raft configures nftables bridge filtering on rfbr0 to prevent communication between workspaces on the same bridge.

Host crash during recovery

If the controller crashes or SSH drops while raft recover is transferring an archive:

  1. Check the Recovery handle: <location>:<name> output printed to stderr when raft recover started.
  2. Inspect the reported handle with raft list --location <location> and confirm container status before removing staging.
  3. Identify the exact staging file (/tmp/raft-recovery.XXXXXXXX) on the host by owner and timestamp.
  4. After confirming container status, remove only that owned file:
    HOST: Remove verified owned staging file
    sudo -n rm -f -- /tmp/raft-recovery.XXXXXXXX

Do not delete by wildcard

Never run rm /tmp/raft-recovery.*. It can delete files from another recovery.

On this page