Troubleshooting
The provider refuses to start without credentials
The provider no longer silently falls back to the fake: with --oci-backend=auto
(the default) and both --region and --compartment unset, it exits with a
clear error rather than coming up on a simulation. Set both flags (and --subnet,
--image) for the real backend. To run the credential-free in-memory fake on
purpose (testing/conformance only), pass --use-fake-backend.
Machines land in FAILED shortly after Create/Configure
A machine moves to FAILED with last_error when an actuator errors or a
transition overruns its timeout. Check:
Createfailing:LaunchInstanceerrors (quota/limits for the shape or AD, bad subnet/image OCID, missingmanage instance-family/use virtual-network-family/read instance-images). The error is inlast_errorandbigfleet_oci_api_calls_total{op="LaunchInstance",outcome="error"}.Createtiming out: the instance didn’t reach RUNNING within the Create timeout (8m). Look at the instance in the OCI Console for a provisioning error.Configurefailing: the Run Command failed — the base image isn’t running the Oracle Cloud Agent Run Command plugin, the bootstrap hook exited non-zero, or the principal lacksuse instance-agent-command-family.Drainfailing:kubectl cordon/drainreturned non-zero on the node (e.g. a PDB that can’t be satisfied within the grace period).
A FAILED machine carries the underlying error in Get(...).last_error.
After a restart, in-flight transitions are FAILED
Expected without durable state. The kit cannot replay a backend actuator (notably
the Configure bootstrap blob) across a restart, so an interrupted transition
surfaces as FAILED (...transition interrupted by a provider restart; needs re-drive) for the shard to re-drive. Enable --state on a PersistentVolume so
inventory, bindings, fence marks, and the idempotency map survive — but in-flight
transitions still resolve to FAILED by design.
Create rejected with FAILED_PRECONDITION
That is fencing, not a fault: a shard sent a token not strictly newer than the
high-water mark (a stale/zombie process). The current shard’s next request, with a
newer (epoch, sequence), is accepted. No action needed.
Inventory looks stale / an orphaned instance appears
Describe/reconcile recovers inventory from the bigfleet-managed=true and
bigfleet-machine-id freeform tags. A managed-but-untagged running instance is
surfaced as an orphan under its OCID so it isn’t lost. If reconcile is erroring
(bigfleet_oci_reconcile_total{outcome="error"}), check API permissions and
throttling; the persisted store is the primary restart path.
A preemptible machine shows a non-zero interruption_probability
That is correct and required — see
Pricing & interruption. A
SPOT machine with 0 would be a correctness bug; the kit rejects such a seed at
startup.
Prices look wrong / stale
price_per_hour is live-refreshed from the public OCI price list on a timer
(--price-refresh, default 45m); prices.yaml (embedded, or --prices-file) is
only the startup seed + fallback. It is a relative ranking signal. Bare-metal
(capacity_type=bare_metal) always reports 0.
- Check freshness:
bigfleet_oci_price_last_success_timestamp_seconds(staleness =time() - this) andbigfleet_oci_price_refresh_total{outcome}. A growingerrorcount or a stale timestamp means refreshes are failing — the provider keeps serving the last live (or seed) prices and logs a warning. - Refresh failing? Confirm egress to the cost-estimator API (or set
--price-list-url). With the fake backend the refresh uses a deterministic, network-free source. - Won’t start, “has no price”: an hourly-billed offering priced at 0 is
rejected (fail closed). Add a
prices.yamlentry / SKU mapping for the shape, or declare the lanecapacity_type=bare_metalif it really is held capacity.