ci: retry transient registry/DNS flakes at the cargo/npm/pip tool level (#105)
The v0.4.0-era CI failure was a runner network blip — `napi build` invokes cargo, whose fetch of index.crates.io hit "Could not resolve host: index.crates.io" and failed the Node-on-macOS job, forcing a manual re-run. The earlier flake-hardening (setup-node/setup-python + rust-cache retries) only covered toolchain download and cache restore, not the registry fetches inside the actual build/publish steps. Set tool-level network retries as workflow env so every cargo/napi/maturin/ wasm-pack/npm/pip invocation in every job inherits them — including the nested cargo calls inside napi/maturin/wasm-pack: - CARGO_NET_RETRY=10 (default 3): cargo classes DNS-resolve / connect / timeout errors as spurious and retries with backoff; 10 attempts ride out a transient blip instead of failing the job. - CARGO_NET_GIT_FETCH_WITH_CLI=true: more robust git-dep fetches. - npm_config_fetch_retries=5 / maxtimeout=120s: npm ci/install registry retries. - PIP_RETRIES=5 / PIP_DEFAULT_TIMEOUT=120: pip install resilience. Applied to ci.yml, release.yml and bench.yml (the workflows that build). No more manual re-runs for transient registry flakes.
This commit is contained in:
@@ -31,6 +31,17 @@ permissions:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
# Network-flake resilience: retry transient registry/DNS failures at the tool
|
||||
# level so a blip fetching crates.io / PyPI inside any build step (cargo,
|
||||
# maturin, pip) retries automatically instead of failing the job. Cargo treats
|
||||
# "couldn't resolve host" / connect / timeout as spurious and retries with
|
||||
# backoff; 10 attempts ride out a transient DNS blip on a runner.
|
||||
CARGO_NET_RETRY: "10"
|
||||
CARGO_NET_GIT_FETCH_WITH_CLI: "true"
|
||||
npm_config_fetch_retries: "5"
|
||||
npm_config_fetch_retry_maxtimeout: "120000"
|
||||
PIP_RETRIES: "5"
|
||||
PIP_DEFAULT_TIMEOUT: "120"
|
||||
|
||||
jobs:
|
||||
cross-library-bench:
|
||||
|
||||
@@ -16,6 +16,19 @@ permissions:
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
RUSTFLAGS: "-D warnings"
|
||||
# Network-flake resilience: retry transient registry/DNS failures at the tool
|
||||
# level so a blip fetching crates.io / npm / PyPI inside any build step (cargo,
|
||||
# napi, maturin, wasm-pack, npm ci, pip) retries automatically instead of
|
||||
# failing the job and needing a manual re-run. Cargo treats "couldn't resolve
|
||||
# host" / connect / timeout as spurious and retries with backoff; 10 attempts
|
||||
# ride out a transient DNS blip on a runner. Complements the setup-action /
|
||||
# cache retries (which only covered toolchain download + cache restore).
|
||||
CARGO_NET_RETRY: "10"
|
||||
CARGO_NET_GIT_FETCH_WITH_CLI: "true"
|
||||
npm_config_fetch_retries: "5"
|
||||
npm_config_fetch_retry_maxtimeout: "120000"
|
||||
PIP_RETRIES: "5"
|
||||
PIP_DEFAULT_TIMEOUT: "120"
|
||||
|
||||
jobs:
|
||||
rust:
|
||||
|
||||
@@ -18,6 +18,17 @@ permissions:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
# Network-flake resilience: retry transient registry/DNS failures at the tool
|
||||
# level so a blip fetching crates.io / npm inside any build or publish step
|
||||
# (cargo, napi, maturin, wasm-pack, npm) retries automatically instead of
|
||||
# failing the job. Cargo treats "couldn't resolve host" / connect / timeout as
|
||||
# spurious and retries with backoff; 10 attempts ride out a transient DNS blip.
|
||||
CARGO_NET_RETRY: "10"
|
||||
CARGO_NET_GIT_FETCH_WITH_CLI: "true"
|
||||
npm_config_fetch_retries: "5"
|
||||
npm_config_fetch_retry_maxtimeout: "120000"
|
||||
PIP_RETRIES: "5"
|
||||
PIP_DEFAULT_TIMEOUT: "120"
|
||||
|
||||
jobs:
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
Reference in New Issue
Block a user