Skip to content

[CON-1511][CON-1510][CON-1624][CON-1661] Align self-test ports, B300 RAM, CUDA diagnostics, and reliability override - #458

Open
jjziets wants to merge 23 commits into
vast-ai:masterfrom
jjziets:CON-1511-self-test-port-range
Open

[CON-1511][CON-1510][CON-1624][CON-1661] Align self-test ports, B300 RAM, CUDA diagnostics, and reliability override#458
jjziets wants to merge 23 commits into
vast-ai:masterfrom
jjziets:CON-1511-self-test-port-range

Conversation

@jjziets

@jjziets jjziets commented Jul 16, 2026

Copy link
Copy Markdown
Contributor

Maintainer action

Merge this PR only after the production image family from
vast-ai/self-test#7 has been
published and its immutable digests verified. Hosts can consume root vast.py
directly from master, so image publication must come first.

Summary

  • Test the complete configured direct-port range over TCP and UDP; there is no
    artificial 64-port cap. DEFAULT_SCAN_WORKERS = 64 limits concurrency only.
  • Remove the obsolete excessive-port warning.
  • Apply min(95% of aggregate VRAM, 2,000,000 MiB) to the system-RAM check in
    both packaged vastai and deprecated root vast.py, fixing 8x B300 hosts.
  • Select the contract-aware vastai/test:self-test-cli-1.2.4-cuda-* production
    family in both launchers.
  • Preserve actionable CUDA failure markers and recursively redact sensitive
    keys in structured support-bundle JSON.
  • Support exact offer pinning and durable exact-instance-ID handoff for safer
    external cleanup.
  • Describe candidate overrides as immutable digest references on both CLI
    surfaces.

Release pair, image names, and pipelines

Companion image/runtime PR:
vast-ai/self-test#7

The repositories have separate release pipelines:

  1. vast-ai/self-test validates, builds, and publishes the Docker images.
  2. vast-ai/vast-cli selects a production image for the host CUDA family,
    launches it, and formats the result. The packaged CLI is released through
    Git/PyPI separately; root vast.py becomes consumable from master.

New images use exactly two tag forms:

  • Production: vastai/test:self-test-cli-<contract>-cuda-<runtime>
  • Candidate: vastai/test:self-test-cli-<contract>-candidate-<12-lowercase-source-SHA>-cuda-<runtime>

For example:

  • Production: vastai/test:self-test-cli-1.2.4-cuda-12.8
  • Candidate: vastai/test:self-test-cli-1.2.4-candidate-0123456789ab-cuda-12.8

vastai/test is the historical Docker Hub repository name, not a staging
channel. cli-1.2.4 is the minimum compatible CLI/image contract, not the PyPI
package version. Historical self-test-cu128, self-test-v2-*, ticket, PR,
feature-prose, and dogfood tags remain intact for compatibility and audit
history; they are not used for new releases.

Released CLI code selects only production tags. Candidate images must be
provided explicitly through --test-image or VAST_SELF_TEST_IMAGE, preferably
as vastai/test@sha256:<OCI-index-digest>.

PR and branch CI validates source without publishing an image. Image build and
publication require a manual self-test workflow dispatch. Merging either source
PR alone does not publish Docker or PyPI artifacts.

Current coordinated source heads:

  • CLI: 9bac1688ac81ed2a9debf981952f8ba39e63a137
  • Self-test: 1150fb6160c15eab0ae50f3aed813bb63531601e
  • CLI/image contract: 1.2.4

Validation

  • Workflow-equivalent local CLI/API/SDK/script suite: 870 passed.
  • Python compilation, Graphify refresh, and git diff --check pass.
  • CON-1661 local CLI/legacy validation passes 635 tests; the current-head
    cross-platform run is linked by the PR checks below.
  • Current-head Ubuntu, macOS, and Windows unit/integration and install-smoke
    validation passed in
    run 32760054253.
  • Self-test current-head four-CUDA validation passed in
    run 32760097650.

Exact paid evidence and limits

Current CON-1661 exact-head qualification (2026-08-24):

  • CLI source: 9bac1688ac81ed2a9debf981952f8ba39e63a137.
  • Self-test source: 1150fb6160c15eab0ae50f3aed813bb63531601e.
  • Contract: 1.2.4.
  • Candidate workflow: run 32760671594.
  • Candidate tag: vastai/test:self-test-cli-1.2.4-candidate-1150fb6160c1-cuda-12.8.
  • OCI index: sha256:4d31e29d4a18802192b4bbc406648bf78d0285ff2e91704d312523eba907c157.
  • Linux/amd64 manifest: sha256:dedc88d9d2d9284fc0d47f76c065c4c8d0c6a2e0f8acaa2f3f9b61a56571f44b.
  • Paid client: macOS 26.5 (25F71), arm64, Python 3.12.12. Linux and Windows source CI passed, but no paid live run was made from those clients.
  • Target: machine 38273, pinned offer 34224259, 1x RTX 3090, reliability 0.7125712, instance 48590573.
  • The no-flag preflight stopped before rental with reliability as its only failed hard check. With --ignore-reliability, reliability remained status=fail, ignored=true, and ignored_by=--ignore-reliability; every other hard preflight passed.
  • The digest-pinned image workload passed system requirements, ResNet18, ECC, NCCL, and simultaneous stress-ng/gpu-burn; the CLI returned success=true, stage=complete.
  • Exact cleanup succeeded and final rental inventory was empty. Account delta/cost: $0.007533352779987013.
  • The successful CLI auto-cleanup removed the container before its structured /progress.jsonl cli_contract event was separately retained. The live result records diagnostics.reliability_ignored=true and the exact image digest; source tests cover image-side ignore_reliability=true, and the paid workload proves the image contract accepted the launch and kept runtime checks enabled.

Paid diagnostic evidence predates the current source head:

  • The machine-145648 cudaErrorTooManyPeers run used CLI
    b3bcd054dc418b194c7ac0a7c146ccc31af4a9cf, self-test
    550ee9bddbef23cf480e88ddd2e80df89bb1d824, and image index
    sha256:998d828782dad8436803ad5857f86f00b8d8c34c94430dd88e32cd1d383aa69f.
  • The machine-41526 cudaErrorContained run used CLI
    f1f4dca4fa6ea0708b0bf61124cc526fb1bf4f7e, self-test
    ad33369ab08f2e18ccd781aab6a999efd1b9f9f8, and image index
    sha256:7b23e4e8619edfe7dde3c54ba2271b8065f4b16a505ee3b8a1084a54470f69e9.
  • Self-test 650f437, index
    sha256:7b1e87fc16438d81391431dcfa15f4e811bd07cb0224f2f559d203f09956d81d,
    was handed unchanged to paid instance 47550175, but image initialization
    timed out before the workload began. That run proves the launch and cleanup
    path, not live rendering of the latest diagnostic.

Current CLI head adds tested parser, support-bundle, offer pinning,
exact-instance handoff/cleanup, and naming guidance after those rentals. It is
not described as an exact-head paid diagnostic run.

The completed B300 pass also used earlier exact heads: CLI b3bcd054, image
88d26e1, on a 2x B300 machine. The current pair retains and regression-tests
those B300 formulas and paths; no new exact-head B300 rental is claimed.

Release order

  1. Validate self-test PR License #7 and manually publish an exact-SHA candidate through
    the candidate channel.
  2. Record its OCI index/platform digests and qualify it by immutable digest.
  3. Merge self-test PR License #7 after review.
  4. From reviewed main, manually dispatch Self-test Images with
    channel=production, cuda_version=all, and push=true.
  5. Verify provenance and immutable digests for all four production tags.
  6. Only then merge this CLI PR, because root vast.py can be consumed directly
    from master.
  7. Publish the packaged CLI through the normal Git/PyPI release pipeline.

At present both PRs remain open, none of the four production tags has been
published, and no PyPI release contains this coordinated CLI mapping.

@jjziets

jjziets commented Jul 16, 2026

Copy link
Copy Markdown
Contributor Author

CI note: the Ubuntu/macOS unit jobs fail in the pre-existing test with for the temporary path; the test fails before any CON-1511 code path. Windows unit/integration and all three install-smoke jobs pass. The focused CON-1511 suite passes locally (105 tests).

@jjziets

jjziets commented Jul 16, 2026

Copy link
Copy Markdown
Contributor Author

Correction to the CI note: Ubuntu/macOS unit jobs fail in the pre-existing cli/test_update_command.py::TestPerformUpdate::test_prune_failure_does_not_fail_the_update test with FileNotFoundError for the temporary env/bin/vastai path. The failure is unrelated to CON-1511. Windows unit/integration and all three install-smoke jobs pass; the focused CON-1511 suite passes locally with 105 tests.

@jjziets
jjziets force-pushed the CON-1511-self-test-port-range branch from 4fca02f to 6756c92 Compare July 16, 2026 16:19
@jjziets jjziets changed the title [CON-1511] Scan the host-configured port range in self-test [CON-1511][CON-1510] Scan host port range and preserve CUDA diagnostics Aug 3, 2026
@jjziets

jjziets commented Aug 3, 2026

Copy link
Copy Markdown
Contributor Author

The paid 10x A4000 validation of the coordinated image exposed one companion-CLI presentation gap, fixed in f493919d63e9da151e11e27b5d0b0681d3d2c49b. The final PR head is 842ec1c0734bf108c4ea730dd57d3f0e50a4bcc6.

Captured paid behavior

  • Machine/offer/instance: 145648 / 46083026 / 46695399
  • Hardware: 10x RTX A4000, bw_nvlink=0
  • Self-test PR/head: vast-ai/self-test#7 at 550ee9bddbef23cf480e88ddd2e80df89bb1d824
  • Image: vastai/test:self-test-cli-1.2.3-con1510-runtime-550ee9b-cuda-13.0
  • Build/provenance: https://github.com/vast-ai/self-test/actions/runs/30806952741

The image correctly emitted:

SELF_TEST_FAILURE[cuda_error_too_many_peers]: All-GPU ResNet18 failed on all 10 visible GPUs. Interpretation: PyTorch DataParallel exhausted CUDA peer mappings (cudaErrorTooManyPeers); GPU 0 may need 9 peers, while non-NVSwitch systems support up to 8 per device. Result: The advertised all-GPU configuration cannot be verified. Next: Split the offer into smaller GPU groups that pass self-test or use an NVSwitch-capable topology, then rerun. Traceback: The full traceback is retained in the container logs.

The CLI revision used for that paid run (b3bcd054) still stored generic failure_code: resnet_failed, left the stage at system_requirements, and echoed the full marker through multiple console paths.

CLI follow-up

This commit:

  • allowlists and preserves all nine stable ResNet failure codes emitted by self-test License #7;
  • recognizes the current ResNet, ECC, NCCL, and stress/gpu-burn stage strings;
  • records this capture as cuda_error_too_many_peers at stage resnet;
  • prints the full marker once while preserving the exact raw line in reason, underlying_error, and support-bundle data; and
  • rejects unknown markers and catalog-known codes that are not valid image markers.

Validation:

  • Focused parser/machine tests: 126 passed.
  • Standard local macOS CLI/API/SDK suite: 601 passed.
  • Python compilation and git diff --check: passed.
  • Independent review: ready, with marker allowlisting and console dedup verified.
  • GitHub Ubuntu/macOS/Windows unit/integration and install-smoke checks: passed in https://github.com/vast-ai/vast-cli/actions/runs/30809455506.

The first macOS run had one unrelated wall-clock-only failure in the existing port-range concurrency test: it reached the expected four active workers but took 0.338s against a brittle 0.2s cutoff. Final head 842ec1c removes only that timing assertion; the deterministic max_active == 4 concurrency guarantee and separate hard-deadline test remain. The fresh three-platform run is green.

No second paid rental was started for this parser/presentation-only follow-up. The live image behavior, traceback retention, cleanup, and approximately $0.05220 visible cost delta are captured in self-test issue #8.

@jjziets jjziets changed the title [CON-1511][CON-1510] Scan host port range and preserve CUDA diagnostics [CON-1511][CON-1510][CON-1624] Align self-test ports, B300 RAM, and CUDA diagnostics Aug 10, 2026
@jjziets
jjziets marked this pull request as ready for review August 10, 2026 09:01
@jjziets

jjziets commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

@robballantyne This is ready for review at exact head 28357a5. It is current with master, conflict-free, and run 31372018919 is green for unit/integration and install-smoke on Ubuntu, macOS, and Windows. The maintainer checklist and evidence limits are now at the top of the PR. Release order matters: merge self-test PR #7 and publish/verify the production self-test-cli-1.2.3-cuda-* images before merging/releasing this CLI mapping.

@jjziets

jjziets commented Aug 12, 2026

Copy link
Copy Markdown
Contributor Author

Naming and evidence clarification:

Future self-test images now have two release channels only:

  • self-test-cli-<contract>-cuda-<runtime> for production
  • self-test-cli-<contract>-candidate-<12-lowercase-source-SHA>-cuda-<runtime>
    for qualification

Candidate runs should pin repository@sha256:digest; ticket numbers, PR
numbers, feature prose, dogfood, and v2 are not release channels.
Historical tags remain intact for compatibility and audit history.

The image and CLI also have separate release pipelines: self-test publishes the
container, while vast-cli selects it and is released independently through
Git/PyPI. Production images must exist and have verified digests before the CLI
mapping is merged or released.

Evidence correction: the paid cudaErrorTooManyPeers result belongs to
self-test 550ee9b; the paid cudaErrorContained result belongs to
ad33369. The later 650f437 image was built and provenance-verified, but its
paid attempt timed out during image initialization before the workload ran.
Therefore we do not claim that the exact latest image completed a paid
diagnostic workload.

This clarification did not merge either PR, publish any production image tag,
or release this work to PyPI.

@jjziets jjziets changed the title [CON-1511][CON-1510][CON-1624] Align self-test ports, B300 RAM, and CUDA diagnostics [CON-1511][CON-1510][CON-1624][CON-1661] Align self-test ports, B300 RAM, CUDA diagnostics, and reliability override Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant