OPBDH

“Of course I can do it, Dave.”

What it does

Open the Pod Bay Door, Hal: launches a RunPod or Prime Intellect GPU pod, runs your model-backed script on it, brings your results home, and shuts the pod down.

For supervised training, opbdh ft turns human-editable examples into automatically sized LoRA, QLoRA, or Full jobs—including multi-GPU runs—while keeping the same cost and cleanup guarantees.

One opbdh launch performs the whole mission. Click a step to watch it in HAL's console →

  1. Verify. Your code is checked statically — Python byte-compiles, shell scripts pass bash -n — before any money is spent.▸ opbdh verify
  2. Select. GPU candidates are picked from a VRAM requirement and an optional hourly-price cap.▸ opbdh gpus
  3. Cache. Optionally create or attach a network volume sized from the model’s actual weight files, so downloads survive across runs.▸ opbdh models size
  4. Launch. The pod is created, your code uploaded, the model pre-downloaded into the cache, your command run.▸ opbdh launch
  5. Monitor. Remote logs/ and results/ sync continuously into a local runpod_results/<run_id>/ directory.▸ live sync
  6. Guard. If the estimated cost of the run crosses max_spend_dollars, the job is stopped.▸ spend guard
  7. Clean up. The pod is deleted when the run finishes — or fails — unless you ask to keep it.▸ teardown

Install

pip install opbdh

Requirements: macOS or Linux, Python ≥ 3.11, and ssh/scp on your PATH.

Pods are reached over SSH: OPBDH uses your existing ~/.ssh/id_ed25519, id_ecdsa, or id_rsa keypair if one exists, generates a dedicated keypair under ~/.config/opbdh/ssh/ otherwise, or uses whatever you point ssh_key / ssh_public_key at in the config.

Configure

Running bare opbdh while unconfigured starts a first-run setup wizard: provider, API token check, and defaults for model, code path, VRAM, and price caps.

Set credentials for your provider:

export RUNPOD_API_TOKEN="..."          # for the default RunPod provider
export PRIME_INTELLECT_API_KEY="..."   # for --provider primeintellect
export HF_TOKEN="..."                  # optional, for private/gated HF models

Create a config interactively:

opbdh config wizard

Or write one directly:

opbdh config write --model Qwen/Qwen2.5-0.5B-Instruct \
  --code "{cwd}/run.py" --vram-gb 24 --max-spend 5

Config is merged from the global ~/.config/opbdh/config.json and a local opbdh.json or .opbdh.json (discovered upward from the working directory), with local values overriding global ones and command-line flags overriding both.

String values support placeholders — {cwd}, {model_id}, {model_slug}, {run_id}, {timestamp}, {config_dir} — and environment variables ($VAR / ${VAR}).

Run

The guided way:

opbdh run wizard

The direct way:

opbdh run now ./run.py \
  --model Qwen/Qwen2.5-0.5B-Instruct \
  --vram-gb 48 \
  --max-dollars-per-hour 2 \
  --max-spend 5

Or the short alias:

opbdh launch ./run.py --model Qwen/Qwen2.5-0.5B-Instruct --dry-run

--dry-run verifies your code and prints the full plan without contacting RunPod. Without --yes, every real launch shows the estimated hourly price and asks for confirmation before any billable compute starts.

Inspecting before you fly

opbdh plan ./run.py --model Org/Model     # show the plan for a run
opbdh verify ./run.py                     # static checks only
opbdh gpus --vram-gb 48                   # GPU candidates + price estimates
opbdh gpus --vram-gb 48 --provider primeintellect   # live PI offers
opbdh models search qwen                  # search Hugging Face models
opbdh models size Qwen/Qwen2.5-0.5B-Instruct   # weight size + suggested volume

Fine-tune

Open the OPBDH:FT guide →

Install the optional fine-tuning extra and run opbdh ft to open an interactive SFT console:

pip install "opbdh[ft]"
opbdh ft

Load TOML, JSON, or JSONL files—or create examples in the console—then organize them with tags and train all or selected groups. The native TOML format handles both plain prompt/completion examples and chat conversations while staying comfortable to edit by hand.

input = "Translate hello to Italian"
output = "Ciao"
tags = ["translation", "short"]

One dataset can drive any number of named LoRA, QLoRA, or Full recipes. Each recipe keeps its own model, technique, selected tags, hyperparameters, provider, spend limits, and GPU count. OPBDH sizes the pod, generates a TRL + Accelerate job, syncs the artifacts home, and keeps the dataset and recipe metadata under .opbdh/ for the next run.

opbdh ft --data ./examples --model Qwen/Qwen3-8B \
  --model-type chat --recipe qlora-2gpu --method qlora --gpu-count 2

opbdh ft:import openai.jsonl --format openai --tag imported

Prime Intellect

Prime Intellect is a marketplace that aggregates GPU offers from many clouds (Hyperstack, Lambda, Nebius, even RunPod itself). Pass --provider primeintellect to any of plan, run now, launch, or gpus — or set "provider": "primeintellect" in your config — and OPBDH will:

  • Query live availability and pick the cheapest offers that satisfy your --vram-gb / --max-dollars-per-hour (spot offers and out-of-stock configurations are skipped).
  • Register your SSH public key with your Prime Intellect account if it isn't there yet.
  • Provision the pod through whichever underlying cloud made the offer, then run the same verified-bundle lifecycle as on RunPod: upload, run, sync results, spend guard, cleanup.
export PRIME_INTELLECT_API_KEY="..."   # PRIME_API_KEY also works
opbdh launch ./run.py --model Org/Model --provider primeintellect \
  --vram-gb 48 --max-spend 5
  • Pricing is live, not estimated from a static catalog — the plan shows the actual cheapest offers, and the spend guard uses the price of the pod you actually got.
  • Images: Prime Intellect uses named environments instead of Docker tags. OPBDH picks a recent PyTorch + CUDA environment supported by the offer; set image to an environment name (e.g. cuda_12_4_pytorch_2_5) to override. RunPod-style Docker image tags are ignored for this provider.
  • Network volumes are RunPod-only for now — combining them with --provider primeintellect is rejected. Every Prime Intellect run downloads the model to the pod's own disk.
  • cloud_type maps to Prime Intellect's security tiers: SECURE (default) → secure_cloud, COMMUNITYcommunity_cloud, ALL → both.

Model cache & volumes

By default the Hugging Face cache lives on the pod's own disk, which is deleted with the pod — so every run re-downloads the model. For anything bigger than a toy model, attach a RunPod network volume: the cache then lives at a persistent /workspace mount that survives across runs, and subsequent launches skip the download entirely.

Two ways to get one:

  • Attach an existing volume: --network-volume-id <id> (or network_volume_id in config).
  • Let OPBDH create one: --auto-network-volume --network-volume-data-center-id EU-RO-1. If a volume named opbdh-{model_slug} already exists in that data center it's reused; otherwise one is created, sized from the model's actual weight files on the Hugging Face Hub (2.5× the weights, minimum 50 GB — room for revisions and pip cache), or from network_volume_size_gb if you set it. Override the name with network_volume_name.

Both wizards walk you through this and show the suggested size; opbdh models size <model> shows it standalone.

Volumes outlive runs by design — and bill by the GB-month. Repeated runs of the same model reuse the same volume, but OPBDH never deletes one; when you're done with a model, remove its volume in the RunPod console.
A volume pins you to its data center. Pods can only attach volumes in the same data center, so GPU availability is constrained to that location. Pick a data center that reliably stocks the GPUs you want.

Your script's environment

On the pod, your code runs from /opbdh-run/user/ with:

VariableMeaning
OPBDH_MODEL_IDThe configured Hugging Face model id
OPBDH_RESULTS_DIRWrite your artifacts here
HF_HOME, HUGGINGFACE_HUB_CACHE, …Pointed at the pod or network-volume cache, with the model already downloaded if pre_download_model is on
HF_TOKENIf set locally — passed to the job as a session environment variable, never written to disk on the pod
  • A requirements.txt next to your code is pip-installed automatically.
  • Anything your job writes to logs/ and results/ under /opbdh-run is synced back to local runpod_results/<run_id>/ while the run is in progress and again at the end (including on failure).
  • .env files and the usual junk (.git, .venv, __pycache__, node_modules, …) are never uploaded.

Safety rails

  • Confirmation gate: real launches require an explicit yes (or --yes).
  • Spend guard: max_spend_dollars caps the estimated cost of a run, measured from pod creation. The estimate is based on OPBDH's built-in price table, which is intentionally conservative — treat it as a guard rail, not an invoice.
  • Cleanup: pods are deleted when the run completes or fails. On failure you get a short window (default 120 s) to opt into keeping the pod for debugging; set keep_pod_on_success to keep it after successful runs.

Development

pip install -e ".[dev]"
ruff check .
pytest

See RELEASING.md for the release process. Licensed under MIT.