All workflows

Reviewed guide · no hardware test recorded

Chat with a model

Plan a single-GPU DeepSeek-R1-Distill-Qwen-7B BF16 session with the existing vLLM preview profile.

Version 2026-09-27.1

chat-deepseek-7b · reviewed 2026-09-27 · review due 2026-10-27

Requirements

Total download
Unknown
Host disk
Unknown
Host RAM
Unknown
Per-device VRAM
Unknown
Task details and compatibility

Model shard sizes are recorded in the model registry; total environment download and host requirements are unknown. Use the shared model preview for conditional memory arithmetic.

precision
BF16
inputTokens
2048
outputTokens
512
concurrency
1

Changing task dimensions invalidates a tested-fit claim for a different task. This version has 0 recorded successful runs.

  • vLLM: 0.11.0. Image digest, flags and support requirements are owned by the shared runtime profile. Runtime overhead and host bounds remain unverified.

Artifacts and prerequisites

DeepSeek-R1-Distill-Qwen-7B

Revision: 916b56a44061fd5cd7d6a8fb632557ed4f724f60

MIT; review upstream model terms. Public checkpoint. Use the pinned revision and tokenizer from the shared model registry.

Shared runtime: vllm-0.11.0-qwen2-bf16-single-gpu-preview-v1. Pinned image: vllm/vllm-openai@sha256:d8d39b59e909d2378ac4feeb191f7e7b6f1342477dc66b7c47cec89e9985ad8a.

Shared model variant: deepseek-r1-distill-qwen-7b-bf16@916b56a44061fd5cd7d6a8fb632557ed4f724f60. Open memory arithmetic preview

Set up the task

On-demand GPU Pod; verify driver and image support before renting.

  1. Review the pinned checkpoint license and vLLM runtime profile. Check GPU driver, CUDA, host architecture and available host memory.
  2. In the provider console, choose one on-demand GPU Pod. Select the pinned vLLM image from the shared runtime profile and configure the persistent mount.
  3. Download the pinned checkpoint revision into the persistent cache. Serve that local snapshot with the profile's eager execution, cache and context settings.
  4. Bind the service privately or use an authenticated tunnel; verify a short response before extending the workload. A retained cache still requires loading and warmup after restart.

Enter your own cold-download, loading and warmup allowance. Warm restart times have not been measured.

Save outputs and shut down

  1. Save outputs and runtime settings outside temporary container storage.
  2. Copy outputs to independent storage and verify the files.
  3. Stop the Pod in the provider console, then review continuing storage charges.

NVGPU does not provision, stop or delete resources.

Plan before you launch

Limits and compatibility
  • Reviewed guide; no successful hardware run recorded.
  • Model/runtime preview remains separately gated. A catalog GPU is a candidate, not a verified fit.
  • No multi-GPU sharding, offload or MoE fit claims.
Sources and review evidence

Official source: Pinned model identity; metadata is shared with the model directory.

Official source: Manual Pod setup and lifecycle actions.