← All model variants

Exact checkpoint / Estimate-only preview

DeepSeek-R1-Distill-Qwen-7B (publisher BF16)

deepseek-ai/DeepSeek-R1-Distill-Qwen-7B · bfloat16 · safetensors

Revision 916b56a44061fd5cd7d6a8fb632557ed4f724f60

Metadata reviewed on 2026-09-27; next review 2026-10-27. Speed not measured. No tested rental configuration.

Model → workload → exact configuration

Find a GPU for DeepSeek-R1-Distill-Qwen-7B (publisher BF16)

Single-GPU text inference. Estimates remain unknown where exact hardware, runtime or charges lack evidence. NVGPU does not reserve stock or stop your instance.

Input includes system instructions, history and tool text. Total context is 8,192 tokens for each request. The current profile supports up to 16,384 tokens and one simultaneous request.

Base and high are your scenarios, not probability bounds. Each launch includes setup, useful work, idle, export and shutdown. Count overlapping startup activities once using their total elapsed allowance. Sessions use one instance at a time.

Session group 1
Export, idle and shutdown stages
Storage copies, transfers and other charges

List each physical copy once. Running, stopped and retained durations for the same copy cannot overlap. Catalog disk capacity does not establish that its price is included. A blank rate uses a current reviewed tariff only for an explicitly selected supported product; its monthly conversion still needs your input.

Include repeated downloads and output exports in the total volume. A free destination tariff does not establish free source-cloud egress.

An explicit zero means you assume no charge. Unknown fees prevent a complete total or within-budget claim.

Confirmed credits and account funding

Ordinary cost and budget checks stay separate from one-time credits. No signup award or eligibility is assumed. Funding is cash placed in an account, not an extra resource charge.

Billing assumption for products without a current policy

A current product-scoped provider policy takes priority. These inputs are assumptions and never change the whole-instance quote.

Runtime allowance assumption

Blank remains unknown. This allowance does not establish hardware compatibility, measured memory or a tested runtime.

Submit your workload to compare exact configurations.
Standalone memory and hourly-rate calculator

Estimate-only preview

Plan with DeepSeek-R1-Distill-Qwen-7B (publisher BF16)

One dense text model on one GPU. Input includes system instructions, conversation history and tool text. Output reserve is added to every simultaneous request.

Total context: 8,192 tokens per request. The initial runtime profile supports at most 16,384 tokens and one request. Larger inputs get an unsupported explanation.

Leave unknown values blank. The runtime allowance applies to each execution phase, in addition to weights and cache. A reserve of the larger of 1 GiB or 10% of usable memory is counted once. Catalog GB and summed multi-GPU memory are not usable-device measurements.

Session schedule and budget assumptions

Startup includes charged download, setup, load and warmup. These stages are sequential in this scenario. Include running/stopped storage, source and destination transfer, taxes and fees in non-compute charges. Shutdown is assumed immediate; NVGPU does not stop your instance.

Submit the workload to see memory requirements and costs.

Artifact and access evidence

DeepSeek derivative of Qwen/Qwen2.5-Math-7B; distilled using DeepSeek-R1 generated examples according to the publisher card. Public ungated metadata and anonymous metadata access observed at review time. This does not establish individual access approval or verify a complete model download. Consult both linked license texts and preserve applicable notices.

Publisher terms: MIT · Upstream terms: Apache-2.0

Gated download: no. Authentication required: no. Access is subject to the linked terms.

Selected weights total 14.19 GiB on disk. Tokenizer, container, runtime workspace and retained data require additional space.

Pinned files and hashes
  • model-00001-of-000002.safetensors · 8,606,596,466 bytes
    SHA-256: 8b27be8d43b81d0cd19b3113ef88caf1a61c2fc602de77e3cce97aedcec184b4
  • model-00002-of-000002.safetensors · 6,624,675,384 bytes
    SHA-256: 1780a9502f56f592cb4d6262c8c76ca8ad36a86b0257e052c672d87f49707e80
Sources and completeness
  • base_model_training_revision
  • locally_verified_weight_checksums
  • runtime_peak_measurements
  • Publisher does not identify the exact base checkpoint training revision; parent repository attribution is known but its training commit remains unknown.
  • SHA256 values are publisher-supplied LFS metadata; no full weights were downloaded or independently hashed.
  • downloadBytes counts only selected safetensors shards; tensorBytes excludes serialization headers. Tokenizer, runtime image and temporary disk requirements belong to the runtime profile.
  • maxPositionEmbeddings describes the publisher configuration limit, not a validated runtime context limit.
  • Metadata review does not establish runtime support, device fit, speed, stock availability or complete cost.
Open this model in the project planner →