NVIDIA: Llama 3.3 Nemotron Super 49B V1.5
nvidia/llama-3.3-nemotron-super-49b-v1.5
Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context. It’s post-trained for agentic workflows (RAG, tool calling) via SFT across math, code, science, and multi-turn chat, followed by multiple RL stages; Reward-aware Preference Optimization (RPO) for alignment, RL with Verifiable Rewards (RLVR) for step-wise reasoning, and iterative DPO to refine tool-use behavior. A distillation-driven Neural Architecture Search (“Puzzle”) replaces some attention blocks and varies FFN widths to shrink memory footprint and improve throughput, enabling single-GPU (H100/H200) deployment while preserving instruction following and CoT quality. In internal evaluations (NeMo-Skills, up to 16 runs, temp = 0.6, top_p = 0.95), the model reports strong reasoning/coding results, e.g., MATH500 pass@1 = 97.4, AIME-2024 = 87.5, AIME-2025 = 82.71, GPQA = 71.97, LiveCodeBench (24.10–25.02) = 73.58, and MMLU-Pro (CoT) = 79.53. The model targets practical inference efficiency (high tokens/s, reduced VRAM) with Transformers/vLLM support and explicit “reasoning on/off” modes (chat-first defaults, greedy recommended when disabled). Suitable for building agents, assistants, and long-context retrieval systems where balanced accuracy-to-cost and reliable tool use matter.
Model specifications
- Input
- text
- Output
- text
- Context
- 131,072 tokens
- Max output
- 131,072 tokens
- Input price
- $0.1 / 1M tokens
- Output price
- $0.4 / 1M tokens
- Released
- 2026-03-19
Capabilities
- Streaming
- Function calling
- Vision
- JSON mode
- Playground
Provider pricing, discounts and data privacy
Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.
Standard service tier
1 available provider · tier input average $0.1 / 1M tokens · tier output average $0.4 / 1M tokens
DeepInfra
Tier: Standard · Region: US · Quantization: fp8
Pricing
- Input
- $0.1 / 1M tokens
- Output
- $0.4 / 1M tokens
No provider discount is currently published.
Data privacy and compliance
- Region
- US
- Zero data retention
- Yes
- Data retention
- Zero retention
- Used for training
- No
- Data collection
- Moderated
- No
- GDPR compliant
- Yes
- HIPAA compliant
- No
- SOC 2 certified
- Yes
- BYOK supported
- Yes
Privacy policy · Terms · Official website · Documentation · Status · Support
Frequently asked questions
- What is NVIDIA: Llama 3.3 Nemotron Super 49B V1.5?
- Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context. It’s post-trained for agentic workflows (RAG, tool calling) via SFT across math, code, science, and multi-turn chat, followed by multiple RL stages; Reward-aware Preference Optimization (RPO) for alignment, RL with Verifiable Rewards (RLVR) for step-wise reasoning, and iterative DPO to refine tool-use behavior. A distillation-driven Neural Architecture Search (“Puzzle”) replaces some attention blocks and varies FFN widths to shrink memory footprint and improve throughput, enabling single-GPU (H100/H200) deployment while preserving instruction following and CoT quality. In internal evaluations (NeMo-Skills, up to 16 runs, temp = 0.6, top_p = 0.95), the model reports strong reasoning/coding results, e.g., MATH500 pass@1 = 97.4, AIME-2024 = 87.5, AIME-2025 = 82.71, GPQA = 71.97, LiveCodeBench (24.10–25.02) = 73.58, and MMLU-Pro (CoT) = 79.53. The model targets practical inference efficiency (high tokens/s, reduced VRAM) with Transformers/vLLM support and explicit “reasoning on/off” modes (chat-first defaults, greedy recommended when disabled). Suitable for building agents, assistants, and long-context retrieval systems where balanced accuracy-to-cost and reliable tool use matter.
- How much does NVIDIA: Llama 3.3 Nemotron Super 49B V1.5 cost?
- Input costs start at $0.1 / 1M tokens and output costs start at $0.4 / 1M tokens. Provider-level prices vary by service tier.
- What is the context length of NVIDIA: Llama 3.3 Nemotron Super 49B V1.5?
- NVIDIA: Llama 3.3 Nemotron Super 49B V1.5 supports a 131,072 token context window and up to 131,072 output tokens.
- What capabilities does NVIDIA: Llama 3.3 Nemotron Super 49B V1.5 support?
- NVIDIA: Llama 3.3 Nemotron Super 49B V1.5 supports Streaming, Function calling, Vision, JSON mode, Playground.
- Which providers offer NVIDIA: Llama 3.3 Nemotron Super 49B V1.5?
- NVIDIA: Llama 3.3 Nemotron Super 49B V1.5 is available from DeepInfra.
- How do providers handle data privacy for NVIDIA: Llama 3.3 Nemotron Super 49B V1.5?
- 1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.