NVIDIA
Llama-3.1-Nemotron-Ultra-253B-v1
Container
NVIDIA
Llama-3.1-Nemotron-Ultra-253B-v1

This container houses the Llama-3.1-Nemotron-Ultra-253B-v1, a reasoning model offering a great tradeoff between accuracy and efficiency. Post-trained for chat, RAG, and tool calling, it supports a 128K context and fits on a single 8xH100 node.

NIM MetadataTo verify your workspace, run the Verify CLI locally and compare the generated hash with the NGC-published hash shown here. Learn more in the NVIDIA Documentation.
GPU Type
GPU Type
  • GPU Count
    GPU Count
  • Precision
    Precision
  • Profile
    Profile
  • LoRA
    LoRA
  • Engine
    Engine
  • TP
    TP
  • 17 Instances
    Columns
  • Profile
    LoRA
    Engine
    Profile ID
    Actions
    H200 NVL233b:10debf16latencytensorrt_llm8
    H100 80GB HBM32330:10debf16throughputtensorrt_llm8
    H200 NVL233b:10defp8throughputtensorrt_llm8
    H2002335:10defp8latencytensorrt_llm8
    H100 80GB HBM32330:10defp8latencytensorrt_llm8
    H200 NVL233b:10defp8latencytensorrt_llm8
    H200 NVL233b:10debf16throughputtensorrt_llm8
    B2002901:10defp8throughputtensorrt_llm8
    H100 NVL2321:10debf16latencytensorrt_llm8
    H100 NVL2321:10defp8throughputtensorrt_llm8