NVIDIA
Llama-3.2-3B-Instruct
Container
NVIDIA
Llama-3.2-3B-Instruct

NVIDIA NIM for GPU accelerated Llama-3.2-3B-Instruct inference through OpenAI compatible APIs

  • NIM MetadataTo verify your workspace, run the Verify CLI locally and compare the generated hash with the NGC-published hash shown here. Learn more in the NVIDIA Documentation.
    GPU Type
    GPU Type
  • GPU Count
    GPU Count
  • Precision
    Precision
  • Profile
    Profile
  • LoRA
    LoRA
  • Engine
    Engine
  • TP
    TP
  • 28 Instances
    Columns
  • Profile
    LoRA
    Engine
    Profile ID
    Actions
    H100 80GB HBM32330:10defp8throughput-loratensorrt_llm1
    L40S26b9:10defp8throughput-loratensorrt_llm1
    H2002335:10defp8throughput-loratensorrt_llm1
    A10G2237:10defp16throughput-loratensorrt_llm1
    H100 NVL2321:10defp8throughput-loratensorrt_llm1
    H2002335:10defp16throughput-loratensorrt_llm1
    A100 SXM4 80GB20b2:10defp16throughputtensorrt_llm1
    H100 NVL2321:10defp16throughputtensorrt_llm1
    H100 NVL2321:10defp16throughput-loratensorrt_llm1
    H100 80GB HBM32330:10defp16throughput-loratensorrt_llm1