NVIDIA
Llama-3.1-70b-instruct
Model
NVIDIA
Llama-3.1-70b-instruct

The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out).

  • 327 Versions
    trtllmapi-pt-runtime-params-h200-nvlx1-throughput-bf16-gjlpxyjp3wSelectedSigned
    01/15/2026 10:33 PM UTC1.26 KB
    trtllmapi-pt-runtime-params-a10gx8-throughput-lora-bf16--osxuwrg2qSigned
    01/15/2026 10:31 PM UTC1.88 KB
    trtllmapi-pt-runtime-params-a10gx8-throughput-bf16-kpwltkwc0aSigned
    01/15/2026 10:31 PM UTC1.55 KB
    trtllmapi-pt-runtime-params-a10gx8-latency-bf16-oqeqt7-ssaSigned
    01/15/2026 10:31 PM UTC1.54 KB
    trtllmapi-pt-runtime-params-a100x2-throughput-bf16-auh9sbukaqSigned
    01/15/2026 10:31 PM UTC1.25 KB
    bf16-tool-calling-fixSigned
    01/15/2026 10:31 PM UTC131.43 GB
    trtllmapi-pt-runtime-params-rtx6000-blackwell-svx4-throughput-lora-bf16-uyoeebitmwSigned
    12/11/2025 10:33 PM UTC1.78 KB
    trtllmapi-pt-runtime-params-rtx6000-blackwell-svx4-latency-fp8-kttskat-baSigned
    12/11/2025 10:33 PM UTC1.6 KB
    trtllmapi-pt-runtime-params-rtx6000-blackwell-svx2-throughput-bf16-msuidgewgqSigned
    12/11/2025 10:33 PM UTC1.62 KB
    trtllmapi-pt-runtime-params-rtx6000-blackwell-svx4-latency-bf16-h0vjef5olqSigned
    12/11/2025 10:33 PM UTC1.78 KB
    ...